
Overview ARK-ASR-3B is a 3-billion-parameter multilingual automatic speech recognition model built by Audio8 that combines a Whisper-style audio encoder with an MLP adapter and a Qwen decoder. The model operates at 16 kHz sampling rate and…
View original source — Hacker Noon ↗


