Bangla & English ASR Studio — Whisper transcription, evaluation and fine-tuning

A local-first speech-to-text studio and Whisper fine-tuning pipeline for Bangla and English: transcription UI, batch jobs, WER/CER benchmarking and guarded LoRA training.

Bangla & English ASR Studio — project screenshot

The problem

Off-the-shelf Whisper models work well in general, but production Bangla and mixed Bangla/English (Banglish) audio raises practical problems: Bangla can be misdetected as another Indic language, and mixed-language speech needs controlled decoding and repeatable evaluation.

Fine-tuning without a fixed baseline creates a grey area — there is no evidence that the new model is actually better than the one it replaces.

Architecture and approach

  • A React operator UI for microphone and file transcription, batch jobs, benchmarks, diagnostics and training controls.
  • A FastAPI backend that handles model loading, audio normalisation, job tracking, exports, evaluation and training subprocesses.
  • CLI scripts for the same operations, so GPU servers and automation do not depend on the UI.
  • Docker Compose stacks for CPU, development with hot reload, and an NVIDIA GPU override.

Key features

  • Benchmarking against ground-truth metadata with word error rate (WER) and character error rate (CER), plus per-sample predictions.
  • Bangla/English decoding profiles and post-processing guards.
  • Training and evaluation directly from Parquet datasets (e.g. SUBAK.KO from Hugging Face) without extracting every clip to WAV, including a streaming mode.
  • LoRA/QLoRA fine-tuning of Whisper, training small adapter layers instead of all model weights.
  • Resumable segmented dataset downloads and a persistent model cache that survives container rebuilds.

Engineering decisions

  • Guarded training workflow: back up the current model, evaluate it on a fixed held-out test set, train, evaluate the new model on the same set, then write a side-by-side comparison with confusion analysis.
  • The pipeline never auto-replaces the production model. The comparison reports a verdict (new better, old better, or mixed metrics) and the final decision stays manual.
  • The test split is kept fixed and untouched so old-vs-new comparisons stay fair.

Technology stack

Speech & ML
  • Whisper
  • faster-whisper
  • Hugging Face Transformers
  • LoRA / QLoRA
  • PyTorch
Backend
  • Python
  • FastAPI
  • FFmpeg
Frontend
  • React
  • Vite
Infrastructure
  • Docker Compose
  • NVIDIA CUDA (GPU override)