Moein Bagheri

Speech & Audio AI · University of Isfahan

I build AI that listens.

I started out writing music and ended up in machine learning. Now I work on speech: models that recognize what people say and how they say it, starting with Persian–English speech, where today's systems still stumble.

Speak, and watch your voice appear on the right.

CVGitHubContact

idle · synthetic speech mel scale · 60 Hz – 8 kHz · FFT 2048
Speech · PyTorch · MFCC

Speech emotion recognition

1D CNNs that recognize angry, happy, sad and neutral speech in IEMOCAP, tested on speakers the model has never heard.

50%55%60%65%70% unseen speakers random split
Accuracy across six runs. Every speaker-independent run scored below the random split.
65.1% → 58.0% on unseen speakers, same model
Persian music · Language models

Melodies as language

Notes from 41 Persian melodies, quarter-tones (koron) included, treated as words. N-gram models finish incomplete phrases, with smoothing tuned by leave-one-out cross-validation.

1-gram2-gram3-gram 6.94.74.0
Mean perplexity on five unfinished phrases (α = 0.5). Lower is better.
C B C D → C D C D C D …
Retrieval · Local LLM

Local research assistant

A fully local RAG system over four retrieval papers (REALM, DPR, RAG, FiD) that cites the PDF page behind every answer.

  1. PDF pages
  2. chunks
  3. MiniLM embeddings
  4. FAISS search
  5. Qwen2.5 1.5B
  6. answer + page refs
Everything runs on one laptop through Ollama. No cloud API.
6 research tools · CLI and Streamlit

Now building: Persian–English code-switched speech recognition, my final-year project.

The path, as a piano roll

Each note is a course. Blue notes lead to speech and audio; the gold line under a note is its grade, like MIDI velocity.

Swipe sideways to see earlier semesters.

A selection of courses from my B.Sc. in Computer Engineering at the University of Isfahan, grouped by how they lead to speech and audio AI. Grades are out of 20.

Before code, there was music.

I played classical piano and produced music in FL Studio through high school, and released an album. Somewhere in that process I realized that what I loved most was building things. Speech and audio AI is where the two meet.

"Lost" is the track I'm still proudest of.

Open in Spotify ↗

Let's talk.

I'm applying for research-based MSc programs in speech and audio machine learning, starting Fall 2027. If you work on speech recognition, low-resource languages or machine listening, I'd be glad to hear from you.

moeinbagheri.ai@gmail.com Write to me

CVCV as PDFGitHub