Hi, I’m Jiaqi.

I’m an AI researcher and engineer specializing in automatic speech recognition (ASR). My work focuses on helping machines understand spoken language, from speech and audio representation learning to efficient, production-ready recognition systems.

More recently, I have been exploring speech generation, audio tokenization, and multimodal voice models. I am especially interested in how modern generative models can produce speech that is natural, expressive, controllable, and responsive in real time.

I enjoy taking complex ideas apart and rebuilding them from first principles—connecting the equations in a paper to architectural choices, training recipes, and real-world engineering trade-offs. I care about systems that are not only capable, but also understandable, reproducible, and useful.

What I work on

  • Automatic speech recognition
  • Speech and audio representation learning
  • Efficient inference, model compression, and deployment
  • Speech generation, audio tokenization, and multimodal voice models

About this blog

This blog is my working notebook for technical deep dives, research notes, and ideas worth thinking through. I write to clarify my own understanding and to make emerging AI systems easier for others to follow—from the intuition behind a method to the details that determine whether it works in practice.

You can also find me on GitHub.