Mengzhe Geng
Research Scientist National Research Council Canada
My research develops and evaluates machine learning for speech, language, audio, and multimodal AI, including accessible and low-resource speech, speaker adaptation, foundation models, generative and reasoning systems, and trustworthy evaluation.
Browse publications Full Google Scholar profile
Google Scholar metrics, last refreshed October 7, 2026
I am open to opportunities in both academia and industry. I can work in Canada, Hong Kong, and Mainland China without needing to apply for an additional work visa. If you know of a good fit, contact me.
Experience
Research Scientist
Digital Technologies, National Research Council Canada · Ottawa, Canada
Education
Ph.D. in Systems Engineering and Engineering Management
The Chinese University of Hong Kong
B.Sc. in Mathematics and Information Engineering
The Chinese University of Hong Kong · First-class honours · ELITE Stream graduate · Minor in Computer Science
Research
Accessible speech recognition
Speech recognition and severity-aware evaluation for dysarthric and older-adult speech under limited data.
Evidence-aware and trustworthy speech AI
Auditable decisions, deepfake detection, and evaluation methods that make speech systems easier to inspect and trust.
Efficient speech foundation models
Evaluation, compression, quantization, and adaptation of speech foundation models for practical deployment, including low-resource languages.
Source-grounded speech generation
Spoken agents, audio language models, and speech generation systems that plan, produce, and evaluate audio with explicit evidence.
Reasoning and evaluation for generative AI
Reasoning systems and evaluation frameworks for generative and multimodal AI in public-sector and research settings.
Author-led publications
These include publications for which I am a first author, co-first author, or (joint) corresponding author.
24 publications
Controlled Acquisition and Abstention in Three-Channel Score Conflicts
Multimodal decision-making
When Does a Spoken Agent Have Enough Evidence to Act? The PACT-SLM Contract Test
Spoken agents
VoxReason: Auditing Source-Grounded Speech Plans Before Synthesis
Spoken agents
SURE-Challenge: Evaluating Speech Evidence Before Speech-LLM Generation
Speech language model evaluation
Auditing Generative Audio Calls for Known-Task Audio-LLM Evaluation
Audio language model evaluation
From Scores to Evidence: Auditable Decisions Can Improve Speech Deepfake Detection
Speech deepfake detection
Text Scores Can Miss Waveform Use: A Qwen2-Audio Quantization Case Study
Multimodal audio evaluation
GlossMATE: Multi-Agent Translator Explanations for Glosses
Language technology
Exploring SSL Discrete Tokens for Multilingual Automatic Speech Recognition
Multilingual speech recognition
V2A-DPO: Omni-Preference Optimization for Video-To-Audio Generation
Video-to-audio generation
Towards Personalized Federated Learning for Dysarthric Speech Recognition
Federated learning
Evaluating Speech Foundation Models for Automatic Speech Recognition in the Low-Resource Kanyen’kéha Language
Indigenous language speech recognition
Supporting SENĆOŦEN Language Documentation Efforts with Automatic Speech Recognition
Indigenous language speech recognition
Homogeneous Speaker Features for on-the-Fly Dysarthric and Elderly Speaker Adaptation and Speech Recognition
Speaker adaptation
Exploring Generative AI Techniques in Government: A Case Study
Generative AI in government
Regularized Federated Learning for Privacy-Preserving Dysarthric and Elderly Speech Recognition
Federated learning
Exploring Cross-Utterance Speech Contexts for Conformer-Transducer Speech Recognition Systems
Speech recognition
Use of speech impairment severity for dysarthric speech recognition
Dysarthric speech recognition
On-the-fly feature based rapid speaker adaptation for dysarthric and elderly speech recognition
Speaker adaptation
Speaker Adaptation Using Spectro-Temporal Deep Features for Dysarthric and Elderly Speech Recognition
Speaker adaptation
Spectro-Temporal Deep Features for Disordered Speech Assessment and Recognition
Disordered speech assessment
Recent Progress in the CUHK Dysarthric Speech Recognition System
Dysarthric speech recognition
Adversarial Data Augmentation for Disordered Speech Recognition
Adversarial data augmentation
Investigation of Data Augmentation Techniques for Disordered Speech Recognition
Disordered speech recognition
View the full publication list All publications on Google Scholar
Selected awards and honours
- Honourable Mention, Outstanding Achievement Award (OAA), NRC Inclusive Innovation Award, National Research Council Canada, 2026
- Award for Excellence in Inclusion, Diversity, Equity and Accessibility, Digital Government Community Awards, Government of Canada, 2026
- Valedictorian, CUHK Postgraduate Class of 2023
- IEEE ICASSP Outstanding Reviewer, 2023
- ISCA INTERSPEECH Travel Grant, 2023
- Finalist, Hong Kong X Foundation FYP+ Project, 2019
- CUHK Academic Excellence Scholarship for Non-local Students, 2019
- CUHK Best Project Award for Undergraduate Research Summer Internship, 2018
- CUHK ELITE Stream Student Scholarship, 2019 and 2016
- CUHK S.H. Ho College Outstanding Student Scholarship, 2018, 2017, 2016, and 2015
- HKSAR Government Reaching Out Award, 2018
- HKSAR Government Talent Development Scholarship (Innovation), 2017
- The IET Prize, the Institute of Engineering and Technology, 2017
- The Soong Ching Ling Foundation Scholarship, China Soong Ching Ling Foundation, 2015
Profiles
For the complete publication record, visit Google Scholar or browse the Publications page.