AI-Driven Cantonese Singing Voice Synthesis
A Cantonese singing voice synthesis initiative using carefully prepared singing and reading data, transcriptions, scales, tempos, and phonetic combinations for precise melodic and phonetic modeling.
Project image 01
Project image 02
Project image 03
Project overview
This research-led voice synthesis project focuses on modelling the Cantonese singing voice of lyricist Chow Yiu Fai. The work begins with a carefully designed recording and annotation process covering both singing and spoken reading material.
Each recording is paired with transcription data and organized around major scales, multiple tempos and targeted phonetic combinations within the performer’s comfortable vocal range. This creates a structured dataset for analysing the relationship between Cantonese pronunciation, pitch, timing and melodic expression.
The resulting training material supports more precise phonetic and musical modelling, providing a technical foundation for expressive Cantonese singing-voice synthesis and future creative applications.