Write the notes. Shape the breath. Hear the difference between playing a melody and performing a phrase.
0Voice samples
0Trained models
100%Procedural audio
01
The score
An original two-phrase miniatureScoreControl pitch
Click a note to edit. The curve shows planned pitch, not measured pitch.
00:00 / 00:12
Preparing the mathematical voice…44.1 kHz · 2× source rate
Edit score & phonemes
One line per note: D4 1 la 0.7 means pitch, beats, syllable, stress. R 0.5 is a rest; _ sustains the preceding syllable on a new pitch. Use one syllable per note; this is not unrestricted text-to-speech. Explicit phones: [s:t:eh:iy]. Vowels: aa ae ah eh ih iy oh oo uh ax er. Consonants: m n ng l r w y h s z sh zh f v th dh p b t d k g. Best register: roughly C3–C5. Maximum score: 90 seconds.
02
Inside the phrase
DRY SIGNAL
Rendered waveformPCM
Phrase effortControl signal
— Hz
/ — /
— %
— ct
Same voice and score in A/B. Dry RMS levels are matched; room is shared. Breath reserve is a phrasing rule, not a physiological measurement.
No network requests. No microphone access.
The plan, implemented
Build the instrument. Then build the performance.
A periodic source can play notes. A singer-shaped performance also needs phrasing, articulation, connected gestures, and coordinated changes in sound. Cantoria keeps those decisions explicit and inspectable.
Neil Thapen’s open glottal-source implementation supplies the LF waveform parameterization. Cantoria adapts that small component, but uses its own acoustic filter bank, score compiler, rendering pipeline, and performance rules. It does not embed Pink Trombone’s complete vocal-tract solver.
A working precedent for performative, parametric singing: a source–filter instrument with controllable pitch, breath, tension, and effort. Its official distribution uses CeCILL. Cantoria references the architecture; no Cantor Digitalis code or Max patches are included.
A rules-based articulatory speech system demonstrating the separation of phonetic planning and acoustic production. It is useful prior art for a larger phonetic front end. Its code is not included in this prototype.
Generated inhalations and a falling breath reserve
Voice color
Same tract and phoneme targets
Effort-linked source tension and subtle resonance motion
Articulation
Same syllables, essential transitions retained
Longer coarticulation and consonant anticipation
Fair comparison
Same score length, voice anatomy, sample rate, and dry RMS target. No added backing track.
Expression is not just random pitch jitter. One phrase-level control coordinates amplitude, source tension, breath, and resonance. Seeded low-amplitude variation adds texture without replacing that structure. The “breath reserve” is a useful control law—not a solved model of lungs, tissue, and fluid dynamics.
Entirely procedural means exactly this.
The app computes its own glottal waveforms from equations, moves its resonators using explicit rules, generates its turbulence with a seeded pseudorandom generator, and creates room reflections with feedback delays. There are no voice recordings, pretrained weights, neural networks, speech services, microphone inputs, or external runtime dependencies. The in-memory PCM buffers are newly synthesized outputs, not sample-library inputs.
Audio is rendered in a Web Worker, then played through Web Audio. Editing a synthesis parameter rerenders the score; this is a score-to-performance instrument, not a low-latency live vocal controller. A/B and room levels switch during playback. A same-thread fallback is available when a browser blocks blob workers.
The buttons below run fresh engineering checks in your browser. They establish that the synthesis works and that the performance layer changes it. They do not establish that the result is perceptually indistinguishable from a human singer.
Not run in this session.
The claim boundary: working, deterministic, non-neural singing synthesis with an audible A/B ablation is testable here. Naturalness, lyric intelligibility, and “sounds like an actual singer” require human listening tests. This prototype has not passed those tests. Consonants are approximations, the dictionary is deliberately small, and extreme pitches can sound vowel-like rather than speech-like.
A useful listening experiment
Start with Vowel aria, Room at Dry, and switch A/B on a held note. Then use Afterglow to judge whether you can understand the words without looking at the score. Finally, set Expression to zero: the two modes should converge to identical output. That last comparison checks that the performance layer really is removable.
For a stronger study, randomize and blind A/B presentation, retain RMS matching, use multiple listeners and previously unheard melodies, and rate naturalness, expressive intention, and lyric intelligibility separately. Add genuine human recordings only as evaluation references, not as synthesis inputs. Do not equate a waveform change or a low pitch error with a human-quality singing voice.
The path to a stronger singer
The next useful work is not adding more random wobble. It is fitting richer rule-based coarticulation, adding nasal zeros and side branches, improving glottal register transitions, expanding stress and syllable planning, and tuning against blinded listening results. The renderer’s explicit control traces make each intervention independently testable. A more complete articulatory backend can replace the current tract filters without discarding the performance planner.