Suzune S1 Speech & Voice AI
An ultra-lightweight 80M-parameter non-autoregressive TTS engine engineered by Darsh Yadav at Sorika Labs. Powered by a 12-layer PL-BERT, continuous AdaLN prosody modeling, and an end-to-end fast iSTFTNet neural vocoder with Neural Source Filter (NSF) harmonic excitation.
SUZUNE S1
80M Non-AR neural TTS engine.
Suzune S1 breathes life into speech. Built on an 80M non-autoregressive architecture with a 12-layer PL-BERT, continuous AdaLN prosody, and an end-to-end fast iSTFTNet neural vocoder with Neural Source Filter (NSF) harmonic excitation.
Directly predicts multi-band STFT magnitude and phase components with strictly constant ~320MB VRAM footprint and zero hallucination risk.

Production Architecture Benchmarks
24kHz Fast iSTFTNet Vocoder
Directly predicts multi-band STFT magnitude and phase components with Neural Source Filter harmonic excitation and Snake periodic activation functions.
12-Layer PL-BERT + AdaLN
Phoneme-level contextual representation with continuous duration and prosody projection. Trained across bilingual English and Hindi acoustic corpora.
Constant ~320 MB VRAM O(1)
Eliminates autoregressive KV caches and stochastic sampling. Prevents phonetic hallucinations, word skipping, and Out-Of-Memory (OOM) crashes.