SUZUNE S1 — 80M NON-AUTOREGRESSIVE TTS Read Research Paper #001

Suzune S1 Speech & Voice AI

An ultra-lightweight 80M-parameter non-autoregressive TTS engine engineered by Darsh Yadav at Sorika Labs. Powered by a 12-layer PL-BERT, continuous AdaLN prosody modeling, and an end-to-end fast iSTFTNet neural vocoder with Neural Source Filter (NSF) harmonic excitation.

06 — SUZUNE S1 SPEECH INTELLIGENCE

SUZUNE S1
80M Non-AR neural TTS engine.

Suzune S1 breathes life into speech. Built on an 80M non-autoregressive architecture with a 12-layer PL-BERT, continuous AdaLN prosody, and an end-to-end fast iSTFTNet neural vocoder with Neural Source Filter (NSF) harmonic excitation.

Fast iSTFTNet Vocoder + AdaLNRTF: 0.018x (GPU)

Directly predicts multi-band STFT magnitude and phase components with strictly constant ~320MB VRAM footprint and zero hallucination risk.

Suzune S1 Voice Architecture Acoustic Preview
Suzune S1 Acoustic Preview
SAMPLING RATE24,000 Hz
VRAM FOOTPRINT~320 MB O(1)
CORPORAEnglish & Hindi
ENGINE SPECIFICATIONS

Production Architecture Benchmarks

24kHz Fast iSTFTNet Vocoder

Directly predicts multi-band STFT magnitude and phase components with Neural Source Filter harmonic excitation and Snake periodic activation functions.

RTF 0.018x (GPU) • RTF 0.087x (x86 CPU)

12-Layer PL-BERT + AdaLN

Phoneme-level contextual representation with continuous duration and prosody projection. Trained across bilingual English and Hindi acoustic corpora.

Bilingual Context • Continuous F0 & Cadence

Constant ~320 MB VRAM O(1)

Eliminates autoregressive KV caches and stochastic sampling. Prevents phonetic hallucinations, word skipping, and Out-Of-Memory (OOM) crashes.

Zero Word Skips • 100% Deterministic
Under Development