Audio Demo

SemBridge

Semantic-Token Anchoring for Continuous-Latent Autoregressive Speech Generation

Zero-shot TTS Score-conditioned SVS Speech–singing generation

Abstract

Continuous-latent autoregressive speech generation has emerged as a promising alternative to discrete-token modeling by avoiding quantization loss and preserving richer acoustic information. However, continuous acoustic targets do not expose linguistic structure as explicit token-level prediction targets. Consequently, the autoregressive language model (LM) must acquire linguistic structure indirectly through acoustic prediction, which can compromise the content fidelity of generated speech. We propose SemBridge, a training-only semantic-token anchoring framework for continuous-latent autoregressive speech generation. SemBridge uses discrete semantic tokens to directly supervise autoregressive LM states and employs a Semantic-Aligned Acoustic VAE to organize the continuous target space under the same semantic reference. The semantic supervision is used only during training, while inference remains entirely continuous. We evaluate SemBridge on zero-shot text-to-speech (TTS) and score-conditioned singing voice synthesis (SVS). Across multiple benchmarks, SemBridge improves content accuracy, as measured by word and character error rates (WER/CER), while maintaining competitive speaker similarity and perceptual quality. Experimental results demonstrate that explicit semantic-token supervision for autoregressive state learning is an effective and general direction for continuous speech generation.

01 / Audio

Zero-shot text-to-speech

Each sample conditions on a short reference recording and generates the target text in the prompt speaker's voice.

Sample
Prompt
Target text
SemBridge
ZH · R001Standard
我箱子里有一套夏装,可是现在穿它也未免太不成体统了。
ZH · R002Standard
即使拉斯维加斯的赌场,在奢华的程度上都很难媲美。
EN · R001Standard
Reading between the lines requires understanding.
EN · R002Standard
At least five species of reptiles have been recorded for the island.

Expressive TTS

Emotion-rich prompts transfer speaking style while preserving the requested content.

Sample
Prompt
Target text
SemBridge
ZH · E001Happiness
太好了,我们终于成功了!今晚一定要好好庆祝一下!
ZH · E002Disgust
这种自以为是的做法,真让人一点也喜欢不起来。
ZH · E003Surprise
什么?你是说我们真的中奖了?这也太不可思议了吧!
ZH · E004Sadness
后来我才明白,有些告别,真的就是最后一面。
ZH · E005Anger
我已经说得很清楚了,别再挑战我的底线!
ZH · E006Fear
等等,门外好像有人,你听见那个声音了吗?

Character and celebrity voices

Selected zero-shot samples from distinctive character and public-figure prompts.

Chinese

Voice
Prompt
Target text
SemBridge
ZH · C001Zhu Bajie
师父您放心,老猪这回绝不偷懒,先把行李放好,再去看看有没有斋饭。
ZH · C002Cai Xukun
舞台上的每一次进步,都来自台下反复的练习和坚持。
ZH · C003Ding Zhen
从雪山脚下走到这里,我最想分享的,还是家乡清晨的风和远处的云。
ZH · C004Hou Yi
太阳落山以后,天空终于安静下来,今晚就让月亮替我值班吧。
ZH · C005Hua Fei
这宫里的花开得再好,也要有人懂得它背后的冷暖。
ZH · C006Jay
这首歌的灵感其实很简单,就是把生活里没说出口的话写进旋律。
ZH · C007Child voice
我昨天发现一颗特别亮的星星,不知道它会不会也在偷偷看着我们。
ZH · C008Taiyi
想当年我云游四海,见过的奇人异事,三天三夜也说不完。
ZH · C009Wu Kong
今天天气这么好,我们去公园散散步,顺便找点好吃的,怎么样?
ZH · C010Yae
今夜月色正好,不如随我沿着山路慢慢走,看看风会把花香带到哪里。

English

Voice
Prompt
Target text
SemBridge
EN · C001Benedict
There are moments when the truth arrives quietly, yet changes everything we thought we understood.
EN · C002Morty
Okay, I know this sounds weird, but maybe we should think this through before opening that door.
EN · C003Rick
Listen, the universe is complicated, but that does not mean we have to make every decision complicated too.
EN · C004Trump
Today we are presenting something bold, clear, and built to deliver a truly remarkable experience.

Narrative voices

Selected zero-shot samples spanning historical, documentary, wuxia, dramatic, and nature narration.

Chinese

Voice
Prompt
Target text
SemBridge
ZH · N001Historical storyteller
汉武帝听完大臣的奏报,沉默片刻,才缓缓问道:这件事究竟是谁的主意?
ZH · N002Documentary narrator
在遥远的星空深处,一束微弱的光,正穿越漫长的时间向我们靠近。
ZH · N003Wuxia narrator
张无忌抬头望去,只见山门之外风雪骤起,来人却始终没有露面。
ZH · N004Dramatic narrator
那只猎犬停在门前,竖起耳朵,仿佛听见森林深处传来一阵异响。
ZH · N005Nature narrator
夜幕降临后,草原上的动物逐渐苏醒,一场关于生存的故事悄然开始。

02 / Audio

Score-conditioned singing voice synthesis

SemBridge generates singing from lyrics, MIDI pitch, note duration, and a short singer prompt.

Sample
Prompt
Lyrics
SemBridge
ZH · R001Standard
像我这样庸俗的人从不喜欢装深沉
ZH · R002Standard
在我的怀里你不用害怕失眠
EN · R001Standard
How I wondered where they'd gone, but they're back again, just like a long lost friend, all the songs I loved so well.
EN · R002Standard
When he's long gone and he's next to me.

03 / Audio

Unified speech–singing generation

A single autoregressive sequence alternates between speech and score-conditioned singing while preserving a shared acoustic trajectory.

ZH · U001 《像我这样的人》

Speech

接下来,送给大家一段《像我这样的人》,也送给每一个还在认真生活的人。

Singing

像我这样孤单的人,像我这样傻的人。

Speech

这一句,也许唱的就是我们自己。愿每一个认真生活的人,都能被世界温柔看见。

Prompt
SemBridge

ZH · U002 《彩虹》

Speech

夜深了,送给大家一段《彩虹》,也送给那个让我一直牵挂的人。

Singing

看不见你的笑,我怎么睡得着。

Speech

愿这段旋律陪你入梦,我们下次再见。

Prompt
SemBridge

ZH · U003 Happiness

Speech

今天心情特别好,先不多说了,给大家唱一句关于幸福的歌。

Singing

幸福开始有预兆。

Speech

原来幸福真的会提前打招呼,今天的好心情也分享给你。

Prompt
SemBridge

ZH · U004 《隐形的翅膀》

Speech

各位观众,俺老孙今天也来唱上一段《隐形的翅膀》,听好了!

Singing

我看见每天的夕阳,也会有变化。

Speech

这一曲唱罢,多谢大家捧场,俺老孙先走一步!

Prompt
SemBridge