ReStyle-TTS: Relative and Continuous Style Control
for Zero-Shot Speech Synthesis

Haitao Li1,2,  Chunxiang Jin3,  Chenglin Li1,2,  Wenhao Guan4,2,  Zhengxing Huang1,  Xie Chen5,2

1Zhejiang University    2Shanghai Innovation Institute    3Ant Group    4Xiamen University    5Shanghai Jiao Tong University

Abstract

Zero-shot text-to-speech models can clone a speaker's timbre from a short reference audio, but they also strongly inherit the speaking style present in the reference. As a result, synthesizing speech with a desired style often requires carefully selecting reference audio, which is impractical when only limited or mismatched references are available. While recent controllable TTS methods attempt to address this issue, they typically rely on absolute style targets and discrete textual prompts, and therefore do not support continuous and reference-relative style control.

We propose ReStyle-TTS, a framework that enables continuous and reference-relative style control in zero-shot TTS. Our key insight is that effective style control requires first reducing the model's implicit dependence on reference style before introducing explicit control mechanisms. To this end, we introduce Decoupled Classifier-Free Guidance (DCFG), which independently controls text and reference guidance, reducing reliance on reference style while preserving text fidelity. On top of this, we apply style-specific LoRAs with Orthogonal LoRA Fusion to enable continuous and disentangled multi-attribute control, and introduce a Timbre Consistency Optimization module to mitigate timbre drift caused by weakened reference guidance.

Experiments show that ReStyle-TTS enables user-friendly, continuous, and relative control over pitch, energy, and multiple emotions while maintaining intelligibility and speaker timbre, and performs robustly in challenging mismatched reference–target style scenarios.

Emotion Style Control Emotion

Drag the slider to continuously control the intensity of each emotion style. At 0.0 the output follows the original reference speaking style (no adjustment). Larger values progressively amplify the target emotion up to a maximum of 4.0. Release the slider (or lift your finger) to hear the result.

Prosodic Style Control Prosodic

Drag the slider to control pitch or energy relative to the reference. At 0.0 the output matches the original reference style. Positive values increase the attribute; negative values decrease it. Range is −2.0 → +2.0.