EI / SCOPUS / CSCD 收录

中文核心期刊

语音合成中的跨语言情感解耦与迁移

Cross-lingual emotion disentanglement and transfer in speech synthesis

  • 摘要: 本文针对低资源条件下合成语音情感表现欠佳的问题, 探索一种跨语言情感解耦与迁移合成方法, 利用情感数据丰富语种的语音作为参考, 提升低资源语种说话人的情感表达能力。针对语音情感与内容信息互相纠缠的问题, 提出结合语音通用特征和超音段情绪韵律特征的情感解耦方法来区分发音相关和内容相关的情感特征。在此基础上, 为了应对不同语言情绪韵律模式差异导致的信息重新耦合困难, 提出去相关投影变换方法, 进一步提高合成语音的质量与情感表现力。主客观评测实验表明, 本文方法在跨语言情感迁移合成方面的性能优于现有基线和最前沿(SOTA)方法, 合成语音情感表现力主观评价分(MOS)提高约0.2, 情感类别混淆概率下降超过10%, 语音音质、内容正确率和说话人相似度也略有提高。

     

    Abstract: This paper addresses the issue of poor emotional expression in speech synthesis under low-resource conditions by exploring emotion disentanglement and transfer methods in a cross-lingual emotional text-to-speech. By utilizing speech data from languages with rich emotional datasets as references, the emotional expression capabilities of speakers in low-resource languages are expected to be enhanced. To tackle the entanglement of speech emotion and content information, an emotional disentanglement method is proposed that combines general speech features with supra-segmental emotional prosody features, distinguishing between articulatory-related and content-related emotional characteristics. Furthermore, to address the challenge of re-encoding information caused by differences in emotional prosody patterns across languages, a decorrelation projection transformation method is introduced, thereby improving the quality and emotion presentation of synthesized speech. Subjective and objective evaluation experiments demonstrate that the proposed method outperforms existing baselines and state-of-the-art (SOTA) methods in cross-lingual emotional transfer synthesis. The mean opinion score (MOS) for emotional expressiveness is increased by approximately 0.2 and the probability of emotional category confusion is reduced by over 10%. In the meantime, there have also been slight improvements in speech quality, content accuracy, and speaker similarity.

     

/

返回文章
返回