短语音说话人识别中域不变特征学习方法
Domain-invariant feature learning method for short utterance speaker recognition
-
摘要: 在文本无关的说话人识别任务中, 识别性能会随语音长度变短而显著降低, 尤其当语音长度在2 s以内时会更加明显。为了解决这一问题, 本文提出了一种域不变特征学习方法。在特征输入阶段将梅尔谱图与色谱图进行融合, 为短语音提供更丰富的输入信息; 在训练阶段引入域对抗神经网络, 通过在说话人特征提取器和语音时长相关的域判别器之间进行对抗训练, 促使网络学习到域不变的说话人特征。在中文说话人识别数据集King-ASR-459和CN-Celeb上的实验结果表明, 所提方法有效提升了短语音上的说话人识别效果。Abstract: In text-independent speaker recognition tasks, performance significantly degrades as speech duration decreases, particularly when the speech is shorter than 2 s. To address this challenge, this paper proposes a domain-invariant feature learning method for short utterance. In the feature input stage, the mel-spectrogram is fused with the chromagram to provide richer input information for short utterance. In the training stage, a domain adversarial neural network is introduced to learn domain-invariant speaker features by performing adversarial training between the speaker feature extractor and the domain discriminator related to speech duration. Experiments on the Chinese speaker recognition datasets King-ASR-459 and CN-Celeb show that the proposed method effectively improves the speaker recognition performance on short utterance.
下载: