节点文献

语音识别置信度特征提取算法研究

A Study of Feature Extraction Algorithm of Speech Recognition Confidence Measure

【作者】 国玉晶

【导师】 刘刚;

【作者基本信息】 北京邮电大学 , 模式识别与智能系统, 2010, 硕士

【摘要】 大规模连续语音识别的研究已经进行了二十多年,虽已取得了显著进展,但距离广泛应用还有相当的距离。在克服识别算法本身缺陷、追求识别性能提升的过程中,研究者们逐渐引入了置信度的概念,用它来衡量语音识别系统所作决策的可信程度。近年来,语音识别置信度在语音错误检测与错误纠正,无监督和半监督训练、多遍搜索技术和语料库中错误语料甄选等应用中都发挥了非常重要的作用。传统的语音识别置信度标注基于不同置信特征或者特征组合进行分类判决,目前常使用的置信特征主要来源于解码信息。但是,方面现有置信度特征对解码信息的挖掘仍局限于孤立和静态,而忽略了词与周围环境之间的关系;另一方面,目前声学特征仍占主要地位,而人类听觉实验表明,人在进行语音理解时,大约有30%的信息来自于语法、语义等知识的指导。因此,在置信度特征提取中,如何挖掘出词与环境之间的关系,同时提炼出词的语法和语义特征,从而提高识别后处理性能,是一个非常值得研究的问题。基于上述目的,本文在搭建传统语音识别置信度标记系统的基础上,提出了两种新的置信度特征,一是环境特征,分为上下文环境、动态环境、句全局环境三类,通过对解码信息的再加工,从空间与时间角度较全面地描述了词与环境之间的关系;二是基于主题相似性的语义层置信特征提取算法TSS (Topic Similarity based Semantic confidence feature extraction algorithm),通过主题模型LDA(Latent Dirichlet Allocation)计算得到识别结果中词的主题分布及其上下文的主题分布,并将二者之间的主题相似性作为词的语义置信特征。实验表明,本文提出的两种特征深入挖掘了解码层的有效信息,又增加了置信特征的信息来源,与解码层置信特征进行组合后能有效地提高置信度标注的精度。

【Abstract】 The large vocabulary continuous speech recognition research has been studied for more than two decades, though significant progress has been made, but there is still a considerable distance from the wide range of applications. In the pursuit of overcoming the deficiencies inside recognition algorithm itself, and improving recognition performance, researchers have gradually introduced the concept of confidence measure, to measure in which degree we could trust the result of speech recognition system. In recent years, speech recognition confidence measure has played a very important role in many applications, including speech error detection and correction, no supervision and semi-supervised training, multi-search technology and corpus selection and verification, etc.Based on different feature-combinations, traditional speech recognition confidence is actually a confidence annotation or classification decisions, with mainly information from decoding messages. However, the current confidence features are still limited to isolated and static, while ignoring the the relationship between the words and their surrounding environment; on the other hand, acoustic features are still dominant, while the experiments show that in speech understanding, human beings depend approximately 30% of the information from the syntax, semantics and other non-acoustic knowledge. Therefore, how to dig out the relationship between words and the environment, to extract the characteristics of syntax and semantics of the word so as to enhance recognition performance of post-processing is a very worthwhile study in the field of feature extraction of confidence measure.For purpose above, in addition to build a traditional baseline system of speech recognition confidence annotation, this paper proposed two new confidence features. The first one is environmental feature, including context, dynamic and the global environment features, which extract more valuable information from the intermediate production of decoding, and provide a more comprehensive description of the relationship between words and the environment from both perspectives of space and time. The second is based on topic similarity of the semantic layer of confidence feature extraction algorithm TSS (Topic Similarity based Semantic confidence feature extraction algorithm), using a new theme Model LDA (Latent Dirichlet Allocation) we could calculated the distribution on theme of first the word in recognition results and then in the context. and distribution similarity between the theme and the word could be figured out as the semantic features of words in context. Experiments show that the two features proposed in this paper deeply excavated valuable decoding information, and, after combined with acoustic features, an significant increase in accuracy of confidence annotation experiment has been seen.

节点文献中: 

本文链接的文献网络图示:

本文的引文网络