Speech recognition method and related device

By using delay neural networks and syllable modeling units in speech recognition technology, and combining the information of preceding and following syllables for recognition, the problem of too fine granularity of phoneme modeling units in the existing technology is solved, and the accuracy and applicable scenarios of speech recognition are improved.

CN114360510AActive Publication Date: 2022-04-15TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
TENCENT TECHNOLOGY (SHENZHEN) CO LTD
Filing Date
2022-01-14
Publication Date
2022-04-15

AI Technical Summary

Technical Problem

Existing speech recognition technology uses phonemes as modeling units, and the granularity is too fine, resulting in high quality requirements for speech data, difficulty in adapting to complex scenarios, low recognition accuracy, and limited applicable scenarios.

Method used

The time-delay neural network is used as the acoustic model, and the syllable is used as the modeling unit. The syllable probability distribution of the speech frame is determined through multiple syllable modeling units in the output layer, and the information of the preceding and following syllables is combined for recognition to improve the recognition accuracy and reduce the need for speech data. quality requirements.

Benefits of technology

It improves the accuracy and robustness of speech recognition, expands the applicable scenarios of speech recognition technology, and can obtain more accurate recognition results under low-quality speech data conditions.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

The embodiment of the invention discloses a speech recognition method and a related device, and at least relates to a speech recognition technology in artificial intelligence, speech data to be recognized are used as input data of a time delay neural network in an acoustic model, and an output layer of the time delay neural network comprises acoustic modeling units corresponding to a plurality of syllables respectively, so that the speech recognition efficiency is improved. And the syllable probability distribution corresponding to the voice frames included in the voice data can be obtained by taking the syllables as the recognition granularity through the time delay neural network. When syllable recognition is carried out through the output layer, auxiliary judgment can be carried out on the syllables to which the voice frames belong on the basis of pronunciation rules in combination with front and back syllable information of the voice frames, so that more accurate syllable probability distribution is output. Moreover, since the syllables are generally composed of one or more phonemes, the method has higher fault-tolerant capability, not only can more accurately determine the speech recognition result based on the probability distribution of the syllables, but also has low requirements for the quality of the speech data to be recognized, and effectively expands the application scenarios of the speech recognition technology.
Need to check novelty before this filing date? Find Prior Art

Citation Information

Patent Citations

  • Speech recognition method and device, computer device and storage medium

    CN108022587A

  • Speech recognition method and device, and electronic equipment

    CN110211588A

  • Wakeup word detection method, device and equipment based on artificial intelligence, and medium

    CN110838289A

  • Voiceprint recognition method and device, computer storage medium and electronic equipment

    CN110970036A

  • Speech recognition model determination method and device, speech recognition method and device, and electronic equipment

    CN111402893A

Cited By

  • Voice input method and system and readable storage medium

    CN115798465A

  • A voice input method, system, and readable storage medium

    CN115798465B

  • Speech recognition method, system and terminal

    CN116189666A

  • Speech recognition method, system and terminal

    CN116189666B

  • Voice recognition method and device, electronic equipment and storage medium

    CN116612783A