Speech recognition model training method and device, equipment and medium

A speech recognition model and training method technology, applied in speech recognition, speech analysis, instruments, etc., can solve the problems of large amount of calculation, low text efficiency, long translation lag time, etc., to achieve the effect of enhancing audio information and saving labor costs

CN113870845APending Publication Date: 2021-12-31PING AN TECH (SHENZHEN) CO LTD
0 Cites 6 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Publication Date
2021-12-31

Smart Images

  • Figure 1
    Figure 1
  • Figure 2
    Figure 2
  • Figure 3
    Figure 3
Patent Text Reader

Abstract

The invention relates to the field of artificial intelligence, and provides a speech recognition model training method and device, equipment and a medium, and the method comprises the steps: obtaining a speech sample set containing speech samples; inputting the speech sample into an initial recognition model; obtaining a to-be-processed audio clip through audio enhancement processing; performing teacher acoustic feature extraction through a teacher network in the initial recognition model to obtain a first feature vector, and performing student acoustic feature extraction through a student network in the initial recognition model to obtain a second feature vector; performing alignment comparison processing in combination with a dynamic queue in a teacher network to obtain a loss value; and when the loss value does not reach a preset convergence condition, carrying out iterative updating until convergence, and obtaining a trained speech recognition model. According to the invention, common speech recognition through the teacher network and the student network is realized, and the training efficiency is improved. The method is suitable for the field of artificial intelligence, and can further promote the construction of a smart city.
Need to check novelty before this filing date? Find Prior Art

Description

technical field

[0001] The invention relates to the field of speech recognition of artificial intelligence, in particular to a speech recognition model training method, device, computer equipment and storage medium. Background technique

[0002] Speech translation is the process of converting one natural language (source language) into another natural language (target language). Unlike traditional machine translation, the input of voice translation is directly voice, and the output is text. With international communication With the increase of the number of people, different languages ​​are used to communicate more and more frequently. In order to overcome language communication barriers, online voice translation based on clients has been widely used.

[0003] Online speech translation generally involves two steps. The first is to perform speech recognition, that is, to convert the speech signal of the first language input by the user into text; the second is to translate th...

Examples

Embodiment Construction

[0029] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some of the embodiments of the present invention, but not all of them. Based on the embodiments of the present invention, all other embodiments obtained by persons of ordinary skill in the art without creative efforts fall within the protection scope of the present invention.

[0030] The speech recognition model training method provided by the present invention can be applied in such as figure 1 In the application environment of , where the client (computer device or terminal) communicates with the server through the network. Wherein, clients (computer devices or terminals) include but are not limited to various personal computers, notebook computers, smart phones, tablet computers and portable wearable devices. The server ...