Speech Recognition Method, Apparatus, Device, and Storage Medium
By inputting speech data into the ASR model in frames, and combining hot word maps for bundle search and matching judgment, the problem of low accuracy of unique noun recognition in the prior art is solved, and a higher accuracy of speech recognition is achieved.
Patent Information
- Application Number
- CN202210834523.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-14
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2042-07-14
AI Technical Summary
The accuracy rate of existing speech recognition technologies is low when identifying unique nouns (such as person names, song names, place names, etc.), especially when it is recognized through hot word matching, the accuracy rate is still not high.
By inputting the voice data into the ASR model in frames for identification processing, and performing bundle search and matching judgment in combination with the hot word map, the target candidate words of each frame are determined. If matched, the alternative words for the next frame are determined from the hot word map; if not matched, the candidate words are determined based on the acoustic probability until the target candidate words for each frame are determined.
The accuracy of speech recognition is improved, especially when identifying unique nouns, the accuracy of recognition is significantly improved by combining the matching of hot word maps and bundle search.
Smart Images

Figure CN115206301B_ABST
Abstract
Claims
1. A speech recognition method, It is characterized in that The method comprises: The speech data is divided into frames and input into the ASR model for recognition processing to obtain multiple candidate words and their corresponding acoustic probabilities; Obtaining a first target candidate word corresponding to the current frame by performing a beam search on the candidate words corresponding to the current frame and their acoustic probabilities; Determine whether the first target candidate word matches a hot word in a hot word graph, wherein the hot word graph is constructed based on a preset hot word table; the hot word graph is constructed based on the preset hot word table, including: splitting the hot words in the preset hot word table to obtain words to be processed; according to the number of words corresponding to each hot word, using the corresponding words to be processed in order from large to small to construct arcs connecting each node in the hot word graph, and setting corresponding arc weights, wherein the words to be processed correspond to the arcs one-to-one, and multiple words to be processed corresponding to the hot words form a closed loop in the hot word graph; If there is a match, based on the first target candidate word, determine the target node in the hot word graph; according to the target node, use the to-be-processed word corresponding to the arc connected to the target node as the candidate word; when the candidate word in the next frame includes the candidate word, determine the number of candidate words included; judge whether the number is less than or equal to the preset number; if the number is less than or equal to the preset number, determine the difference between the number and the preset number, and determine the remaining candidate words based on the acoustic probability corresponding to the candidate word in the next frame according to the difference; if the number is greater than the preset number, use the first preset number of candidate words with the highest acoustic probability as the second target candidate word; If there is no match, the second target candidate word is determined based on the acoustic probability corresponding to the candidate word in the next frame, until the target candidate words corresponding to each frame are determined; Based on the target candidate words corresponding to each frame, a plurality of sentence combinations and their corresponding acoustic scores are obtained, and the sentence combinations are used to search a hot word graph to obtain hot word scores; Based on the acoustic score and the hot word score, a recognition result is determined.
2. The speech recognition method according to claim 1, It is characterized in that After the corresponding arc weights are set, the method further includes: A fallback arc is set on each node of the hot word graph, the fallback arc is an arc connecting each node with the initial node, and the weight corresponding to the fallback arc is the opposite number of the existing weight of each node; When the hot word constructed later is the prefix of the hot word constructed previously, the fallback arc weight of the node corresponding to the prefix of the constructed hot word is reset to zero.
3. The speech recognition method according to claim 1, It is characterized in that The step of obtaining multiple sentence combinations and their corresponding acoustic scores based on the target candidate words corresponding to each frame includes: Determining a plurality of sentence combinations based on the target candidate words corresponding to each frame; According to the target candidate words contained in the sentence combination, obtaining the acoustic probability corresponding to the target candidate words; The acoustic score is obtained by multiplying the acoustic probabilities corresponding to the target candidate characters.
4. The speech recognition method according to claim 1, It is characterized in that Determining the recognition result based on the acoustic score and the hot word score includes: Determine the total score of each sentence combination through the acoustic score and the hot word score; The sentences with the highest total scores are combined as the recognition result.
5. A speech recognition device, It is characterized in that The device comprises: The recognition module is used to input the speech data into frames and input them into the ASR model for recognition processing to obtain multiple candidate words and their corresponding acoustic probabilities; A search module, configured to obtain a first target candidate word corresponding to the current frame by performing a beam search on the candidate words corresponding to the current frame and their acoustic probabilities; A matching judgment module is used to judge whether the first target candidate word matches a hot word in a hot word graph, wherein the hot word graph is constructed based on a preset hot word table; wherein the matching judgment module includes: a splitting submodule, used to split the hot words in the preset hot word table to obtain words to be processed; a graph construction submodule, used to construct arcs connecting the nodes in the hot word graph using the corresponding words to be processed in order from large to small according to the number of words corresponding to each hot word, and set corresponding arc weights, wherein the words to be processed correspond to the arcs one-to-one, and the multiple words to be processed corresponding to the hot words form a closed loop in the hot word graph; A corresponding processing module is used to determine the candidate word of the next frame from the hot word graph based on the first target candidate word if there is a match, and when the candidate word of the next frame includes the candidate word, the candidate word is used as the second target candidate word; if there is no match, the second target candidate word is determined based on the acoustic probability corresponding to the candidate word in the next frame, until the target candidate words corresponding to each frame are determined; wherein the corresponding processing module includes: a node determination submodule, an alternative word determination submodule, a quantity determination submodule, a judgment submodule and a processing submodule; the node determination submodule is used to determine the target node in the hot word graph based on the first target candidate word; the alternative word determination submodule The block is used to use the to-be-processed word corresponding to the arc connected to the target node as a candidate word according to the target node; the quantity determination submodule is used to determine the number of candidate words when the candidate words of the next frame include the candidate word; the judgment submodule is used to judge whether the number is less than or equal to the preset number; the processing submodule is used to determine the difference between the number and the preset number if the number is less than or equal to the preset number, and determine the remaining candidate words based on the acoustic probability corresponding to the candidate words in the next frame according to the difference; if the number is greater than the preset number, the first preset number of candidate words with the highest acoustic probability are used as the second target candidate words; A score calculation module, used to obtain multiple sentence combinations and their corresponding acoustic scores based on the target candidate words corresponding to each frame, and use the sentence combinations to search the hot word graph to obtain the hot word scores; The sentence determination module is used to determine the recognition result based on the acoustic score and the hot word score.
6. A computer device, It is characterized in that The computer device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores computer-readable instructions, and the processor implements the speech recognition method according to any one of claims 1 to 4 when executing the computer-readable instructions.
7. A computer-readable storage medium, It is characterized in that The computer-readable storage medium stores computer-readable instructions, and when the computer-readable instructions are executed by a processor, the speech recognition method according to any one of claims 1 to 4 is implemented.
Citation Information
Patent Citations
Speech recognition method, device and apparatus
CN109523991A
End-to-end streaming keyword spotting
CN112368769A