Speech Recognition Method, Apparatus, Electronic Device, and Storage Medium

By generating a collection of difficult and easy-to-recognize words and determining their hot word weights, the problem of insufficient recognition rate of the speech recognition engine in different scenarios is solved, and the overall accuracy and adaptability of speech recognition is improved.

CN114360544BActive Publication Date: 2025-07-25BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210141226.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-16
Publication Date
2025-07-25
Estimated Expiration
2042-02-16

AI Technical Summary

Technical Problem

The existing voice recognition engine cannot achieve a recognition rate of 100%, resulting in poor hot word recognition effect in different scenarios and difficult to meet diversified needs.

Method used

By generating a set of difficult-to-recognized candidate words and a set of easy-to-recognized candidate words, their hot word weights are determined separately, and automatic speech recognition is performed based on these weights to improve the recognition rate of difficult-to-recognized words.

Benefits of technology

The overall recognition rate of speech recognition, especially the recognition rate of difficult-to-recognize words, is enhanced, and the adaptability and accuracy of speech recognition engines in different scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114360544B_ABST
    Figure CN114360544B_ABST
Patent Text Reader

Abstract

The present disclosure provides a speech recognition method, apparatus, electronic device, and storage medium. A specific implementation of the method includes: for a candidate word in a candidate word set, determining the recognition rate of speech recognition of the candidate word; generating a difficult-to-recognize candidate word set based on the candidate words in the candidate word set whose recognition rates are less than a first preset recognition rate threshold; generating an easy-to-recognize candidate word set based on the candidate words in the candidate word set whose recognition rates are greater than a second preset recognition rate threshold, where the second preset recognition rate threshold is greater than the first preset recognition rate threshold. This implementation provides difficult-to-recognize and easy-to-recognize words for the hot word technology; respectively determining the hot word weights of the candidate words in the difficult-to-recognize candidate word set and the hot word weights of the candidate words in the easy-to-recognize candidate word set, and performing automatic speech recognition based on the determined hot word weights.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the field of speech recognition technology, and particularly to speech recognition methods, devices, electronic devices, and storage media. Background Art

[0002] With the development of speech recognition technology, numerous speech recognition engines have emerged. Here, an Automatic Speech Recognition (ASR) engine refers to an application program used to recognize speech data as text.

[0003] Due to the limitations of the prior art, the recognition rate of speech recognition engines cannot yet reach 100%. To meet the needs of different scenarios, most ASR engines support hotword input, that is, it is desired to improve the recognition probability of hotwords by inputting hotwords or hotwords and corresponding hotword speech data into the ASR engine. Hotwords are an important means of intervening in ASR recognition results. Summary of the Invention

[0004] Embodiments of the present disclosure propose speech recognition methods, devices, electronic devices, and storage media.

[0005] In a first aspect, embodiments of the present disclosure provide a speech recognition method, the method including: for a candidate word in a candidate word set, determining the recognition rate of speech recognition of the candidate word; generating a difficult-to-recognize candidate word set based on candidate words in the candidate word set whose recognition rate is less than a first preset recognition rate threshold; generating an easy-to-recognize candidate word set based on candidate words in the candidate word set whose recognition rate is greater than a second preset recognition rate threshold, where the second preset recognition rate threshold is greater than the first preset recognition rate threshold; respectively determining the hotword weight of the candidate words in the difficult-to-recognize candidate word set and the hotword weight of the candidate words in the easy-to-recognize candidate word set, and performing automatic speech recognition based on the determined hotword weights.

[0006] In some optional embodiments, the candidate word set is obtained in the following manner: performing keyword extraction on a reference text word segmentation sequence to generate a candidate word set, where the reference text word segmentation sequence is obtained by performing word segmentation on a reference text corresponding to target speech data, and the reference text is used to represent the actual speech text content corresponding to the target speech data.

[0007] In some optional embodiments, the performing keyword extraction on the reference text word segmentation sequence to generate a candidate word set includes: generating the candidate word set based on at least one of the part of speech of the reference text word segmentation in the reference text word segmentation sequence, the word frequency in the reference text word segmentation sequence, and whether it belongs to a preset stop word set.

[0008] In some alternative embodiments, for the candidate words in the candidate word set, determining the recognition rate of the speech recognition of the candidate word includes: performing word segmentation on the recognition text corresponding to the target speech data to obtain a word segmentation sequence of the recognition text; for the candidate words in the candidate word set, determining the recognition rate of the speech recognition of the candidate word according to the ratio of the occurrence frequency of the candidate word in the word segmentation sequence of the recognition text to the occurrence frequency in the word segmentation sequence of the reference text.

[0009] In some alternative embodiments, separately determining the hot word weights of the candidate words in the difficult-to-recognize candidate word set and the hot word weights of the candidate words in the easy-to-recognize candidate word set, and performing automatic speech recognition based on the determined hot word weights includes: for the candidate words in the candidate word set, determining the hot word weight of the candidate word in the first preset speech recognition engine according to the recognition rate of the candidate word, where the hot word weight of the candidate word is negatively correlated with the recognition rate of the candidate word; and using the candidate words in the candidate word set as hot words and inputting them into the first preset speech recognition engine according to the determined corresponding hot word weights to achieve automatic speech recognition.

[0010] In some alternative embodiments, separately determining the hot word weights of the candidate words in the difficult-to-recognize candidate word set and the hot word weights of the candidate words in the easy-to-recognize candidate word set, and performing automatic speech recognition based on the determined hot word weights includes: for the candidate words in the difficult-to-recognize candidate word set, determining the hot word weight of the candidate word when input as a hot word into the second preset speech recognition engine; and using the candidate words in the difficult-to-recognize candidate word set as hot words and inputting them into the second preset speech recognition engine according to the determined corresponding hot word weights to achieve automatic speech recognition.

[0011] In some alternative embodiments, before using the candidate words in the difficult-to-recognize candidate word set as hot words and inputting them into the second preset speech recognition engine according to the determined corresponding hot word weights to achieve automatic speech recognition, the method further includes: generating a recognizable candidate word set based on the candidate words in the candidate word set whose recognition rate is greater than or equal to the second preset recognition rate threshold and less than or equal to the second preset recognition rate threshold; for the candidate words in the recognizable candidate word set, determining the hot word weight of the candidate word when input as a hot word into the second preset speech recognition engine according to the recognition rate of the candidate word, where the hot word weight of the candidate word is negatively correlated with the recognition rate of the candidate word; and using the candidate words in the difficult-to-recognize candidate word set as hot words and inputting them into the second preset speech recognition engine according to the determined corresponding hot word weights to achieve automatic speech recognition includes: using the candidate words in the difficult-to-recognize candidate word set and the recognizable candidate word set as hot words and inputting them into the second preset speech recognition engine according to the determined hot word weights to achieve automatic speech recognition.

[0012] In a second aspect, an embodiment of the present disclosure provides a speech recognition device, which includes: a recognition rate determination unit configured to determine the recognition rate of speech recognition of a candidate word in a candidate word set; a difficult-to-recognize word generation unit configured to generate a difficult-to-recognize candidate word set based on candidate words in the candidate word set whose recognition rates are less than a first preset recognition rate threshold; an easy-to-recognize word generation unit configured to generate an easy-to-recognize candidate word set based on candidate words in the candidate word set whose recognition rates are greater than a second preset recognition rate threshold, where the second preset recognition rate threshold is greater than the first preset recognition rate threshold; and a speech recognition unit configured to respectively determine the hot word weights of the candidate words in the difficult-to-recognize candidate word set and the hot word weights of the candidate words in the easy-to-recognize candidate word set, and perform automatic speech recognition based on the determined hot word weights.

[0013] In some alternative embodiments, the candidate word set is obtained in the following manner: keyword extraction is performed on a reference text tokenization sequence to generate a candidate word set, where the reference text tokenization sequence is obtained by tokenizing a reference text corresponding to target speech data, and the reference text is used to represent the actual speech text content corresponding to the target speech data.

[0014] In some alternative embodiments, the keyword extraction from the reference text tokenization sequence to generate a candidate word set includes: generating the candidate word set based on at least one of the part of speech of the reference text tokens in the reference text tokenization sequence, the word frequency in the reference text tokenization sequence, and whether it belongs to a preset stop word set.

[0015] In some alternative embodiments, for a candidate word in the candidate word set, determining the recognition rate of speech recognition of the candidate word includes: performing tokenization processing on the recognition text corresponding to the target speech data to obtain a recognition text tokenization sequence; for a candidate word in the candidate word set, determining the recognition rate of speech recognition of the candidate word according to the ratio of the occurrence frequency of the candidate word in the recognition text tokenization sequence to the occurrence frequency in the reference text tokenization sequence.

[0016] In some alternative embodiments, the speech recognition unit is further configured to: for a candidate word in the candidate word set, determine the hot word weight of the candidate word in a first preset speech recognition engine according to the recognition rate of the candidate word, where the hot word weight of the candidate word is negatively correlated with the recognition rate of the candidate word; and use the candidate words in the candidate word set as hot words and input them into the first preset speech recognition engine according to the determined corresponding hot word weights to achieve automatic speech recognition.

[0017] In some alternative embodiments, the above voice recognition unit is further configured to: for the candidate words in the above set of hard-to-recognize candidate words, determine the hot word weight of the candidate word as a hot word input into the second preset voice recognition engine; and input the candidate words in the above set of hard-to-recognize candidate words as hot words into the above second preset voice recognition engine according to the determined corresponding hot word weights, so as to implement automatic speech recognition.

[0018] In some alternative embodiments, the above voice recognition unit is further configured to: generate a set of recognizable candidate words based on the candidate words in the above candidate word set whose recognition rate is greater than or equal to the above second preset recognition rate threshold and less than or equal to the above second preset recognition rate threshold; for the candidate words in the above set of recognizable candidate words, determine the hot word weight of the candidate word as a hot word input into the above second preset voice recognition engine based on the recognition rate of the candidate word, wherein the hot word weight of the candidate word is negatively correlated with the recognition rate of the candidate word; and input the candidate words in the above set of recognizable candidate words as hot words into the above second preset voice recognition engine according to the determined hot word weights, so as to implement automatic speech recognition.

[0019] In some alternative embodiments, the above device further includes: a recognizable candidate word determination unit, configured to: before inputting the candidate words in the above set of hard-to-recognize candidate words as hot words into the above second preset voice recognition engine according to the determined corresponding hot word weights to implement automatic speech recognition, generate a set of recognizable candidate words based on the candidate words in the above candidate word set whose recognition rate is greater than or equal to the above second preset recognition rate threshold and less than or equal to the above second preset recognition rate threshold; for the candidate words in the above set of recognizable candidate words, determine the hot word weight of the candidate word as a hot word input into the above second preset voice recognition engine based on the recognition rate of the candidate word, wherein the hot word weight of the candidate word is negatively correlated with the recognition rate of the candidate word; and the above inputting the candidate words in the above set of hard-to-recognize candidate words as hot words into the above second preset voice recognition engine according to the determined corresponding hot word weights to implement automatic speech recognition includes: inputting the candidate words in the above set of hard-to-recognize candidate words and the above set of recognizable candidate words as hot words into the above second preset voice recognition engine according to the determined hot word weights, so as to implement automatic speech recognition.

[0020] In a third aspect, an embodiment of the present disclosure provides an electronic device, including: one or more processors; a storage device, on which one or more programs are stored, and when the above one or more programs are executed by the above one or more processors, the above one or more processors are caused to implement the method described in any implementation manner in the first aspect.

[0021] Fourthly, an embodiment of the present disclosure provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by one or more processors, the method described in any implementation manner of the first aspect is implemented.

[0022] To better apply the hotword technology in the speech recognition process, the applicant has found through practical research that since the ASR engine itself has different recognition probabilities for different words. For example, some words are very easy to recognize, and these words are called "easy-to-recognize words"; conversely, some words are very difficult to recognize, called "difficult-to-recognize words". For easy-to-recognize words, since the ASR engine itself can recognize them well, these words as hotwords themselves do not have much meaning, and on the contrary, they may bring side effects, resulting in other words with similar sounds being easily misrecognized as this hotword. An effective hotword should be a difficult-to-recognize word with a low recognition probability by the ASR engine itself. Therefore, it is necessary to obtain the difficult-to-recognize and easy-to-recognize words of the ASR engine through some methods.

[0023] The speech recognition method, device, electronic device and storage medium provided by the embodiments of the present disclosure generate a set of difficult-to-recognize candidate words based on candidate words in the candidate word set with a recognition rate less than a first preset recognition rate threshold, and generate a set of easy-to-recognize candidate words based on candidate words in the candidate word set with a recognition rate greater than a second preset recognition rate threshold, where the second preset recognition rate threshold is greater than the first preset recognition rate threshold, that is, generate a set of easy-to-recognize candidate words based on candidate words in the candidate word set with a higher recognition rate (greater than the second preset recognition rate threshold), and generate a set of difficult-to-recognize candidate words based on candidate words in the candidate word set with a lower recognition rate (less than the first preset recognition rate threshold). Thus, the difficult-to-recognize and easy-to-recognize words of the ASR engine can be obtained, and then the hotword weights of the generated difficult-to-recognize and easy-to-recognize words are determined, and speech recognition is performed based on the corresponding hotword weights, thereby improving the recognition rate of difficult-to-recognize words and then improving the overall speech recognition rate. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] By reading the detailed description of the non-limiting embodiments with reference to the following drawings, other features, objects and advantages of the present disclosure will become more apparent. The drawings are only for the purpose of showing the specific implementation manners and are not considered as a limitation of the present invention. In the drawings:

[0025] Figure 1 is an exemplary system architecture diagram to which an embodiment of the present disclosure can be applied;

[0026] Figure 2A is a flowchart of an embodiment of the speech recognition method according to the present disclosure;

[0027] Figure 2B is a decomposed flowchart of an embodiment of step 204 according to the present disclosure;

[0028] Figure 2C is a decomposed flowchart of an embodiment of step 204 according to the present disclosure;

[0029] Figure 3 is a schematic structural diagram of an embodiment of the speech recognition device according to the present disclosure;

[0030] Figure 4 is a schematic structural diagram of a computer system of an electronic device suitable for implementing the embodiments of the present disclosure. Detailed implementation manners

[0031] The present disclosure will be further described in detail below with reference to the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the related invention, rather than limiting the invention. Additionally, it should be noted that for the convenience of description, only parts related to the relevant invention are shown in the drawings.

[0032] It should be noted that, without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other. The present disclosure will be described in detail below with reference to the drawings and embodiments.

[0033] Figure 1 An exemplary system architecture 100 is shown, which can apply the embodiments of the speech recognition method, device, electronic device, and storage medium of the present disclosure.

[0034] As Figure 1 shown, the system architecture 100 may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0035] Users can use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications may be installed on the terminal devices 101, 102, 103, such as text processing applications, speech recognition applications, short video social applications, web conferencing applications, web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.

[0036] The terminal devices 101, 102, and 103 can be either hardware or software. When the terminal devices 101, 102, and 103 are hardware, they can be various electronic devices with a display screen, including but not limited to smartphones, tablet computers, laptop portable computers, desktop computers, and so on. When the terminal devices 101, 102, and 103 are software, they can be installed in the terminal devices listed above. They can be implemented as multiple software or software modules (for example, to provide speech recognition services), or can be implemented as a single software or software module. Specific limitations are not made here.

[0037] In some cases, the speech recognition method provided by the present disclosure can be executed by the terminal devices 101, 102, and 103. Correspondingly, the speech recognition device can be set in the terminal devices 101, 102, and 103. At this time, the system architecture 100 may not include the server 105 either.

[0038] In some cases, the speech recognition method provided by the present disclosure can be jointly executed by the terminal devices 101, 102, and 103 and the server 105. The present disclosure does not make a limitation on this. Correspondingly, the speech recognition device can also be respectively set in the terminal devices 101, 102, and 103 and the server 105.

[0039] In some cases, the speech recognition method provided by the present disclosure can be executed by the server 105. Correspondingly, the speech recognition device can also be set in the server 105. At this time, the system architecture 100 may not include the terminal devices 101, 102, and 103 either.

[0040] It should be noted that the server 105 can be either hardware or software. When the server 105 is hardware, it can be implemented as a distributed server cluster composed of multiple servers, or can be implemented as a single server. When the server 105 is software, it can be implemented as multiple software or software modules (for example, to provide distributed services), or can be implemented as a single software or software module. Specific limitations are not made here.

[0041] It should be understood that Figure 1 the numbers of the terminal devices, the network, and the server in

[0042] Continue to refer to Figure 2A which shows a flow 200 of an embodiment of the speech recognition method according to the present disclosure. The speech recognition method includes the following steps:

[0043] Step 201, for the candidate words in the candidate word set, determine the recognition rate of the speech recognition of the candidate word.

[0044] In this embodiment, the execution subject of the speech recognition method (such as Figure 1 the terminal devices 101, 102, and 103 shown) can first determine the recognition rate of speech recognition for the candidate words in the candidate word set in various implementation manners.

[0045] As an example, the candidate word set can be dynamically learned from a large amount of corpus by using machine learning or data mining algorithms, or can be manually formulated by technicians according to experience.

[0046] Here, the recognition rate of speech recognition of the candidate word is used to characterize the recognition rate of the candidate word in the recognition process by at least one ASR engine. Specifically, for example, for the speech data of the same user or different users saying the candidate word, the above speech data can be sent to at least one ASR engine for recognition, and the recognition rate of speech recognition of the candidate word can be obtained by dividing the number of correct recognitions by the total number of speech data.

[0047] In some alternative embodiments, the candidate word set can be obtained in the following manner:

[0048] First, perform word segmentation on the reference text corresponding to the target speech data to obtain a reference text word segmentation sequence.

[0049] Here, the reference text corresponding to the target speech data is used to characterize the actual speech text content corresponding to the target speech data. For example, the target speech data can be the speech data recorded by the user according to the corresponding reference text. Another example is that the reference text corresponding to the target speech data can also be the content text obtained by manually annotating the speech content of the target speech data. Still another example is that the reference text corresponding to the target speech data can also be the text obtained by manually adjusting on the basis of the speech recognition result text of the target speech data.

[0050] Here, various currently known or future-developed word segmentation methods can be used to perform word segmentation on the reference text corresponding to the target speech data to obtain a reference text word segmentation sequence, and the present disclosure does not make specific limitations on this. For example, a word segmentation method based on string matching, a word segmentation method based on understanding, or a word segmentation method based on statistics, etc. can be used.

[0051] Then, extract keywords from the reference text word segmentation sequence to generate a candidate word set.

[0052] Here, various methods can be used to extract keywords from the reference text word segmentation sequence to generate a candidate word set.

[0053] In one embodiment, for example, methods such as TextRank, TF-IDF (Term Frequency / Inverse Document Frequency), Latent Dirichlet Allocation, keyword extraction based on semantics, and keyword extraction based on word vectors (such as word2vec) can be adopted.

[0054] In some alternative embodiments, the candidate word set can also be obtained in the following manner:

[0055] Based on at least one of the part-of-speech of the reference text segmentation in the reference text segmentation sequence, the word frequency in the reference text segmentation sequence, and whether it belongs to a preset stop word set, a candidate word set is generated. For example, the candidate word set can be generated by the segmentation words in the reference text segmentation sequence whose word frequency is greater than the preset word frequency threshold, does not belong to the preset stop word set, and whose part-of-speech does not belong to the preset set of function word part-of-speech. Furthermore, candidate words with relatively low occurrence frequencies, stop words, prepositions, conjunctions, auxiliary words, modal particles and other function words can be filtered out.

[0056] In some alternative embodiments, based on the alternative embodiments of the reference text segmentation sequence, step 201 can be executed as follows:

[0057] First, the recognition text corresponding to the target speech data is segmented to obtain a recognition text segmentation sequence.

[0058] Here, the recognition text corresponding to the target speech data can be the speech recognition result text obtained by performing automatic speech recognition on the target speech data.

[0059] Second, for the candidate words in the candidate word set, according to the ratio of the occurrence frequency of the candidate word in the recognition text segmentation sequence to the occurrence frequency in the reference text segmentation sequence, the recognition rate of the speech recognition of the candidate word is determined.

[0060] Here, it can be considered that the occurrence frequency of the candidate word in the recognition text segmentation sequence is the frequency of the candidate word being correctly recognized by automatic speech recognition, and the occurrence frequency of the candidate word in the reference text segmentation sequence can be considered as the frequency that the candidate word should appear or should be correctly recognized. Furthermore, the ratio of the occurrence frequency of the candidate word in the recognition text segmentation sequence to the occurrence frequency in the reference text segmentation sequence can represent the recognition rate of the speech recognition of the candidate word.

[0061] Here, a candidate word set can be generated based on at least one target speech data. This disclosure only takes the target speech data as an example for illustration.

[0062] Step 202: Generate a difficult-to-recognize candidate word set based on the candidate words in the candidate word set whose recognition rates are less than the first preset recognition rate threshold.

[0063] In this embodiment, the above-mentioned execution entity can adopt various implementation methods to generate a difficult-to-recognize candidate word set based on the candidate words in the candidate word set whose recognition rates are less than the first preset recognition rate threshold. Among them, the first preset recognition rate threshold can be a relatively small recognition rate threshold set in advance, which can be set to be constant or can be customized according to the actual situation. That is, if the recognition rate of a candidate word in the candidate word set is small, it can be considered that the possibility of this candidate word being a difficult-to-recognize word is relatively large.

[0064] For example, the above-mentioned execution entity can use the candidate words in the candidate word set whose recognition rates are less than the first preset recognition rate threshold to generate a difficult-to-recognize candidate word set.

[0065] For another example, the above-mentioned execution entity can first use the candidate words in the candidate word set whose recognition rates are less than the first preset recognition rate threshold to generate a first candidate word set, and then perform filtering processing on the first candidate word set to obtain a difficult-to-recognize candidate word set. For example, filter out infrequent words such as function words and auxiliary words and the words in the preset stop word set.

[0066] For still another example, the above-mentioned execution entity can use the candidate words in the candidate word set whose recognition rates are less than the first preset recognition rate threshold to merge with the manually specified difficult-to-recognize words to obtain a difficult-to-recognize word set.

[0067] Step 203: Generate an easy-to-recognize candidate word set based on the candidate words in the candidate word set whose recognition rates are greater than the second preset recognition rate threshold.

[0068] In this embodiment, the above-mentioned execution entity can adopt various implementation methods to generate an easy-to-recognize candidate word set based on the candidate words in the candidate word set whose recognition rates are greater than the second preset recognition rate threshold. Among them, the second preset recognition rate threshold can be a relatively large recognition rate threshold set in advance and greater than the first preset recognition rate threshold, which can be set to be constant or can be customized according to the actual situation. That is, if the recognition rate of a candidate word in the candidate word set is high, it can be considered that the possibility of this candidate word being an easy-to-recognize word is relatively large.

[0069] For example, the above-mentioned execution entity can use the candidate words in the candidate word set whose recognition rates are greater than the second preset recognition rate threshold to generate an easy-to-recognize candidate word set.

[0070] For another example, the above-mentioned execution entity can first use the candidate words in the candidate word set whose recognition rates are greater than the second preset recognition rate threshold to generate a second candidate word set, and then perform filtering processing on the second candidate word set to obtain an easy-to-recognize candidate word set. For example, filter out infrequent words such as function words and auxiliary words and the words in the preset stop word set.

[0071] For another example, the above-mentioned execution entity may merge the candidate words with recognition rates greater than the second preset recognition rate threshold in the candidate word set with the easily recognizable words specified manually to obtain an easily recognizable word set.

[0072] Step 204: Determine the hot word weights of the candidate words in the difficult-to-recognize candidate word set and the hot word weights of the candidate words in the easily recognizable candidate word set respectively, and perform automatic speech recognition based on the determined hot word weights.

[0073] In this embodiment, the above-mentioned execution entity may first adopt various implementation manners to determine the hot word weights of the candidate words in the difficult-to-recognize candidate word set and the hot word weights of the candidate words in the easily recognizable candidate word set respectively, and then perform automatic speech recognition based on the determined hot word weights.

[0074] Automatic Speech Recognition (ASR) is an algorithm technology that converts speech signals into text. During the process of speech recognition, it is necessary to model acoustic units with a neural network, abstract acoustic signals into acoustic feature vectors, and send them to a decoding network, where a language model is used to correct the decoding process online to determine the optimal decoding path, thereby obtaining the text of the speech to be recognized. In practice, it is very difficult for any ASR engine to correctly recognize all speech in all scenarios. To improve the recognition rate of speech recognition, a hot word (or called a feature word) mechanism is usually adopted to assist decoding.

[0075] Here, performing automatic speech recognition based on the determined hot word weights means that during the process of decoding the speech to be recognized, it is judged whether the decoding path contains candidate words in the difficult-to-recognize candidate word set or the easily recognizable candidate word set. When the decoding path contains such candidate words, the corresponding decoding path is weight-stimulated according to the hot word weights of the corresponding candidate words to improve the recognition accuracy of the candidate words, and further improve the recognition rate of speech recognition.

[0076] Specifically, it can be carried out as follows:

[0077] First, perform recognition processing on the speech to be recognized to obtain at least one candidate character corresponding to each time unit of the speech to be recognized, and the acoustic model scores of each candidate character.

[0078] The acoustic model is a differentiated representation of acoustic, phonetic, environmental variables, speaker gender, accent, etc. Acoustic models may include, but are not limited to, acoustic models based on Hidden Markov Model (HMM) and end-to-end acoustic models. Among them, the acoustic model of HMM may include, for example, Gaussian HMM and deep neural network HMM, and the end-to-end acoustic model may include, for example, connectionist temporal classification (CTC) model, long-short term memory (LSTM) model and attention model, etc. The acoustic model may perform recognition processing on the speech to be recognized, and obtain at least one candidate word corresponding to each time unit of the speech to be recognized, and the acoustic model score of each candidate word. Among them, each time unit corresponds to a word of the text, each time unit may include one or more candidate words, and each candidate word corresponds to a decoding path. One or more candidate words corresponding to each time unit may be words with the same or similar pronunciation. For example, the candidate characters corresponding to a certain time unit include "动", "栋", and "洞", and these three candidate characters correspond to three different decoding paths.

[0079] Then, the recognized text may be obtained according to the hot word weight of each candidate word in the difficult-to-recognize candidate word set and the easy-to-recognize candidate word set, at least one candidate word corresponding to each time unit, and the acoustic model score of each candidate word.

[0080] Specifically, the language model score of each candidate word corresponding to the time unit i may be first obtained, where i is 1, 2, 3, ..., n, n is the number of words in the text, and each time unit corresponds to one word in the text.

[0081] After obtaining the language model score of each candidate word, the hot word weight of each candidate word corresponding to the time unit i is obtained according to the hot word weight of each candidate word in the difficult-to-identify candidate word set and the easy-to-identify candidate word set.

[0082] After determining the hot word weight of each candidate word, the i-th word of the text can be determined in each candidate word corresponding to time unit i according to the acoustic model score, language model score and hot word weight of each candidate word corresponding to time unit i. Then, the recognized text corresponding to the speech to be recognized is obtained to complete automatic speech recognition.

[0083] It should be noted that for easily recognizable words, since the ASR engine itself can recognize them well, these words are not very meaningful as hot words. On the contrary, they may cause side effects and lead to other words with similar pronunciations being easily misrecognized as this hot word. An effective hot word should be a difficult-to-recognize word with a low recognition probability by the ASR engine itself. Therefore, here, when determining the hot word weights of the candidate words in the difficult-to-recognize candidate word set and the hot word weights of the candidate words in the easily recognizable candidate word set, the hot word weights of the candidate words in the difficult-to-recognize candidate word set should be greater than the hot word weights of the candidate words in the easily recognizable candidate word set. Since the recognition rate of the candidate words in the original difficult-to-recognize candidate word set is relatively low, that is, in the case of similar pronunciations, the probability of being recognized as a candidate word in the difficult-to-recognize candidate word set is relatively low. However, due to setting a relatively high hot word weight for the candidate words in the difficult-to-recognize candidate word set, in the case of words with similar pronunciations, the probability of being recognized as a candidate word in the difficult-to-recognize candidate word set will be increased, and thus the recognition rate of the speech recognition of the candidate words in the difficult-to-recognize candidate word set can be improved.

[0084] In some alternative embodiments, step 204 may include steps 2041a and 2042a as Figure 2B shown:

[0085] Step 2041a, for a candidate word in the candidate word set, determine the hot word weight of the candidate word in the first preset speech recognition engine according to the recognition rate of the candidate word.

[0086] In this embodiment, the execution entity of the speech recognition method (such as Figure 1 the terminal devices 101, 102, 103 shown) may, for a candidate word in the candidate word set, adopt various implementation manners to determine the hot word weight of the candidate word in the first preset speech recognition engine according to the recognition rate of the candidate word. Among them, the hot word weight of the candidate word is negatively correlated with the recognition rate of the candidate word. That is, the higher the recognition rate of the candidate word, the easier it is for the candidate word to be recognized, which means the smaller the meaning of the candidate word as a hot word, that is, the lower the hot word weight input to the first preset speech recognition engine. On the contrary, the lower the recognition rate of the candidate word, the more difficult it is for the candidate word to be recognized, which means the greater the meaning of the candidate word as a hot word, that is, the higher the hot word weight input to the first preset speech recognition engine. The hot word weight of the candidate word may be linearly negatively correlated or non-linearly negatively correlated with the corresponding recognition rate.

[0087] Suppose the value range of the recognition rate P1 of the candidate word is between 0 and N1, and the value range of the hot word weight P2 of the candidate word is between 0 and N2. Then, for example, the hot word weight of the candidate word can be calculated according to the following formula 1:

[0088] P2 = N2 × (1 - P1 / N1) (Formula 1)

[0089] Step 2042a: Use the candidate words in the candidate word set as hot words and input them into the first preset speech recognition engine according to the determined corresponding hot word weights to achieve automatic speech recognition.

[0090] It should be noted that the first preset speech recognition engine here can be a speech recognition engine that supports the input of hot words and corresponding hot word weights. By inputting the hot words and corresponding hot word weights into the first preset speech recognition engine together, in the process of using the first preset speech recognition engine for automatic speech recognition, since the hot word weights of candidate words with higher recognition rates are lower, while the hot word weights of candidate words with lower recognition rates are higher, that is, lower hot word weights are assigned to words that are originally easier to recognize, and higher hot word weights are assigned to words that are originally more difficult to recognize, the recognition rate of words that are originally more difficult to recognize can be improved, thus achieving the effect of using hot words to improve the recognition rate of speech recognition.

[0091] In some alternative embodiments, step 204 may include steps 2041b and 2042b as shown in Figure 2C the following:

[0092] Step 2041b: For the candidate words in the difficult-to-recognize candidate word set, determine the hot word weight for inputting this candidate word as a hot word into the second preset speech recognition engine.

[0093] Here, for the candidate words in the difficult-to-recognize candidate word set generated in step 202, various implementation methods can be used to determine the hot word weight for inputting this candidate word as a hot word into the second preset speech recognition engine.

[0094] Optionally, the candidate words in all difficult-to-recognize candidate word sets can be set with the same and relatively high hot word weights.

[0095] For example, under the assumption of the alternative embodiments of the above steps 2041a and 2042b, the recognition rate P1 of the candidate words ranges from 0 to N1, the hot word weight P2 of the candidate words ranges from 0 to N2, and the average recognition rate of each candidate word in the difficult-to-recognize candidate words is P 1-mean , the hot word weight of the candidate words in all difficult-to-recognize candidate words can be calculated according to the following formula 2:

[0096] P2 = N2 × (1 - P 1-mean ) (Formula 2)

[0097] According to this alternative method, since the recognition rates of the candidate words in the difficult-to-recognize candidate word set are relatively low, the average recognition rate P 1-mean of each candidate word is also relatively low, and thus the hot word weights calculated according to formula 2 will be relatively high.

[0098] In some alternative embodiments, it is also possible to determine that the hotword weight of the candidate word is negatively correlated with the recognition rate of the candidate word when determining the hotword weight of the candidate word. For example, Figure 2B The implementation manner shown in Formula 1 in [reference] is not elaborated here again.

[0099] Step 2042b: Use the candidate words in the set of difficult-to-recognize candidate words as hotwords and input them into the second preset speech recognition engine according to the determined corresponding hotword weights to achieve automatic speech recognition.

[0100] Here, the second preset speech recognition engine can be a speech recognition engine that supports the input of hotwords and corresponding hotword weights. By inputting the hotwords and corresponding hotword weights into the second preset speech recognition engine together, in the process of using the second preset speech recognition engine for automatic speech recognition, since only the candidate words with a lower recognition rate and the corresponding hotword weights are input, while the words with a higher original recognition rate, that is, the words that are relatively easy to recognize, are filtered out and not input into the second preset speech recognition engine as hotwords. Therefore, only the difficult-to-recognize words will be input into the second preset speech recognition engine, which can improve the recognition rate of the difficult-to-recognize words, and thus achieve the effect of improving the recognition rate of speech recognition by using hotwords.

[0101] In some alternative embodiments, before step 2042b, the following steps 2043b and 2044b may further be included:

[0102] Step 2043b: Generate a set of recognizable candidate words based on the candidate words in the candidate word set whose recognition rate is greater than or equal to the first preset recognition rate threshold and less than or equal to the second preset recognition rate threshold.

[0103] Here, the above-mentioned execution subject can adopt various implementation manners to generate a set of recognizable candidate words based on the candidate words in the candidate word set whose recognition rate is greater than or equal to the first preset recognition rate threshold and less than or equal to the second preset recognition rate threshold.

[0104] For example, the above-mentioned execution subject can generate a set of recognizable candidate words with the candidate words in the candidate word set whose recognition rate is greater than or equal to the first preset recognition rate threshold and less than or equal to the second preset recognition rate threshold.

[0105] For another example, the above-mentioned execution subject can first generate a second candidate word set with the candidate words in the candidate word set whose recognition rate is greater than or equal to the first preset recognition rate threshold and less than or equal to the second preset recognition rate threshold, and then perform filtering processing on the second candidate word set to obtain a set of recognizable candidate words, such as filtering out infrequent words such as function words and auxiliary words and the words in the preset stop word set.

[0106] For another example, the above-mentioned execution entity may combine the candidate words with recognition rates greater than or equal to the first preset recognition rate threshold and less than or equal to the second preset recognition rate threshold in the candidate word set with the manually specified recognizable words to obtain a recognizable word set.

[0107] Step 2044b: For the candidate words in the recognizable candidate word set, determine the hot word weight of the candidate word as a hot word input into the second preset speech recognition engine based on the recognition rate of the candidate word.

[0108] Here, the above-mentioned execution entity may determine the hot word weight of the candidate word as a hot word input into the second preset speech recognition engine based on the recognition rate of the candidate word in the recognizable candidate word set, where the hot word weight of the candidate word in the recognizable candidate word is negatively correlated with the recognition rate of the candidate word. The specific method for determining the hot word weight may refer to the relevant records in the optional implementation manners of Step 2041a and Step 2042a, which will not be elaborated here.

[0109] Correspondingly, Step 2042b may be executed as follows:

[0110] Use the hard-to-recognize candidate words and the candidate words in the recognizable candidate word set as hot words, and input them into the second preset speech recognition engine according to the determined corresponding hot word weights to achieve automatic speech recognition.

[0111] By using the hard-to-recognize candidate words and the candidate words in the recognizable candidate word set as hot words and inputting them into the second preset speech recognition engine together according to the determined corresponding hot word weights, in the process of using the second preset speech recognition engine for automatic speech recognition, since hard-to-recognize words with low recognition rates and recognizable words with medium recognition rates are input, the recognition rates of the hard-to-recognize words with originally low recognition rates and the recognizable words with medium recognition rates can both be improved, thus achieving the effect of improving the recognition rate of speech recognition by using hot words.

[0112] The speech recognition method provided by the above embodiments of the present disclosure generates a set of difficult-to-recognize candidate words based on the candidate words in the candidate word set with a recognition rate less than the first preset recognition rate threshold, and generates a set of easy-to-recognize candidate words based on the candidate words in the candidate word set with a recognition rate greater than the second preset recognition rate threshold, where the second preset recognition rate threshold is greater than the first preset recognition rate threshold, that is, a set of easy-to-recognize candidate words is generated based on the candidate words in the candidate word set with a higher recognition rate (greater than the second preset recognition rate threshold), and a set of difficult-to-recognize candidate words is generated based on the candidate words in the candidate word set with a lower recognition rate (less than the first preset recognition rate threshold). Thus, the difficult and easy-to-recognize words of the ASR engine can be obtained, and then the hot word weights of the generated difficult and easy-to-recognize words are determined, and speech recognition is performed based on the corresponding hot word weights, thereby improving the recognition rate of difficult-to-recognize words and then improving the overall speech recognition rate. Optionally, by filtering out the set of easy-to-recognize candidate words and only inputting the difficult-to-recognize candidate words or only the difficult-to-recognize candidate words and recognizable candidate words into the speech recognition engine, when the hot words are input to the speech recognition engine according to the corresponding hot word weights, the recognition rate of the candidate words that were originally difficult to recognize or had a medium recognition rate can be improved, further improving the recognition rate of using hot words to improve speech recognition.

[0113] Further referring to Figure 3 , as an implementation of the methods shown in the above figures, the present disclosure provides an embodiment of a speech recognition device. This device embodiment corresponds to Figure 2A the method embodiment shown, and this device can be specifically applied to various electronic devices.

[0114] As Figure 3 shown, the speech recognition device 300 of this embodiment includes: a recognition rate determination unit 301, a difficult-to-recognize word generation unit 302, and an easy-to-recognize word generation unit 303. Among them, the recognition rate determination unit 301 is configured to determine the speech recognition rate of a candidate word in the candidate word set; the difficult-to-recognize word generation unit 302 is configured to generate a set of difficult-to-recognize candidate words based on the candidate words in the candidate word set with a recognition rate less than the first preset recognition rate threshold; and the easy-to-recognize word generation unit 303 is configured to generate a set of easy-to-recognize candidate words based on the candidate words in the candidate word set with a recognition rate greater than the second preset recognition rate threshold, where the second preset recognition rate threshold is greater than the first preset recognition rate threshold; and a speech recognition unit 304 is configured to respectively determine the hot word weights of the candidate words in the set of difficult-to-recognize candidate words and the hot word weights of the candidate words in the set of easy-to-recognize candidate words, and perform automatic speech recognition based on the determined hot word weights.

[0115] In this embodiment, for the specific processing of the recognition rate determination unit 301, difficult word generation unit 302, easy word generation unit 303, and speech recognition unit 304 of the speech recognition device 300 and the technical effects brought thereby, reference can be respectively made to Figure 2A the relevant descriptions of steps 201, step 202, step 203, and step 204 in the corresponding embodiment, which will not be elaborated here.

[0116] In some alternative embodiments, the above candidate word set can be obtained in the following manner: keyword extraction is performed on the reference text segmentation sequence to generate a candidate word set, where the above reference text segmentation sequence is obtained by performing word segmentation processing on the reference text corresponding to the target speech data, and the above reference text is used to represent the actual speech text content corresponding to the above target speech data.

[0117] In some alternative embodiments, the above keyword extraction is performed on the reference text segmentation sequence to generate a candidate word set, which may include: generating the above candidate word set based on at least one of the part of speech of the reference text segmentation in the above reference text segmentation sequence, the word frequency in the above reference text segmentation sequence, and whether it belongs to a preset stop word set.

[0118] In some alternative embodiments, for the candidate words in the above candidate word set, determining the recognition rate of speech recognition of the candidate word may include: performing word segmentation processing on the recognition text corresponding to the above target speech data to obtain a recognition text segmentation sequence; for the candidate words in the above candidate word set, determining the recognition rate of speech recognition of the candidate word according to the ratio of the occurrence frequency of the candidate word in the above recognition text segmentation sequence to the occurrence frequency in the above reference text segmentation sequence.

[0119] In some alternative embodiments, the above speech recognition unit 304 may be further configured to: for the candidate words in the above candidate word set, determine the hot word weight of the candidate word in the first preset speech recognition engine according to the recognition rate of the candidate word, where the hot word weight of the candidate word is negatively correlated with the recognition rate of the candidate word; and use the candidate words in the above candidate word set as hot words and input them into the first preset speech recognition engine according to the determined corresponding hot word weights to implement automatic speech recognition.

[0120] In some alternative embodiments, the above speech recognition unit 304 may be further configured to: for the candidate words in the above difficult-to-recognize candidate word set, determine the hot word weight of the candidate word as a hot word input into the second preset speech recognition engine; and use the candidate words in the above difficult-to-recognize candidate word set as hot words and input them into the second preset speech recognition engine according to the determined corresponding hot word weights to implement automatic speech recognition.

[0121] In some alternative embodiments, the above-mentioned speech recognition unit 304 may be further configured to: generate a set of recognizable candidate words based on the candidate words in the above-mentioned candidate word set whose recognition rates are greater than or equal to the above-mentioned second preset recognition rate threshold and less than or equal to the above-mentioned second preset recognition rate threshold; for the candidate words in the above-mentioned set of recognizable candidate words, determine the hot word weight of the candidate word as a hot word input into the second preset speech recognition engine based on the recognition rate of the candidate word, wherein the hot word weight of the candidate word is negatively correlated with the recognition rate of the candidate word; and input the candidate words in the above-mentioned set of recognizable candidate words as hot words into the second preset speech recognition engine according to the determined hot word weights, so as to achieve automatic speech recognition.

[0122] In some alternative embodiments, the above-mentioned apparatus 300 may further include: a recognizable candidate word determination unit ( Figure 3 not shown in the figure), configured to: before inputting the candidate words in the above-mentioned set of difficult-to-recognize candidate words as hot words into the second preset speech recognition engine according to the determined corresponding hot word weights to achieve automatic speech recognition, generate a set of recognizable candidate words based on the candidate words in the above-mentioned candidate word set whose recognition rates are greater than or equal to the above-mentioned second preset recognition rate threshold and less than or equal to the above-mentioned second preset recognition rate threshold; for the candidate words in the above-mentioned set of recognizable candidate words, determine the hot word weight of the candidate word as a hot word input into the second preset speech recognition engine based on the recognition rate of the candidate word, wherein the hot word weight of the candidate word is negatively correlated with the recognition rate of the candidate word; and the above-mentioned inputting the candidate words in the above-mentioned set of difficult-to-recognize candidate words as hot words into the second preset speech recognition engine according to the determined corresponding hot word weights to achieve automatic speech recognition includes: inputting the candidate words in the above-mentioned set of difficult-to-recognize candidate words and the above-mentioned set of recognizable candidate words as hot words into the second preset speech recognition engine according to the determined hot word weights to achieve automatic speech recognition.

[0123] It should be noted that the implementation details and technical effects of each unit in the speech recognition apparatus provided in the embodiments of the present disclosure may refer to the descriptions of other embodiments in the present disclosure, and will not be elaborated herein.

[0124] Next, refer to Figure 4 , which shows a schematic structural diagram of a computer system 400 suitable for implementing the electronic device of the present disclosure. Figure 4 The shown computer system 400 is only an example and should not impose any limitations on the functions and usage scopes of the embodiments of the present disclosure.

[0125] As Figure 4As shown, the computer system 400 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 401, which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 402 or a program loaded from a storage device 408 into a random access memory (RAM) 403. In the RAM 403, various programs and data required for the operation of the computer system 400 are also stored. The processing device 401, the ROM 402, and the RAM 403 are connected to each other through a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.

[0126] Generally, the following devices may be connected to the I / O interface 405: an input device 406 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, etc.; an output device 407 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 408 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 409. The communication device 409 may allow the computer system 400 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 4 a computer system 400 of an electronic device with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices may be implemented or had.

[0127] Specifically, according to an embodiment of the present disclosure, the process described above with reference to the flowchart may be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for performing the method shown in the flowchart. In such an embodiment, the computer program may be downloaded and installed from a network through the communication device 409, or installed from the storage device 408, or installed from the ROM 402. When the computer program is executed by the processing device 401, the above functions defined in the method of the embodiment of the present disclosure are executed.

[0128] It should be noted that the above-mentioned computer-readable medium in the present disclosure can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the present disclosure, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable signal medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0129] The above-mentioned computer-readable medium can be included in the above-mentioned electronic device; or it can exist separately and not be assembled into the electronic device.

[0130] The above-mentioned computer-readable medium carries one or more programs. When the above one or more programs are executed by the electronic device, the electronic device implements the speech recognition method shown in the embodiment and its optional implementation manners shown in FIG. 2, and / or, as Figure 3 shown in the embodiment and its optional implementation manners shown in the figure.

[0131] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or it may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0132] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a portion of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system that performs the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.

[0133] The units involved in the embodiments described in the present disclosure may be implemented in software or in hardware. Among them, the name of the unit does not, in some cases, constitute a limitation on the unit itself. For example, the recognition rate determination unit may also be described as "a unit that determines the recognition rate of speech recognition for a candidate word in a candidate word set".

[0134] The above description is only a preferred embodiment of the present disclosure and an explanation of the technical principles applied. Those skilled in the art should understand that the scope of the disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with other technical features (but not limited to) having similar functions disclosed in the present disclosure.

Claims

1. A speech recognition method, comprising: For a candidate word in a candidate word set, determining a recognition rate of speech recognition of the candidate word; Generating a difficult-to-recognize candidate word set based on candidate words in the candidate word set whose recognition rates are less than a first preset recognition rate threshold; Generating an easy-to-recognize candidate word set based on candidate words in the candidate word set whose recognition rates are greater than a second preset recognition rate threshold, where the second preset recognition rate threshold is greater than the first preset recognition rate threshold; Respectively determining a hot word weight of a candidate word in the difficult-to-recognize candidate word set and a hot word weight of a candidate word in the easy-to-recognize candidate word set, and performing weight excitation on corresponding decoding paths according to the hot word weights of the candidate words to improve the accuracy of candidate word recognition, where the hot word weight of the candidate word is negatively correlated with the recognition rate of the candidate word.

2. The method according to claim 1, wherein, The candidate word set is obtained by the following method: Performing keyword extraction on a reference text segmentation sequence to generate a candidate word set, where the reference text segmentation sequence is obtained by performing segmentation processing on a reference text corresponding to target speech data, and the reference text is used to represent the actual speech text content corresponding to the target speech data.

3. The method according to claim 2, wherein, The performing keyword extraction on the reference text segmentation sequence to generate a candidate word set includes: Generating the candidate word set based on at least one of the part of speech of a reference text segmentation in the reference text segmentation sequence, the word frequency in the reference text segmentation sequence, and whether it belongs to a preset stop word set.

4. The method according to claim 3, wherein The determining a recognition rate of speech recognition of a candidate word in the candidate word set includes: Performing segmentation processing on a recognition text corresponding to the target speech data to obtain a recognition text segmentation sequence; For a candidate word in the candidate word set, determining the recognition rate of speech recognition of the candidate word according to the ratio of the occurrence frequency of the candidate word in the recognition text segmentation sequence to the occurrence frequency in the reference text segmentation sequence.

5. The method according to claim 1, wherein The respectively determining a hot word weight of a candidate word in the difficult-to-recognize candidate word set and a hot word weight of a candidate word in the easy-to-recognize candidate word set, and performing automatic speech recognition based on the determined hot word weights includes: For a candidate word in the candidate word set, determining the hot word weight of the candidate word in a first preset speech recognition engine according to the recognition rate of the candidate word; and Taking the candidate words in the candidate word set as hot words and inputting them into the first preset speech recognition engine according to the determined corresponding hot word weights to implement automatic speech recognition.

6. The method according to claim 1, wherein The respectively determining a hot word weight of a candidate word in the difficult-to-recognize candidate word set and a hot word weight of a candidate word in the easy-to-recognize candidate word set, and performing automatic speech recognition based on the determined hot word weights includes: For a candidate word in the difficult-to-recognize candidate word set, determining the hot word weight of the candidate word as a hot word input into a second preset speech recognition engine; and Taking the candidate words in the difficult-to-recognize candidate word set as hot words and inputting them into the second preset speech recognition engine according to the determined corresponding hot word weights to implement automatic speech recognition.

7. The method according to claim 6, wherein, Before using the candidate words in the set of hard-to-recognize candidate words as hot words and inputting them into the second preset speech recognition engine according to the determined corresponding hot word weights to achieve automatic speech recognition, the method further includes: Generating a set of recognizable candidate words based on the candidate words in the candidate word set whose recognition rates are greater than or equal to the second preset recognition rate threshold and less than or equal to the second preset recognition rate threshold; For the candidate words in the set of recognizable candidate words, determining the hot word weights for inputting the candidate words as hot words into the second preset speech recognition engine based on the recognition rates of the candidate words, where the hot word weights of the candidate words are negatively correlated with the recognition rates of the candidate words; and The step of using the candidate words in the set of hard-to-recognize candidate words as hot words and inputting them into the second preset speech recognition engine according to the determined corresponding hot word weights to achieve automatic speech recognition includes: Using the candidate words in the set of hard-to-recognize candidate words and the set of recognizable candidate words as hot words and inputting them into the second preset speech recognition engine according to the determined hot word weights to achieve automatic speech recognition.

8. A speech recognition device, comprising: A recognition rate determination unit configured to determine the recognition rate of speech recognition of a candidate word for a candidate word in a candidate word set; A hard-to-recognize word generation unit configured to generate a set of hard-to-recognize candidate words based on the candidate words in the candidate word set whose recognition rates are less than a first preset recognition rate threshold; An easy-to-recognize word generation unit configured to generate a set of easy-to-recognize candidate words based on the candidate words in the candidate word set whose recognition rates are greater than a second preset recognition rate threshold, where the second preset recognition rate threshold is greater than the first preset recognition rate threshold; A speech recognition unit configured to respectively determine the hot word weights of the candidate words in the set of hard-to-recognize candidate words and the hot word weights of the candidate words in the set of easy-to-recognize candidate words, and perform weight excitation on the corresponding decoding paths according to the hot word weights of the candidate words to improve the accuracy of candidate word recognition, where the hot word weights of the candidate words are negatively correlated with the recognition rates of the candidate words.

9. An electronic device, comprising: One or more processors; A storage device having one or more programs stored thereon, When the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the method according to any one of claims 1-7.

10. A computer-readable storage medium having a computer program stored thereon, wherein, The computer program, when executed by one or more processors, implements the method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Audio call assessment method, assessment equipment and computer storage medium

    CN109427327A

  • Speech recognition method, device and apparatus

    CN109523991A

  • Editing support device, editing support method, and program

    CN110136720A