Speech recognition method and related products thereof

By determining the whole-word score of hot words and constructing a new lexicon, the problem of whole-word excitation failure in the speech recognition model was solved, thereby improving the accuracy of hot-word speech recognition and the model's recognition ability.

CN115312041BActive Publication Date: 2025-11-07IFLYTEK CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210945100.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-08
Publication Date
2025-11-07
Estimated Expiration
2042-08-08

AI Technical Summary

Technical Problem

Existing speech recognition models are prone to failure in hot word activation due to the failure of activation of single characters or sub-words, which reduces the accuracy of hot word speech recognition.

Method used

By acquiring the acoustic features of speech data and combining them with hot words in the hot word library, the whole word score of the hot words is determined, and the whole word score is used to stimulate the hot words, avoiding stimulating single characters or sub-words one by one. A new word library is constructed to expand the word library of the speech recognition model, and the recognition accuracy is improved through hard example mining and model updates.

Benefits of technology

It improves the accuracy of hot word speech recognition, reduces the amount of computation, enhances the model's ability to distinguish similar hot words, and avoids the problem of whole word excitation failure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115312041B_ABST
    Figure CN115312041B_ABST
Patent Text Reader

Abstract

The application discloses a speech recognition method and related products. The method can include: obtaining speech data and a hot word library; the hot word library includes hot words; determining acoustic features of the speech data according to the speech data; determining whole word scores of the hot words based on the hot words in the hot word library and the acoustic features; and performing hot word excitation on the speech data using the whole word scores of the hot words. By determining the whole word scores of the hot words in the hot word library, the hot word excitation can be directly performed according to the whole word scores when the hot word excitation is performed, thereby avoiding the problem of whole word excitation failure caused by excitation one by one according to the scores of single characters or sub-words, and improving the accuracy of hot word speech recognition.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of speech processing, and in particular, to a speech recognition method and a related product thereof. BACKGROUND

[0002] At present, speech recognition technology has gradually become an important way of human-computer interaction and has been widely applied in intelligent mobile devices, intelligent customer service, smart home and other fields. In actual application, due to the differences in age, occupation, social network, interest and hobby of different users, and the emergence of new hot topics and corresponding hot words, some words with user characteristics and timeliness often appear in the speech recognition scene, which are referred to as "hot words".

[0003] Based on the above speech recognition requirements for hot words, the existing speech recognition model usually uses the way of hot word excitation of speech data to realize recognition, that is, if the decoding result of the speech data matches the pre-set hot word, the score of the hot word is excited to increase the output probability of the hot word, so that the hot word appears in the optimal output path. However, the process of hot word excitation needs to excite each single word or sub-word in the hot word according to the score, and in this case, if a single word or sub-word excitation fails, or the decoding of the first single word or sub-word is stopped and the decoding is interrupted, the problem of whole word excitation failure will occur. Therefore, once the above way of exciting hot words according to the score of single word or sub-word in the hot word occurs, the accuracy of speech recognition of hot words will be insufficient. SUMMARY

[0004] Embodiments of the present application provide a speech recognition method and a related product thereof to improve the accuracy of speech recognition of hot words.

[0005] In a first aspect, embodiments of the present application provide a speech recognition method, comprising:

[0006] obtaining speech data and a hot word library; the hot word library comprises hot words;

[0007] determining acoustic features of the speech data according to the speech data;

[0008] determining whole word scores of the hot words in the hot word library based on the hot words and the acoustic features;

[0009] exciting hot words of the speech data by using the whole word scores of the hot words.

[0010] Optionally, the determining the whole word scores of the hot words in the hot word library based on the hot words and the acoustic features comprises:

[0011] determine a word vector of the hotword according to hotword information corresponding to the hotword;

[0012] obtain a first acoustic semantic vector corresponding to the acoustic feature through an attention mechanism;

[0013] calculate a similarity between the word vector of the hotword and the first acoustic semantic vector as a whole-word score of the hotword.

[0014] Optionally, the method is implemented through a speech recognition model; the method further includes:

[0015] obtain a basic vocabulary of the speech recognition model; the basic vocabulary includes single characters and / or subwords;

[0016] construct a new vocabulary based on the single characters and / or subwords in the basic vocabulary and the hotwords in the hotword library;

[0017] perform hotword activation on the speech data based on the new vocabulary.

[0018] Optionally, the constructing of the new vocabulary based on the single characters and / or subwords in the basic vocabulary and the hotwords in the hotword library includes:

[0019] obtain scores of the single characters and / or subwords;

[0020] perform splicing processing on the scores of the single characters and / or subwords and the whole-word score of the hotword to generate the new vocabulary; the new vocabulary is constructed with the single characters and / or subwords and the hotwords as elements.

[0021] Optionally, the performing of the hotword activation on the speech data based on the new vocabulary includes:

[0022] match a decoding result of the speech data with the new vocabulary;

[0023] when a matching degree of a hotword in the new vocabulary with the decoding result is greater than or equal to a preset matching degree, perform hotword activation on the speech data based on scores of elements in the new vocabulary.

[0024] Optionally, the method further includes:

[0025] obtain a preset hotword dictionary for speech recognition;

[0026] determine a similarity between each two preset hotwords in the preset hotword dictionary, and sort the similarity between each two preset hotwords in a descending order;

[0027] determine a similar hotword of each preset hotword from the preset hotword dictionary based on a sorting result.

[0028] If the hot word corresponding to the voice data exists in both the hot word library and the preset hot word dictionary, one of similar hot words of the hot word corresponding to the voice data is added into the hot word library.

[0029] Optionally, the method further comprises:

[0030] obtaining a decoded result of the voice data before a current time point;

[0031] obtaining a second acoustic semantic vector corresponding to the decoded result through an attention mechanism;

[0032] analyzing the second acoustic semantic vector to obtain first syllable information of the second acoustic semantic vector;

[0033] based on the first syllable information, deleting a hot word that does not match the first syllable information from the hot word library, and performing voice recognition based on the hot word library after the hot word that does not match the first syllable information is deleted.

[0034] Optionally, the method is implemented through a voice recognition model; and the method further comprises:

[0035] determining a target function based on a Softmax loss function;

[0036] updating the voice recognition model by using the target function.

[0037] Optionally, the determining the target function based on the Softmax loss function comprises:

[0038] determining a number of training samples of the voice recognition model; the training samples comprise hot words in the hot word library;

[0039] calculating the Softmax loss function according to the number of training samples and a whole-word score of the hot words in the hot word library, and taking the Softmax loss function as the target function.

[0040] In a second aspect, an embodiment of the present application provides a voice recognition device, comprising:

[0041] a data acquisition module configured to acquire voice data and a hot word library; the hot word library comprises hot words;

[0042] an acoustic feature determination module configured to determine acoustic features of the voice data according to the voice data;

[0043] a whole-word score determination module configured to determine a whole-word score of the hot words based on the hot words in the hot word library and the acoustic features;

[0044] A hotword excitation module is configured to excite the speech data using the hotword score of the hotword.

[0045] In a third aspect, an embodiment of the present application provides a speech recognition device, the device comprising: a processor, a memory, a system bus;

[0046] The processor and the memory are connected through the system bus;

[0047] The memory is configured to store one or more programs, the one or more programs comprising instructions that, when executed by the processor, cause the processor to perform the method described above.

[0048] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, the computer-readable storage medium storing instructions, when the instructions are run on a terminal device, causing the terminal device to perform the method described above.

[0049] From the above technical solutions, it can be seen that the embodiments of the present application have the following advantages:

[0050] In the embodiments of the present application, after obtaining the speech data and the hotword library, the acoustic features of the speech data can be determined according to the speech data, and the hotword score of the hotword can be determined based on the hotword and the acoustic features stored in the hotword library, and then the speech data can be excited by the hotword based on the hotword score. It can be seen that by determining the hotword score of the hotword in the hotword library, the hotword can be excited directly according to the hotword score during hotword excitation, so that the problem of failure of whole-word excitation caused by exciting each word or sub-word one by one can be avoided, thereby improving the accuracy of hotword speech recognition. BRIEF DESCRIPTION OF DRAWINGS

[0051] Figure 1 A flowchart of a speech recognition method provided by an embodiment of the present application;

[0052] Figure 2a A flowchart of an implementation of constructing a new word library provided by an embodiment of the present application;

[0053] Figure 2b A schematic diagram of an implementation of constructing a new word library provided by an embodiment of the present application;

[0054] Figure 3 A flowchart of a difficult example mining method provided by an embodiment of the present application;

[0055] Figure 4 A flowchart of an update method of a speech recognition model provided by an embodiment of the present application;

[0056] Figure 5A structural schematic diagram of a speech recognition system provided by an embodiment of the present application is shown in FIG. 1.

[0057] Figure 6 A structural schematic diagram of a speech recognition device provided by an embodiment of the present application is shown in FIG. 2. DETAILED DESCRIPTION

[0058] As described above, the inventors have found in the research on speech recognition of hot words that the existing speech recognition model usually implements recognition by using speech data to stimulate hot words, that is, if the decoding result of the speech data matches the pre-set hot word, the score of the hot word is stimulated to increase the output probability of the hot word, so that the hot word appears in the optimal output path. However, the process of hot word stimulation needs to stimulate each single word or sub-word in the hot word according to the score of the single word or sub-word, in which case, if the stimulation of a certain single word or sub-word fails, or the decoding is stopped after the first single word or sub-word is decoded, resulting in interruption of the stimulation, the problem of whole-word stimulation failure will occur. Therefore, once the above-mentioned way of stimulating hot words according to the scores of single words or sub-words in the hot word one by one has a problem, it will lead to insufficient accuracy of speech recognition of hot words, affecting the speech recognition of hot words.

[0059] To solve the above-mentioned problem, an embodiment of the present application provides a speech recognition method, which comprises: after obtaining speech data and a hot word library, the acoustic features of the speech data can be determined according to the speech data, and the whole-word score of the hot word is determined based on the hot word saved in the hot word library and the acoustic features, and then the whole-word score of the hot word is used to stimulate the hot word of the speech data.

[0060] As can be seen, by determining the whole-word score of the hot word in the hot word library, the hot word stimulation can be directly performed according to the whole-word score when the hot word stimulation is performed, so that the problem of whole-word stimulation failure caused by stimulating each single word or sub-word can be avoided, thereby improving the accuracy of speech recognition of hot words.

[0061] It should be noted that the embodiment of the present application does not limit the subject performing the speech recognition method, for example, the speech recognition method of the embodiment of the present application can be applied to a terminal device or a server and the like data processing device. The terminal device can be a smart phone, a computer, a smart dictionary, a recording pen, a vehicle-mounted device, a tablet computer, a smart home device, etc. The server can be a stand-alone server, a cluster server or a cloud server.

[0062] In order to make the purposes, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the protection scope of the present application.

[0063] Figure 1 A flowchart of a speech recognition method provided by an embodiment of the present application is shown in FIG. 1. As shown in FIG. 1, the speech recognition method provided by the embodiment of the present application can include the following steps. Figure 1

[0064] S101: Obtain speech data and a hotword library.

[0065] The speech data refers to speech data in a general speech data set. Here, the general speech data set is one or more of, for example, a VoxForg, a CHIME, a TED-LIUM, and the like. The obtaining manner of the general speech data set is not specifically limited in the embodiments of the present application. For example, the general speech data set can be stored in a data processing device for speech recognition, and the data processing device obtains the speech data by a local reading manner when speech recognition is needed. Alternatively, the general speech data set can be stored in another data storage device, and the data processing device can obtain the speech data in the general speech data set by accessing the data storage device when needed.

[0066] The hotword library refers to a database in which hotwords are stored. The obtaining manner of the hotword library is not specifically limited in the embodiments of the present application. For the convenience of understanding, a possible implementation manner will be described below.

[0067] ​In a possible implementation, a hotword can be randomly extracted from the annotated text of the voice data, and a hotword library is constructed. Specifically, the implementation process of randomly extracting a hotword and constructing a hotword library can include: randomly generating a starting extraction position s of the hotword from the annotated text of the voice data; randomly generating a hotword word number n based on the starting extraction position s and the word number c of the annotated text of the voice data; extracting the hotword from the annotated text of the voice data according to the starting extraction position s and the hotword word number n, and constructing a hotword library. The starting extraction position s refers to the starting position of extracting the hotword after the first several words of the annotated text of the voice data, and the value range of the starting extraction position s is 0≤s<n; the value range of the hotword word number n is 1≤n≤min(c,1-s). For example, the annotated text of the voice data is, for example, "I call Xiaoming", and the word number c of the annotated text of the voice data is 4. If the randomly generated starting extraction position s of the hotword is 2 and the randomly generated hotword word number n is 2, it means that the hotword is extracted from the second word "call" in the annotated text of the voice data "I call Xiaoming", and the word number of the extracted hotword is 2, so the hotword is "Xiaoming", and then the hotword "Xiaoming" can be added to the hotword library.

[0068] In addition, the representation form of the hotword library is not specifically limited in the embodiments of the present application, and for the convenience of understanding, the following is described in the form of Table 1.

[0069] Table 1

[0070] Hot word Identification information of the hot word Xiao Ming v1 Xiao Liu V1+1 …… …… Li Si V1+v2-2 Li Hua v1+v2-1

[0071] In combination with Table 1, the hot word library can also embody the identification information of the hot word. In actual application, the speech recognition model can perform general speech recognition through the basic word library, which can be composed of single characters and / or sub-words as elements, for example, the elements in the basic word library can be embodied as "ah", "zui", …, "look", "ed", "ing", etc. It can be understood that if the size of the basic word library is set as v1, the identification information of the first element in the basic word library, i.e., the single character or sub-word, can be set as 0, and the identification information of the last element can be set as V1-1. Correspondingly, the size of the hot word library can be set as v2, and accordingly, the identification information of the first element in the hot word library, i.e., the hot word, can be v1, and the identification information of the last element can be v1+v2-1. In this way, based on the size of the basic word library and the size of the hot word library, the identification information of the hot word is set, which can distinguish the elements in the basic word library and the hot word library, thereby facilitating the distinction of the elements in the new word library after the new word library is constructed based on the basic word library and the hot word library. That is, in the embodiments of the present application, the speech recognition of the hot word can be performed based on the construction of the new word library based on the basic word library and the hot word library, so as to expand the word library for the speech recognition of the hot word. The construction method of the new word library in the embodiments of the present application can not be specifically limited, and in order to facilitate understanding, the present application can provide a possible implementation manner, and the technical details will be introduced below.

[0072] S102: determining the acoustic feature of the speech data according to the speech data.

[0073] Here, the acoustic feature of the speech data, for example, is the Filter Bank feature, the Mel frequency cepstrum coefficient feature, the perceptual linear prediction coefficient feature, etc., which can not be specifically limited in the embodiments of the present application.

[0074] S103: determining the whole-word score of the hot word based on the hot word in the hot word library and the acoustic feature.

[0075] The whole-word score of the hot word can represent the score of the whole hot word. For example, taking "xiaoming" as an example, the whole-word score can represent the score of the whole hot word "xiaoming", rather than the scores of the two single characters "xia" and "ming". Here, the determination process of the whole-word score of the hot word can not be specifically limited in the embodiments of the present application, and in order to facilitate understanding, the following will be described in combination with a possible implementation manner.

[0076] In a possible implementation, S103 can specifically include: determining a word vector of the hotword according to the hotword information corresponding to the hotword; obtaining a first acoustic semantic vector corresponding to the acoustic feature through the attention mechanism; and calculating a similarity between the word vector of the hotword and the first acoustic semantic vector as a whole-word score of the hotword. Since the first acoustic semantic vector corresponding to the acoustic feature only contains information corresponding to single characters or subwords, directly determining the whole-word score of the hotword based on the first acoustic semantic vector cannot effectively improve the accuracy of the speech recognition model, and therefore, in the embodiments of the present application, the whole-word score of the hotword is determined based on the word vector of the hotword and the first acoustic semantic vector, which can improve the accuracy of speech recognition of the hotword. In addition, the first acoustic semantic vector can be used for general speech recognition by means of a deep learning network, for example, a long short-term memory artificial neural network based on the attention mechanism, thereby avoiding the problem of degradation of the effect of general speech recognition caused by training of the speech recognition of the hotword.

[0077] In the embodiments of the present application, the hotword information can be embodied as subword information and phoneme information of the hotword, and correspondingly, the determination of the word vector of the hotword can specifically include: encoding the subwords of the hotword by using a subword encoder module to obtain a subword encoding result, and encoding the phonemes of the hotword by using a phoneme encoder module to obtain a phoneme encoding result; and performing concatenation processing and dimensionality reduction processing on the subword encoding result and the phoneme encoding result to obtain a hotword vector of the hotword.

[0078] In addition, in the embodiments of the present application, before the whole-word score of the hotword is determined, the acoustic feature of the speech data can be encoded by using an audio encoder module to obtain an acoustic feature encoding result. In this way, the acoustic feature encoding result can be subjected to attention operation by the attention mechanism to obtain the first acoustic semantic vector corresponding to the acoustic feature.

[0079] In addition, the process of calculating the similarity between the word vector of the hotword and the first acoustic semantic vector can specifically include: performing attention operation on the word vector of the hotword and the first acoustic semantic vector by using the attention mechanism; and calculating the similarity between the word vector of the hotword and the first acoustic semantic vector based on the operation result. In the embodiments of the present application, the calculation manner of the similarity between the word vector of the hotword and the first acoustic semantic vector is not limited. For example, the cosine similarity between the two can be calculated, or the similarity between the two can be obtained by calculating the Euclidean distance between the two, or the similarity between the two can be obtained by calculating the Manhattan distance between the two. In the embodiments of the present application, since the word vector of the hotword and the first acoustic semantic vector are first subjected to attention operation, and the cosine similarity between the two is a form of multiplicative attention operation, the preferred calculation manner is to calculate the cosine similarity between the two.

[0080] Here, the inventor creatively finds that in the existing speech recognition scheme for hot words, the process of hot word excitation all needs to be performed one by one according to the scores of single words or subwords in the hot words. In the embodiment of the present application, by determining the whole word score of the hot word, the hot word excitation can be directly performed according to the whole word score when the hot word excitation is performed, thereby avoiding the problem of whole word excitation failure caused by performing excitation one by one according to the scores of single words or subwords, so as to improve the accuracy of hot word speech recognition.

[0081] S104: Perform hot word excitation on the speech data by using the whole word score of the hot word.

[0082] In addition, in the process of hot word speech recognition, the decoding result of the speech data at the current time needs to be calculated with all the hot words in the hot word library respectively, so that when there are many hot words in the hot word library, the calculation amount will greatly increase. Based on this, in the embodiment of the present application, in order to reduce the calculation amount of hot word speech recognition, the size of the hot word library can be reduced by hot word screening. Specifically, the process of hot word screening can include: obtaining the decoded result of the speech data before the current time; obtaining the second acoustic semantic vector corresponding to the decoded result by using the attention mechanism; analyzing the second acoustic semantic vector to obtain the first syllable information of the second acoustic semantic vector; based on the first syllable information, deleting the hot words in the hot word library that do not match the first syllable information, and performing speech data recognition on the hot word library after deleting the hot words that do not match the first syllable information. Here, the process of analyzing the second acoustic semantic vector specifically includes: using a syllable classification module to classify the second acoustic semantic vector into syllables to obtain the first syllable information of the second acoustic semantic vector; wherein when the second acoustic semantic vector is a vector corresponding to Chinese speech data, the first syllable information is pinyin information; when the second acoustic semantic vector is a vector corresponding to English speech data, the first syllable information is word information. For example, the first syllable information of the second acoustic semantic vector is "xiao", and the hot words in the hot word library are "Xiaoming" and "Lihua" for example, correspondingly, the first syllable information of "Xiaoming" is "xiao", and the first syllable information of "Lihua" is "li", so "Lihua" can be deleted from the hot word library.

[0083] Based on the above related content of S101-S104, in the embodiment of the present application, after obtaining the speech data and the hot word library, the acoustic features of the speech data can be determined according to the speech data, and the whole word score of the hot word can be determined based on the hot words and acoustic features saved in the hot word library, and then the hot word excitation is performed on the speech data by using the whole word score of the hot word. It can be seen that by determining the whole word score of the hot word in the hot word library, the hot word excitation can be directly performed according to the whole word score when the hot word excitation is performed, so that the problem of whole word excitation failure caused by performing excitation one by one according to the scores of single words or subwords can be avoided, thereby improving the accuracy of hot word speech recognition.

[0084] Since the above methods are implemented through a speech recognition model, in order to expand the vocabulary of the speech recognition model, an embodiment of the present application provides a possible implementation manner for constructing a new vocabulary, which can specifically include S201-S203. S201-S203 are described below in combination with embodiments and drawings.

[0085] Figure 2a A flowchart of an implementation manner for constructing a new vocabulary provided by an embodiment of the present application is shown in FIG. 2. As shown in FIG. 2, S201-S203 can specifically include: Figure 2a

[0086] S201: Obtain a basic vocabulary of a speech recognition model.

[0087] The basic vocabulary refers to a vocabulary used for general speech recognition. The basic vocabulary can include single words and / or subwords. The obtaining manner of the basic vocabulary is not specifically limited in the embodiment of the present application. For example, the basic vocabulary can be saved in a data processing device used for speech recognition, and the data processing device obtains the basic vocabulary through local reading when needed. Alternatively, the basic vocabulary can be saved in another data storage device, and the data processing device can obtain the basic vocabulary by accessing the data storage device when needed.

[0088] S202: Construct a new vocabulary based on the single words and / or subwords in the basic vocabulary and the hot words in the hot vocabulary.

[0089] The constructing manner of the new vocabulary is not specifically limited in the embodiment of the present application. For ease of understanding, a possible implementation manner is described below.

[0090] In a possible implementation manner, S202 can specifically include: obtaining scores of the single words and / or subwords; performing splicing processing on the scores of the single words and / or subwords and the whole-word scores of the hot words to generate a new vocabulary; and the new vocabulary is constructed with the single words and / or subwords and the hot words as elements. Specifically, the process of performing splicing processing on the scores of the single words and / or subwords and the whole-word scores of the hot words can include: processing the scores of the single words and / or subwords according to a pre-set first score formula, and processing the whole-word scores of the hot words according to a pre-set second score formula; and taking the processed new scores as the scores of the corresponding elements in the new vocabulary.

[0091] In the embodiment of the present application, the first score formula can be expressed as:

[0092]

[0093] ​Among them, s1’ is the score after single-word and / or sub-word processing, s1 is the score of single-word and / or sub-word, p is the probability of the absence of hot words, and k is a preset hot-word coefficient.

[0094] It should be noted that the probability p of the absence of hot words refers to the score of a special hot word "nobias", which can represent the absence of hot words.

[0095] Correspondingly, the second score formula can be expressed as:

[0096]

[0097] Among them, s2’ is the whole-word score after processing the hot words in the hot-word library, s2 is the whole-word score of the hot words in the hot-word library, and k is a preset hot-word coefficient.

[0098] In addition, the larger the preset hot-word coefficient, the larger the whole-word score of the processed hot words, and the better the subsequent excitation effect of the hot words. However, correspondingly, the misrecognition rate of the hot words will increase. Therefore, in the embodiments of the present application, the value of the preset hot-word coefficient is 0.5.

[0099] For ease of understanding, in combination with Figure 2b as shown, the embodiments of the present application can provide a schematic diagram of a new method for constructing a word library. In Figure 2b , the single words and / or sub-words in the basic word library are taken as examples from "啊" to "最", and the hot words in the hot-word library are taken as examples of "nobias", "张三", and "李四". Correspondingly, these single words and / or sub-words and the hot words in the hot-word library can first be respectively processed by Softmax, so as to represent the scores of single words and / or sub-words and the whole-word scores of hot words in the form of probabilities, and then the two word libraries are fused according to the above first score formula and second score formula. Specifically, in Figure 2b , the scores of the single words and / or sub-words in the basic word library after being processed by Softmax can be multiplied by [1-(1-p)×k], and the whole-word scores of the hot words in the hot-word library after being processed by Softmax can be multiplied by k, so as to obtain the scores of the elements in the new word library and realize the construction of the new word library.

[0100] S203: Perform hot-word excitation on the speech data based on the new word library.

[0101] For the method of hot-word excitation, the embodiments of the present application do not specifically limit it. For ease of understanding, the following will be described in combination with a possible implementation manner.

[0102] In a possible implementation, S203 can specifically include: matching the decoding result of the voice data with the new vocabulary; and when the matching degree between the hot word in the new vocabulary and the decoding result is greater than or equal to a preset matching degree, performing hot word excitation on the voice data based on the scores of the elements in the new vocabulary.

[0103] Based on the related content of S201-S203, it can be known that the new vocabulary is constructed based on the basic vocabulary and the hot word vocabulary, the vocabulary of the voice recognition model can be expanded, and therefore, the voice recognition accuracy of the hot word can be improved, and the effect of general voice recognition can be avoided.

[0104] In the construction process of the hot word vocabulary, the hot words used for constructing the hot word vocabulary are mostly randomly extracted from the annotated text of the voice data, and therefore, the similarity between the hot words is not high, the difference is large, and the voice recognition model can easily complete the voice recognition of the hot word. However, this can cause the voice recognition model to have difficulty in accurately recognizing the hot word with high similarity. To solve this problem, the embodiment of the present application can provide a difficult example mining method, which improves the training difficulty of the voice recognition model through difficult example mining, thereby improving the accuracy of the voice recognition model. The difficult example mining method will be described below in combination with the embodiments and the accompanying drawings.

[0105] Figure 3 A flowchart of a difficult example mining method provided by an embodiment of the present application is shown in FIG. 3. Figure 3 As shown in FIG. 3, the difficult example mining method provided by the embodiment of the present application can include the following steps.

[0106] S301: Obtain a preset hot word dictionary for voice recognition.

[0107] The preset hot word dictionary for voice recognition is, for example, a general Chinese hot word dictionary, an English hot word dictionary, or the like.

[0108] S302: Determine the similarity between each two preset hot words in the preset hot word dictionary, and sort the similarity between each two preset hot words in descending order.

[0109] Here, the determination process of the similarity between each two preset hot words in the preset hot word dictionary can specifically include: determining the word vector of each preset hot word in the preset hot word dictionary; and determining the similarity between the word vectors of each two preset hot words. Wherein, the calculation manner of the similarity between the word vectors of each two preset hot words is not limited in the embodiments of the present application. For example, the cosine similarity of the two can be calculated, or the similarity of the two can be obtained by calculating the Euclidean distance of the two, or the similarity of the two can be obtained by calculating the Manhattan distance of the two. In the embodiments of the present application, since the similarity between the word vector of the hot word and the first acoustic semantic vector needs to be calculated, when calculating the similarity between the word vectors of each two preset hot words, the preferred calculation manner can be the same as the calculation manner of the similarity between the word vector of the hot word and the first acoustic semantic vector.

[0110] S303: determining the similar hot word of each preset hot word from the preset hot word dictionary based on the ranking result.

[0111] The determination manner of the similar hot word of each preset hot word is not limited in the embodiments of the present application. In order to facilitate understanding, the following will be described in combination with a possible implementation manner.

[0112] In a possible implementation manner, S303 can specifically include: determining the first target hot word from the preset hot word dictionary; determining the similarity between the first target hot word and each preset hot word except the first target hot word in the preset hot word dictionary based on the ranking result; and determining the similar hot word of the first target hot word according to the pre-set selection rule and based on the similarity between the first target hot word and each preset hot word except the first target hot word in the preset hot word dictionary. Here, the pre-set selection rule is, for example, selecting the top 5 preset hot words with the highest similarity as the similar hot words, or selecting the top 10 preset hot words with the highest similarity as the similar hot words, which is not limited herein. In actual application, taking the first target hot word “interesting” as an example, some preset hot words are selected from the preset hot word dictionary, and the similarity is calculated to obtain the ranking result shown in Table 2. It can be understood that the representation form of the ranking result is not limited in the embodiments of the present application.

[0113] Table 2

[0114] First target hot word Part of the preset hot words selected from the preset hot word dictionary Similarity of the two Interesting Interested 0.9459 Interesting Interestingly 0.9292 Interesting Interests 0.9238 Interesting Interest 0.9235 Interesting Interest rate 0.8601 Interesting Intracity 0.7557 Interesting Intrastate 0.7198 Interesting Intrasystem 0.697

[0115] If the pre-set selection rule is to select the top 5 similar hot words as the similar hot words, the similar hot words of the first target hot word "interesting" are "interested", "interestingly", "interests", "interest" and "interestrate".

[0116] In addition, in the process of determining the similar hot words, not only the first target hot word can be operated, but also any pre-set hot word in the pre-set hot word dictionary can be operated, so as to determine the similar hot words of each pre-set hot word in the pre-set hot word dictionary. In order to facilitate understanding of the determination method of the similar hot words of a specific hot word in the pre-set hot word dictionary, the first target hot word is taken as an example for detailed description in the embodiments of the present application.

[0117] S304: If the hot word library and the pre-set dictionary both exist the hot word corresponding to the voice data, one of the similar hot words of the hot word corresponding to the voice data is added to the hot word library.

[0118] Based on the above related content of S301-S304, by determining the similar hot words of each word in the pre-set dictionary and selecting one of the similar hot words of the hot word corresponding to the voice data to add to the hot word library, the similar hot word interference situation can be effectively simulated, the training difficulty of the voice recognition model is increased, so as to improve the discrimination degree of the voice recognition model for similar hot words, and improve the recognition accuracy of the voice recognition model for hot words.

[0119] At present, when a fixed word library is used, the existing voice recognition model can achieve good recognition effect, but since the hot words have user characteristics and timeliness, the hot word library is a changeable word library rather than a fixed word library. In this case, the recognition accuracy of the voice recognition model will be affected. To solve this problem, the embodiments of the present application can provide an updating method of a voice recognition model, which improves the accuracy of the voice recognition model by updating the voice recognition model. The updating method of the voice recognition model will be described below in combination with embodiments and drawings.

[0120] Figure 4 The flowchart of the updating method of the voice recognition model provided by the embodiments of the present application is shown in FIG. 4. In combination with FIG. 4, the updating method of the voice recognition model provided by the embodiments of the present application can include the following steps. Figure 4

[0121] S401: Determine a target function based on a Softmax loss function.

[0122] The determination method of the target function is not limited in the embodiments of the present application. In order to facilitate understanding, a possible implementation manner will be described below. ​

[0123] In a possible implementation, S401 can specifically include: determining a number of training samples of the speech recognition model; the training samples include hot words in a hot word library; calculating a Softmax loss function according to the number of training samples and whole-word scores of the hot words in the hot word library, and taking the Softmax loss function as an objective function.

[0124] In the embodiment of the application, the hot word in the current decoding result of the speech data can be determined as a second target hot word, and other hot words in the hot word library except the second target hot word can be determined, and correspondingly, the Softmax loss function can be implemented by the following formula:

[0125]

[0126] wherein, is the Softmax loss function, N is the number of training samples, i represents the second target hot word, s i is the whole-word score of the second target hot word, s j is the whole-word score of the other hot words, and C is the number of the other hot words.

[0127] S402: updating the speech recognition model by using the objective function.

[0128] In the embodiment of the application, the smaller the objective function determined by the Softmax loss function is, the better the speech recognition model fits, and the higher the accuracy of the speech recognition model is.

[0129] According to the related content of S401-S402 described above, the speech recognition model can be updated by using the Softmax loss function, so that the speech recognition model can assign a higher whole-word score to the hot word in the current decoding result of the speech data, thereby solving the problem of hot word misrecognition and improving the recognition accuracy of the speech recognition model.

[0130] Based on the speech recognition method provided in the above embodiments, the embodiment of the application further provides a speech recognition system. The speech recognition system will be described below in combination with embodiments and drawings.

[0131] Figure 5 FIG. 1 is a structural schematic diagram of a speech recognition system provided in the embodiment of the application. As shown in FIG. 1, the speech recognition system 500 provided in the embodiment of the application can include: Figure 5

[0132] The decoder 501 can be used to obtain the decoded result of the speech data before the current time. Specifically, the decoder 501 can be constructed by using an LSTM (Long Short-Term Memory) network combined with an attention mechanism.​

[0133] The first attention layer 502 can be connected with the decoder 501. The first attention layer 502 is configured to perform attention operation on the decoded result output by the decoder 501 through an attention mechanism to obtain a second acoustic semantic vector. In this way, general speech recognition can be performed by using the second acoustic semantic information, and hotword screening can also be performed based on the second acoustic semantic vector to reduce the size of the hotword library.

[0134] Correspondingly, the first score layer 503 can be connected with the decoder 501 and the first attention layer 502 respectively. The first score layer 503 can be configured to determine the scores of the characters and / or subwords in the basic word library according to the decoded result output by the decoder 501 and the second acoustic semantic information output by the first attention layer 502. In this way, general speech recognition can be performed by using the basic word library subsequently.

[0135] The audio encoder 504 can be configured to encode the acoustic features of the speech data to obtain an acoustic feature encoding result.

[0136] The second attention layer 505 can be connected with the audio encoder 504 and the decoder 501 respectively. The second attention layer 505 can be configured to perform attention operation on the acoustic feature encoding result output by the audio encoder 504 through an attention mechanism in combination with the decoded result output by the decoder 501 to obtain a first acoustic semantic vector corresponding to the acoustic features.

[0137] The subword encoder 506 can be configured to encode the subwords of the hotword to obtain a subword encoding result. The subword encoder 506 can be constructed based on an LSTM (Long Short-Term Memory).

[0138] The phoneme encoder 507 can be configured to encode the phonemes of the hotword to obtain a phoneme encoding result. The phoneme encoder 507 can also be constructed based on an LSTM.

[0139] The conversion layer 508 can be connected with the subword encoder 506 and the phoneme encoder 507 respectively. The conversion layer 508 can be configured to perform concatenation processing and dimension reduction processing on the subword encoding result output by the subword encoder 506 and the phoneme encoding result output by the phoneme encoder 507 to obtain a hotword vector of the hotword.

[0140] The third attention layer 509 can be connected with the second attention layer 505 and the conversion layer 508 respectively. The third attention layer 509 can be configured to perform attention operation on the hotword vector output by the conversion layer 508 and the first acoustic semantic vector output by the second attention layer 505 through an attention mechanism. In this way, the scores of the whole words of the hotword can be further determined according to the operation result subsequently.

[0141] Correspondingly, the second score layer 510 can be connected with the third attention layer 509. The second score layer 510 can be used to calculate the operation result output by the third attention layer 509 to obtain the whole-word score of the hot word.

[0142] In addition, in the embodiment of the present application, the speech recognition system 500 can further include a fusion layer 511. The fusion layer 511 can be connected with the first score layer 503 and the second score layer 510, and is used to fuse the scores of the single characters and / or sub-words in the basic word library output by the first score layer 503 and the whole-word scores of the hot words in the hot word library output by the second score layer 510 to construct a new word library. In this way, by fusing the scores through the fusion layer 511, that is, constructing a new word library based on the basic word library and the hot word library, the word library of the speech recognition model can be expanded, so that the speech recognition accuracy of the hot word can be improved while avoiding the decline of the effect of general speech recognition.

[0143] Further, in the embodiment of the present application, the speech recognition system 500 described above can be specifically used for a speech recognition model. Correspondingly, the speech recognition system 500 can further include a Softmax function layer 512. The Softmax function layer 512 is connected with the second score layer 510, and can be used to calculate a Softmax loss function as an objective function according to the number of training samples of the speech recognition model and the whole-word scores of the hot words output by the second score layer 510, and update the speech recognition model by using the objective function.

[0144] In the embodiment of the present application, in order to realize hot word screening, the speech recognition system can further include a syllable classification layer 513. The syllable classification layer 513 can be connected with the first attention layer 502 and the second score layer 510 respectively. Specifically, the syllable classification layer 513 is used to analyze the second acoustic semantic vector output by the first attention layer 502 to obtain first syllable information of the second acoustic semantic vector, and delete the hot words in the hot word library that do not match the first syllable information based on the first syllable information, and input the deletion result to the second score layer 510, so that the second score layer 510 determines the whole-word scores of the hot words according to the hot word library after deleting the hot words that do not match the first syllable information, thereby realizing hot word screening to reduce the calculation amount of hot word speech recognition and reduce the size of the hot word library.

[0145] Based on the speech recognition method provided in the above embodiment, the embodiment of the present application further provides a speech recognition device. The speech recognition device will be described below in combination with the embodiments and the drawings.

[0146] Figure 6 FIG. 1 is a structural schematic diagram of a speech recognition device provided in the embodiment of the present application. The speech recognition device will be described in combination with the above embodiments and the drawings. Figure 6As shown, the voice recognition device 600 provided by the embodiments of the present application can include:

[0147] The data acquisition module 601 is configured to acquire voice data and a hotword library, and the hotword library includes hotwords.

[0148] The acoustic feature determination module 602 is configured to determine acoustic features of the voice data according to the voice data.

[0149] The whole-word score determination module 603 is configured to determine whole-word scores of the hotwords based on the hotwords in the hotword library and the acoustic features.

[0150] The hotword excitation module 604 is configured to excite the voice data with the hotword based on the whole-word scores of the hotwords.

[0151] As an implementation form, in order to improve the accuracy of voice recognition of the hotword, the whole-word score determination module 603 can specifically include:

[0152] The word vector determination module is configured to determine a word vector of the hotword according to hotword information corresponding to the hotword.

[0153] The first acoustic semantic vector module is configured to obtain a first acoustic semantic vector corresponding to the acoustic features by using an attention mechanism.

[0154] The similarity calculation module is configured to calculate a similarity between the word vector of the hotword and the first acoustic semantic vector as the whole-word score of the hotword.

[0155] As an implementation form, in order to improve the accuracy of voice recognition of the hotword, the voice recognition device 600 can be implemented by using a voice recognition model. Correspondingly, the voice recognition device 600 can further include:

[0156] The basic library acquisition module is configured to acquire a basic library of the voice recognition model, and the basic library includes single words and / or subwords.

[0157] The new library construction module is configured to construct a new library based on the single words and / or subwords in the basic library and the hotwords in the hotword library.

[0158] The first hotword excitation module is configured to excite the voice data with the hotword based on the new library.

[0159] As an implementation form, in order to improve the accuracy of voice recognition of the hotword, the new library construction module can specifically include:

[0160] The score acquisition module is configured to acquire scores of the single words and / or subwords.

[0161] The score splicing module is configured to splice scores of single characters and / or sub-words and whole-word scores of hot words to generate a new vocabulary, and the new vocabulary is constructed by taking single characters and / or sub-words and hot words as elements.

[0162] As an implementation, to improve the accuracy of speech recognition of hot words, the first hot word excitation module can specifically include:

[0163] The matching module is configured to match the decoding result of the speech data with the new vocabulary.

[0164] The second hot word excitation module is configured to, when the matching degree between the hot word in the new vocabulary and the decoding result is greater than or equal to a preset matching degree, excite the speech data based on the scores of the elements in the new vocabulary.

[0165] As an implementation, to improve the accuracy of speech recognition of hot words, the speech recognition device 600 can further include:

[0166] The dictionary acquisition module is configured to acquire a preset hot word dictionary for speech recognition.

[0167] The similarity sorting module is configured to determine the similarity between each two preset hot words in the preset hot word dictionary, and sort the similarity between each two preset hot words in descending order.

[0168] The similar hot word determination module is configured to determine, based on the sorting result, the similar hot word of each preset hot word from the preset hot word dictionary.

[0169] The hot word library updating module is configured to, if the hot word corresponding to the speech data exists in both the hot word library and the preset hot word dictionary, add one of the similar hot words of the hot word corresponding to the speech data to the hot word library.

[0170] As an implementation, to improve the accuracy of speech recognition of hot words, the speech recognition device 600 can further include:

[0171] The decoding result acquisition module is configured to acquire the decoded result of the speech data before the current time.

[0172] The second acoustic semantic vector acquisition module is configured to acquire, by using an attention mechanism, a second acoustic semantic vector corresponding to the decoded result.

[0173] The second acoustic semantic vector analysis module is configured to analyze the second acoustic semantic vector to obtain first syllable information of the second acoustic semantic vector.

[0174] The hot word deletion module is configured to delete, based on the first syllable information, the hot word that does not match the first syllable information from the hot word library, and perform speech recognition based on the hot word library after the hot word that does not match the first syllable information is deleted.

[0175] As an implementation form, in order to improve the accuracy of hot word speech recognition, the speech recognition device 600 can be implemented by using a speech recognition model. Correspondingly, the speech recognition device 600 can further include:

[0176] a first target function determination module configured to determine a target function based on a Softmax loss function;

[0177] a speech recognition model updating module configured to update the speech recognition model by using the target function.

[0178] As an implementation form, in order to improve the accuracy of hot word speech recognition, the first target function determination module can specifically include:

[0179] a training sample number determination module configured to determine the number of training samples of the speech recognition model; the training samples include hot words in a hot word library;

[0180] a second target function determination module configured to calculate the Softmax loss function according to the number of training samples and the whole word score of the hot words in the hot word library, and take the Softmax loss function as the target function.

[0181] Further, the embodiments of the present application also provide a device, which includes a processor, a memory, a system bus, and the processor and the memory are connected through the system bus.

[0182] The processor and the memory are connected through the system bus;

[0183] The memory is configured to store one or more programs, and the one or more programs include instructions, which, when executed by the processor, cause the processor to perform any one of the implementation methods of the speech recognition method.

[0184] Further, the embodiments of the present application also provide a computer readable storage medium, and the computer readable storage medium stores instructions, when the instructions run on a terminal device, cause the terminal device to perform any one of the implementation methods of the speech recognition method.

[0185] Those skilled in the art can clearly understand that all or part of the steps of the above-mentioned method can be implemented by means of software and necessary universal hardware platforms through the description of the above embodiments. Based on such understanding, the technical solutions of the present application can be embodied in the form of a software product in essence or in the form of a part of the prior art that makes contributions to the prior art. The computer software product can be stored in a storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions to make a computer device (which can be a personal computer, a server, or a network communication device such as a media gateway, etc.) execute the method described in each embodiment or some part of the embodiments of the present application.

[0186] It should be noted that the various embodiments described in the specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between embodiments can be referred to each other. For the device disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and the related parts can be referred to the method part.

[0187] It should also be noted that in this document, relationship terms such as first and second are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply that there is any such actual relationship or order between these entities or operations. Moreover, the terms "include", "contain" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or device. Without more limitations, the element defined by the statement "including a" does not exclude the presence of other identical elements in the process, method, article or device including the element.

[0188] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to the embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A voice recognition method, characterized by, The method comprises the following steps: acquiring voice data and a hotword library; the hotword library comprises hotwords; determining acoustic features of the voice data according to the voice data; determining a whole-word score of the hotwords based on the hotwords in the hotword library and the acoustic features; carrying out hotword excitation on the voice data by using the whole-word score of the hotwords; wherein the method further comprises: acquiring a decoded result of the voice data before a current time; obtaining a second acoustic semantic vector corresponding to the decoded result through an attention mechanism; analyzing the second acoustic semantic vector to obtain first syllable information of the second acoustic semantic vector; based on the first syllable information, deleting hotwords that do not match the first syllable information from the hotword library, and carrying out voice recognition based on the hotword library after the deletion of the hotwords that do not match the first syllable information.

2. The method of claim 1, wherein, The method further comprises: determining a word vector of the hotwords according to hotword information corresponding to the hotwords; obtaining a first acoustic semantic vector corresponding to the acoustic features through an attention mechanism; calculating a similarity between the word vector of the hotwords and the first acoustic semantic vector as the whole-word score of the hotwords.

3. The method of claim 1, wherein, The method is implemented by a voice recognition model; the method further comprises: acquiring a basic word library of the voice recognition model; the basic word library comprises single words and / or subwords; constructing a new word library based on the single words and / or subwords in the basic word library and the hotwords in the hotword library; carrying out hotword excitation on the voice data based on the new word library.

4. The method of claim 3, wherein, The method further comprises: acquiring scores of the single words and / or subwords; performing splicing processing on the scores of the single words and / or subwords and the whole-word scores of the hotwords to generate the new word library; the new word library is constructed by taking the single words and / or subwords and the hotwords as elements.

5. The method of claim 3, wherein, The method further comprises: matching a decoding result of the voice data with the new word library; when a matching degree of the hotwords in the new word library and the decoding result is greater than or equal to a preset matching degree, carrying out hotword excitation on the voice data based on scores of elements in the new word library.

6. The method according to any one of claims 1 to 5, characterized in that, The method further comprises: acquiring a preset hotword dictionary for voice recognition; determining similarities between every two preset hotwords in the preset hotword dictionary and sorting the similarities between every two preset hotwords in descending order; based on the sorting result, determining similar hotwords of each preset hotword from the preset hotword dictionary; if the voice data corresponds to a hotword that exists in both the hotword library and the preset hotword dictionary, adding the voice data to the hotword library from the similar hotwords of the hotword corresponding to the voice data.

7. The method according to any one of claims 1 to 5, characterized in that, The method is implemented by a voice recognition model; the method further comprises: determining a target function based on a Softmax loss function; updating the voice recognition model by using the target function.

8. The method of claim 7, wherein, The target function is determined based on the Softmax loss function, and the target function comprises: determining the number of training samples of the speech recognition model; the training samples comprise hot words in the hot word library; According to the number of training samples and the whole word score of the hot words in the hot word library, the Softmax loss function is calculated, and the Softmax loss function is taken as the target function.

9. A speech recognition apparatus, characterized by comprising: Comprise: Data acquisition module, for acquiring voice data and hot word library; the hot word library comprises hot words; Acoustic feature determination module, for determining the acoustic feature of the voice data according to the voice data; Whole word score determination module, for determining the whole word score of the hot words based on the hot words in the hot word library and the acoustic feature; Hot word excitation module, for exciting the voice data with the whole word score of the hot words; Wherein, the speech recognition device further comprises: Decoding result acquisition module, for acquiring the decoded result of the voice data before the current time; Second acoustic semantic vector acquisition module, for acquiring the second acoustic semantic vector corresponding to the decoded result through attention mechanism; Second acoustic semantic vector analysis module, for analyzing the second acoustic semantic vector to obtain the first syllable information of the second acoustic semantic vector; Hot word deletion module, for deleting the hot words in the hot word library which do not match the first syllable information based on the first syllable information, and performing speech recognition on the hot word library after deleting the hot words which do not match the first syllable information.

10. A speech recognition device, characterized by The device comprises a processor, a memory and a system bus; The processor and the memory are connected through the system bus; The memory is used to store one or more programs, the one or more programs include instructions, the instructions are executed by the processor to make the processor execute the method of any one of claims 1 to 8.

11. A computer readable storage medium, characterized in that, The computer readable storage medium stores instructions, when the instructions run on the terminal device, make the terminal device execute the method of any one of claims 1 to 8.

Citation Information

Patent Citations

  • Lexicon updating method and device

    CN107180084A

  • Hot word recognition method, device and system

    CN110879839A

  • Speech recognition method, device and apparatus and storage medium

    CN111583909A