A voice wake-up method and apparatus, an electronic device, and a storage medium
By penalizing the wake word after the voice wake-up operation, the recognition competitiveness of the wake word is reduced, the problem of multiple wake word crosstalk is solved, the accuracy of voice wake-up is improved, and the recognition accuracy of multiple wake words is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- IFLYTEK CO LTD
- Filing Date
- 2022-10-24
- Publication Date
- 2026-07-24
Smart Images

Figure CN115881110B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of speech recognition, and in particular to a voice wake-up method, device, electronic device, and storage medium. Background Technology
[0002] Voice, as one of the most commonly used communication methods, has been a key focus of research since 1960 as a natural human-computer interaction method, and has made significant progress in the past decade. Simultaneously, with the rapid development of hardware performance and the increasing maturity of artificial intelligence technologies, various smart devices have emerged and entered the lives of ordinary users. Voice interaction, as a hands-free interaction method, is highly favored by users and has become one of the most frequently used interaction methods, such as voice wake-up and voice recognition. This has also driven the development of industries such as smartphones and smart homes, demonstrating huge demand and market potential. As the entry point for voice recognition, voice wake-up also places higher demands on the effectiveness of wake-up and preventing false wake-ups.
[0003] Currently, with the increasing market demand, it is becoming more and more common for a product to have several or even dozens of wake words, which brings greater challenges to the crosstalk between wake words, and this is also a problem in the industry. Summary of the Invention
[0004] This application provides at least one voice wake-up method, device, electronic device, and storage medium, which can effectively reduce crosstalk between multiple wake-up words.
[0005] The first aspect of this application provides a voice wake-up method, comprising: acquiring first voice data; performing voice recognition on the first voice data to obtain a first wake-up word represented by the first voice data; performing a first wake-up operation according to the first wake-up word; and performing a preset penalty operation on the first wake-up word within a preset time after the first wake-up operation, wherein the preset penalty operation is used to reduce the probability of recognizing second voice data acquired within the preset time as the first wake-up word.
[0006] A second aspect of this application provides a voice wake-up device, comprising: an acquisition module for acquiring first voice data; a voice recognition module for performing voice recognition on the first voice data to obtain a first wake-up word represented by the first voice data; a decoding module for performing a first wake-up operation according to the first wake-up word; and a preset penalty operation on the first wake-up word within a preset time after the first wake-up operation, wherein the preset penalty operation is used to reduce the probability of recognizing second voice data acquired within the preset time as the first wake-up word.
[0007] A third aspect of this application provides an electronic device including a memory and a processor coupled to each other, the processor being configured to execute program instructions stored in the memory to implement the voice wake-up method described in the first aspect above.
[0008] The fourth aspect of this application provides a computer-readable storage medium having program instructions stored thereon, which, when executed by a processor, implement the voice wake-up method described in the first aspect above.
[0009] The above-mentioned scheme first performs speech recognition on the acquired first speech data to obtain the first wake-up word represented by it, and then performs a first wake-up operation on the first wake-up word. If the recognition of the first wake-up word is incorrect, that is, crosstalk occurs between wake-up words, the user may re-enter the speech data to correct the first wake-up operation. Therefore, the scheme of this application performs a preset penalty operation on the first wake-up word within a preset time after the first wake-up operation. The preset penalty operation is used to reduce the probability of recognizing the second speech data acquired within the preset time as the first wake-up word, thereby reducing the recognition competitiveness of the first wake-up word in the speech recognition process within the preset time after the first wake-up operation, reducing the possibility of the first wake-up word crosstalking to other wake-up words within the preset time, indirectly improving the recognition competitiveness of other wake-up words, so that the correct wake-up word can be accurately recognized. Therefore, this application can effectively reduce the crosstalk rate of wake-up words, thereby improving the accuracy of voice wake-up.
[0010] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this application. Attached Figure Description
[0011] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the specification, serve to explain the technical solutions of this application.
[0012] Figure 1 This is a flowchart illustrating an embodiment of the voice wake-up method provided in this application;
[0013] Figure 2 This is a flowchart illustrating step S120 in another embodiment of the voice wake-up method of this application;
[0014] Figure 3 This is a flowchart illustrating step S121 in another embodiment of the voice wake-up method of this application;
[0015] Figure 4 This is a schematic diagram of the punishment process in one embodiment of the voice wake-up method provided in this application;
[0016] Figure 5 This is a schematic diagram of the framework of an embodiment of the voice wake-up system provided in this application;
[0017] Figure 6 This is a flowchart illustrating another embodiment of the voice wake-up method provided in this application;
[0018] Figure 7 This is a schematic diagram of the framework of an embodiment of the voice wake-up device provided in this application;
[0019] Figure 8 This is a schematic diagram of the framework of an embodiment of the electronic device provided in this application;
[0020] Figure 9 This is a schematic diagram of a framework of an embodiment of the computer-readable storage medium provided in this application. Detailed Implementation
[0021] The embodiments of this application will now be described in detail with reference to the accompanying drawings.
[0022] In the following description, specific details such as particular system architectures, interfaces, and technologies are presented for illustrative purposes rather than for limiting purposes, in order to provide a thorough understanding of this application.
[0023] In this document, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " generally indicates that the preceding and following related objects have an "or" relationship. Furthermore, "many" in this document means two or more. Moreover, the term "at least one" in this document means any combination of at least two of any one or more of a plurality of objects. For example, including at least one of A, B, and C can mean including any one or more elements selected from the set consisting of A, B, and C.
[0024] Please see Figure 1 , Figure 1 This is a flowchart illustrating an embodiment of the voice wake-up method provided in this application. This method can be executed by any electronic device capable of data processing, such as a computer, mobile phone, or home appliance. Specifically, the voice wake-up method of this application includes the following steps:
[0025] Step S110: Obtain the first voice data.
[0026] The first voice data in this document can be voice information spoken by a user, or voice data recorded or stored in real time by electronic devices such as speakers, computers, and mobile phones. This first voice data can be acquired by a device containing microphone or other sound pickup hardware. It is understood that the device acquiring this first voice data can be an electronic device executing this method, or other devices capable of communicating with that electronic device.
[0027] Step S120: Perform speech recognition on the first speech data to obtain the first wake-up word represented by the first speech data.
[0028] In some embodiments, a speech recognition model can be used to perform speech recognition on the first speech data to obtain the first wake-up word. The speech recognition model can be a neural network model. The training process of this neural network model is as follows: acquiring audio data labeled with the wake-up word as training data, or acquiring the speech features of the audio data labeled with the wake-up word as training data, continuously feeding these training data into the neural network model, and updating the model's weights through backpropagation to reduce the error between its prediction result and the target result, thereby enabling the model to achieve a sufficient fit. The resulting neural network model is then a well-trained neural network model. It is understood that this speech recognition model is a model with deep learning capabilities, and no specific limitation is made here. Of course, in other embodiments, other speech recognition algorithms can also be used to achieve speech recognition, and no limitation is made here.
[0029] In another embodiment, reference may be made to Figure 2 Achieve speech recognition, Figure 2 This is a flowchart illustrating step S120 in another embodiment of the voice wake-up method of this application. Specifically, step S120 includes the following sub-steps:
[0030] Step S121: Obtain the acoustic state posterior probability of the first speech data.
[0031] Speech data typically contains syllable information for each word segment. The posterior probability of the acoustic state is used to characterize the probability that the pronunciation contained in the speech data belongs to various acoustic states. Specifically, it can be obtained by performing acoustic state prediction on the speech data.
[0032] For example, in conjunction with reference Figure 3 , Figure 3 This is a partial flowchart of step S121 in another embodiment of the voice wake-up method of this application. Step S121 specifically includes the following steps S1211 and S1212.
[0033] Step S1211: Extract the spectral features of the first speech data.
[0034] Step S1212: Use a speech recognition model to predict the acoustic state of the spectral features, and obtain the posterior probability of the acoustic state of the first speech data.
[0035] Speech data is generally a time-series signal of variable length and is not suitable as the input of deep learning. Usually, the original speech signal waveform needs to be converted into a specific feature vector representation. In this embodiment, spectral features of the first speech data are extracted. For example, features can be extracted from the first speech data according to Mel Frequency Cepstrum Coefficient (MFCC), filter banks, etc., and corresponding spectral features such as MFCC features are obtained. No specific limitation is made here.
[0036] In some embodiments, a speech recognition model is used to perform acoustic state prediction based on the pronunciation information of each segmented word included in the spectral features, and obtain the posterior probability of the acoustic state of the first speech data. Among them, the posterior probability of the acoustic state is used to represent the posterior probability corresponding to the spectral features. That is, it can be understood that the posterior probability of the acoustic state is used to represent the importance of the spectral features for speech recognition. The greater the posterior probability of the acoustic state, the more important the spectral features are for speech recognition, that is, the more accurate the recognition result obtained based on the spectral features for speech recognition.
[0037] Specifically, the posterior probability of the acoustic state can be a probability matrix, and each element in this probability matrix is the pronunciation probability of each segmented word that may be recognized in the spectral features. For example, the input first speech data is "twenty-five degrees", which includes segmented words "two, ten, five, degrees". Its spectral features are extracted, and the acoustic state of its spectral features is predicted. The speech recognition model predicts each segmented word respectively, and obtains the pronunciation state probability of each segmented word. For example, for the segmented word "two", the probabilities of each pronunciation state such as "er1", "er2", "er4", etc. are predicted. For the segmented word "ten", the probabilities of each pronunciation state such as "sh i2", "s i2", etc. are predicted, and so on. The pronunciation state probabilities of each segmented word are used as matrix elements respectively to form the posterior probability of the acoustic state of this first speech data. It can be understood that speech data usually contains various data information such as semantics, grammar, phonemes, spectra, etc. Different data information can be selected for feature extraction according to the usage situation of the speech data. No specific limitation is made here.
[0038] Step S122: Decode based on the posterior probability of the acoustic state of the first speech data to obtain the first wake-up word.
[0039] In some embodiments, the Viterbi algorithm can be used to decode the acoustic state posterior probability of the first speech data. The Viterbi algorithm is a dynamic programming algorithm used to find the most likely pronunciation sequence in speech recognition. Specifically, the Viterbi algorithm is used to decode the acoustic state posterior probability of the first speech data to determine the state transition probabilities between adjacent pronunciation states in several second candidate pronunciation sequences corresponding to the first speech data. Based on the state transition probabilities between adjacent pronunciation states in the several second candidate pronunciation sequences, a second optimal candidate pronunciation sequence is selected, and a first wake-up word is obtained based on the second optimal candidate pronunciation sequence.
[0040] In some specific embodiments, when the input first speech data is "twenty-five degrees", after obtaining the acoustic state posterior probability of the first speech data, the pronunciation states predicted for each word in the acoustic state posterior probability are used to form different candidate pronunciation sequences. For example, the acoustic state posterior probability includes: the pronunciation state predicted for the word "two" is "er4", the pronunciation states predicted for the word "ten" are "sh i2" and "si2", the pronunciation states predicted for the word "five" are "wu3" and "liu4", and the pronunciation state predicted for the word "degree" is "du4". Then, the second candidate pronunciation sequence can be decoded to include "er4 shi2 liu4 du4", "er4 si2 liu4 du4", "er4 sh i2 wu3 du4", and "er4 si2 wu3". In the sequence “du4”, the state transition probability between adjacent pronunciation states in each second candidate pronunciation sequence can be obtained based on the probability of adjacent pronunciation states in the acoustic state posterior probability. In the acoustic state posterior probability, the probability of “er4” is 0.8 and the probability of “sh i2” is 0.9. Therefore, the state transition probability of the adjacent pronunciation states “er4 sh i2” can be 0.8 * 0.9, which is 0.72. Based on the state transition probabilities between adjacent pronunciation states in each of the several second candidate pronunciation sequences, the optimal candidate pronunciation sequence is selected from the several second candidate pronunciation sequences in the acoustic state posterior probability. The several second candidate pronunciation sequences can be equivalent to several second candidate paths. The optimal path can be selected from these several second candidate paths to serve as the second optimal candidate pronunciation sequence. Specifically, the Viterbi algorithm can be used to select the optimal path from several second candidate paths. Each pronunciation state in the second candidate pronunciation sequence is a path point on the corresponding second candidate path, and the state transition probability between adjacent pronunciation states is the path score of the corresponding adjacent path point. The scores of several second candidate paths are calculated using the Viterbi algorithm, and the optimal path is selected based on the scores of each second candidate path (e.g., selecting the second candidate path with the highest score as the optimal path). The candidate pronunciation sequence corresponding to this optimal path is the second optimal candidate pronunciation sequence. The text corresponding to this second optimal candidate pronunciation sequence is then obtained as the first wake word.
[0041] It is understood that this embodiment is not limited to using the Viterbi algorithm to obtain the first wake word based on the acoustic state posterior probability. Other methods of obtaining the wake word using the acoustic state posterior probability can also be used. Of course, it is not limited to obtaining the wake word using the acoustic state posterior probability. This application can use any speech recognition method to obtain the wake word, so no specific limitation is made here.
[0042] Step S130: Perform the first wake-up operation according to the first wake-up word.
[0043] In this embodiment, the electronic device typically pre-stores wake-up operation instructions corresponding to different wake-up words. After determining the first wake-up word, it can search the pre-stored information to obtain the wake-up operation instruction corresponding to the first wake-up word, and then execute the wake-up operation instruction to realize the first wake-up operation. For example, if the electronic device is an air conditioner and the first wake-up word is "26 degrees", then the corresponding first wake-up operation is to adjust the current temperature to 26 degrees, or to send an instruction to its associated device, such as an air conditioner in another room, to adjust the current temperature to 26 degrees.
[0044] In one specific embodiment, after obtaining the first wake-up word, the recognition accuracy of the first wake-up word can be determined first. If the recognition accuracy is relatively high, then step S130 can be executed. For example, the score of the second optimal candidate pronunciation sequence can be obtained, which can represent the accuracy of the second optimal candidate pronunciation sequence, specifically as the score of the optimal path calculated using the Viterbi algorithm as described above. In response to the score of the second optimal candidate pronunciation sequence exceeding the threshold score of the second optimal candidate pronunciation sequence, the first wake-up operation is performed according to the first wake-up word.
[0045] And step S140: Perform a preset penalty operation on the first wake-up word within a preset time after the first wake-up operation.
[0046] The preset penalty operation is used to reduce the probability of recognizing the second voice data acquired within a preset time as the first wake word.
[0047] During speech recognition, crosstalk between wake words may occur. This means that the speech data to be recognized is one wake word, but the speech is recognized as a different wake word. In other words, when the first speech data input by the user results in a wake word that is not the actual wake word entered by the user, the user will usually choose to repeat the input within a short period. Therefore, when a user repeatedly uses the same wake word within a short period, there is a high probability that crosstalk between multiple wake words has occurred. Therefore, a low-crosstalk scheme can be implemented for the first wake word. This involves reducing the recognition competitiveness of the first wake word for a preset period after the first speech data is used to obtain the first wake word, thus avoiding crosstalk from the first wake word again. Specifically, a penalty coefficient can be set for the first wake word to reduce the probability of recognizing the second speech data acquired within the preset period as the first wake word, thereby reducing the recognition competitiveness of the first wake word.
[0048] In this embodiment, the first voice data is first processed by speech recognition to obtain the first wake-up word represented by it, and a first wake-up operation is performed on the first wake-up word. If the first wake-up word is misidentified, i.e., crosstalk occurs between wake-up words, the user may repeatedly input voice data to correct the first wake-up operation. Therefore, the present application performs a preset penalty operation on the first wake-up word within a preset time after the first wake-up operation. The preset penalty operation is used to reduce the probability of recognizing the second voice data acquired within the preset time as the first wake-up word, thereby reducing the recognition competitiveness of the first wake-up word in the speech recognition process within the preset time after the first wake-up operation, reducing the possibility of the first wake-up word crosstalking to other wake-up words within the preset time, and indirectly improving the recognition competitiveness of other wake-up words, so that the correct wake-up word can be accurately identified. Therefore, the present application can effectively reduce the crosstalk rate of wake-up words, thereby improving the accuracy of voice wake-up.
[0049] In one embodiment, the implementation of a low crosstalk scheme can be flexibly set for different wake words. Therefore, after step S130 and before step 140, it can be determined whether the first wake word applies the low crosstalk scheme. If the first wake word applies the low crosstalk scheme, step 140 is executed to apply a preset penalty to the first wake word. Furthermore, the implementation of the low crosstalk scheme can also be determined by combining the recognition confidence of the first wake word. For example, if the first wake word applies the low crosstalk scheme, it can be further determined whether the speech recognition confidence of the first wake word exceeds a confidence threshold, wherein the confidence threshold can be preset; if the speech recognition confidence of the first wake word does not exceed the confidence threshold, step 140 is executed.
[0050] In some embodiments, the user can determine which wake words require a low-crosstalk scheme based on their needs, or the device can determine which wake words require a low-crosstalk scheme based on the crosstalk probability of different wake words. Wake words determined to require a low-crosstalk scheme can be labeled for differentiation; for example, labeled wake words use a low-crosstalk scheme, while unlabeled wake words do not. During decoding, if a labeled wake word is obtained, its confidence level is assessed. This confidence level can be the accuracy of obtaining the optimal path for that wake word, such as the optimal path score.
[0051] In other embodiments, step S140 above may specifically include the following steps:
[0052] Step S141: In response to acquiring the second speech data within a preset time, the second speech data is used to perform speech recognition using the penalty coefficient of the first wake-up word to obtain the second wake-up word represented by the second speech data.
[0053] The penalty coefficient for the first wake word is used to reduce the probability that the second speech data is identified as the first wake word.
[0054] In some embodiments, the penalty coefficient of the first wake-up word varies with time, specifically as a distribution function that varies with time. In a specific embodiment, the penalty coefficient of the first wake-up word decreases as time increases within a preset time period. For example, if the preset time period is 3 minutes, the penalty coefficient of the first wake-up word is 0.9 in the first minute, 0.7 in the second minute, and 0.5 in the third minute; or the penalty coefficient is a linear or non-linear function that decreases as time increases within the 3 minutes. It is understood that the penalty coefficient of the first wake-up word is a valid value within the preset time period. If the preset time is exceeded, the penalty coefficient becomes invalid, for example, it is 0, meaning that the first wake-up word is no longer penalized after the preset time has elapsed. Specifically, during the decoding process of the second speech data, the penalty coefficient can be used to reduce the state transition probability of the pronunciation states related to the first wake-up word in several first candidate pronunciation sequences of the second speech data. This reduces the probability of recognizing the second speech data as the first wake-up word. The pronunciation states related to the first wake-up word can be some or all of the pronunciation states corresponding to the first wake-up word.
[0055] In some specific embodiments, the above-described method of using the penalty coefficient of the first wake-up word to perform speech recognition on the second speech data to obtain the second wake-up word represented by the second speech data may include the following specific steps:
[0056] Step S1411: Obtain the acoustic state posterior probability of the second speech data.
[0057] This step is the same as the method described above for obtaining the posterior probability of the acoustic state of the first speech data, and will not be elaborated further here.
[0058] Step S1412: Decode the second wake word based on the acoustic state posterior probability of the second speech data and the penalty coefficient of the first wake word.
[0059] The penalty coefficient is used to reduce the state transition probability of at least some of the pronunciation states of the second speech data during the decoding process. The decoding of this second speech data can be implemented using the Viterbi algorithm, as detailed in the preceding description.
[0060] In some embodiments, based on the posterior probability of the acoustic state of the second speech data, the state transition probability between adjacent pronunciation states in each of a plurality of first candidate pronunciation sequences corresponding to the second speech data is determined. At least one set of sub-pronunciation sequences is selected from the plurality of first candidate pronunciation sequences, and the state transition probability between adjacent pronunciation states in each set of sub-pronunciation sequences is reduced using a penalty coefficient. Among them, the sub-pronunciation sequence consists of at least two consecutive pronunciation states in the first candidate pronunciation sequence, and the pronunciation states except the first and last ones in the sub-pronunciation sequence are the pronunciation states included in the first wake-up word. Specifically, the pronunciation states except the first and last ones in the sub-pronunciation sequence can be the crosstalk-prone pronunciation states in the first wake-up word, and the crosstalk-prone pronunciation states can be some or all of the pronunciation states included in the first wake-up word. Specifically, the crosstalk-prone pronunciation states in the first wake-up word can be set for the first wake-up word according to prior experience. For example, for the first wake-up word "twenty-six degrees", according to prior experience, its corresponding crosstalk-prone pronunciation state is "l iu4".
[0061] In some specific embodiments, for each group of adjacent pronunciation states in each group of sub-pronunciation sequences, the state transition probability of the adjacent pronunciation states is multiplied by the penalty coefficient corresponding to the adjacent pronunciation states to obtain the current state transition probability between the adjacent pronunciation states. Among them, the penalty coefficient is less than 1, and the penalty coefficients corresponding to different groups of adjacent pronunciation states are the same or different.
[0062] Specifically, for example, the first wake-up word "twenty-six degrees" is obtained from the first speech data and the first wake-up operation is performed. According to prior experience, it is known that the first wake-up word "twenty-six degrees" is prone to crosstalk with the wake-up word "twenty-five degrees", and the difference between the pronunciation sequences of "twenty-six degrees" and "twenty-five degrees" lies in "l iu4" and "w u3". Therefore, the pronunciation state "l iu4" of the first wake-up word "twenty-six degrees" can be used as the crosstalk-prone pronunciation state. Therefore, within a preset time after the first wake-up operation, when the second speech data is input, the "l iu4" and its adjacent pronunciation states before and after in the plurality of first candidate pronunciation sequences corresponding to the second speech data are used as the sub-pronunciation sequences that need to be penalized by the penalty coefficient. For specific details, please refer to Figure 4The second speech data is input at a certain moment within a preset time after the execution of the first wake-up operation, and the penalty coefficient obtained at that moment is 0.8. The first candidate sequences corresponding to the second speech data obtained through speech recognition include "er4 sh i2 liu4 du4" and "er4 sh i2 w u3 du4". According to the previously determined easily crosstalked pronunciation state in the first wake-up word, "liu4", it is checked whether each first candidate sequence contains "liu4". In the first candidate sequence "er4sh i2 liu4 du4" which contains "liu4", the pronunciation states containing "liu4" and its adjacent states are selected to form the sub-pronunciation sequence 410 "i2 liu4 d" corresponding to the first candidate sequence. The penalty coefficient of 0.8 corresponding to the input moment is used to penalize the sub-pronunciation sequence 410. Specifically, the state transition probability between each adjacent pronunciation state in the sub-pronunciation sequence 410 is multiplied by the penalty coefficient of 0.8 to reduce the competitiveness of "26 degrees" and at the same time increase the recall rate of the wake-up word "25 degrees". Regarding the penalty coefficient, the penalty coefficients corresponding to different groups of adjacent pronunciation states in the sub-pronunciation sequence can be different. For example, the penalty coefficient multiplied by the pronunciation state "i2" to the pronunciation state "l" is 0.7, while the penalty coefficients multiplied by the pronunciation state "l" to the pronunciation state "iu4" and the pronunciation state "iu4" to the pronunciation state "d" are both 0.8. It is understood that the penalty coefficient depends on the specific situation and is not specifically limited here.
[0063] Step S142: Perform the second wake-up operation according to the second wake-up word.
[0064] For details, please refer to the relevant description of step S130, which will not be limited here.
[0065] In some embodiments, before performing the second wake-up operation according to the second wake-up word, the method further includes: obtaining a score of the first optimal candidate pronunciation sequence; and, in response to the score of the first optimal candidate pronunciation sequence exceeding a threshold score, performing the second wake-up operation according to the second wake-up word. For details, please refer to the description of the first wake-up word above, which will not be elaborated further here.
[0066] Please see Figure 5 , Figure 5This is a schematic diagram of the framework of the voice wake-up system provided in this application. First voice data is input into the voice system. The voice system uses a speech recognition model to recognize and calculate the first voice data, obtaining the acoustic state posterior probability of the first voice data, and then decodes it based on the acoustic state posterior probability. Specifically, a second optimal candidate pronunciation sequence can be obtained based on the acoustic state posterior probability using the Viterbi algorithm. A first wake-up word is obtained based on the second optimal candidate pronunciation sequence, and a score is calculated for the second optimal candidate pronunciation sequence. Whether to wake up is determined based on the score of the second optimal candidate pronunciation sequence. If the score of the second optimal candidate pronunciation sequence is greater than or equal to a threshold score, the first wake-up word is emitted, i.e., wake-up is executed; if the score of the second optimal candidate pronunciation sequence is less than the threshold score, the first wake-up word is not emitted, i.e., wake-up is not executed. After emitting the first wake-up word, it is determined whether the first wake-up word is a pre-labeled wake-up word that requires a low crosstalk scheme. If the first wake-up word is a pre-labeled wake-up word that requires a low crosstalk scheme, then the confidence level of this wake-up is first determined. If the confidence level of the first wake-up word does not exceed its confidence threshold, it is considered that crosstalk may occur in this wake-up, and a decoding penalty is applied to subsequent wake-ups.
[0067] Within a preset time after the first wake-up word is executed, second speech data is input. A speech recognition model is used to recognize and calculate the second speech data, obtaining its acoustic state posterior probability. Decoding is then performed based on this probability. During the decoding process, the pronunciation sequence of the first wake-up word is penalized by multiplying the pronunciation states associated with the first wake-up word by a penalty coefficient. This reduces the competitiveness of the first wake-up word, thereby improving recall and making crosstalk less likely to occur during this wake-up. Subsequent steps are the same as those for the first speech data and will not be elaborated further.
[0068] In this embodiment, typical user behavior occurring under crosstalk conditions is used as a priori and applied to the decoding process to preemptively penalize the user, thereby recalling the correct wake-up word and reducing the crosstalk rate. Simultaneously, due to the confidence level determination and the fact that the penalty is zero after a certain time, normal wake-up is guaranteed even with continuous wake-ups. This achieves a very limited impact on the overall effect while reducing the crosstalk rate and improving the recall rate.
[0069] Please see Figure 6 , Figure 6 This is a flowchart illustrating another embodiment of the voice wake-up method provided in this application. The specific steps are as follows:
[0070] Step S210: Obtain the first voice data.
[0071] This step can be referred to in the description of the relevant steps in the above embodiments, and will not be repeated here.
[0072] Step S220: Preprocess the first speech data and perform speech recognition to obtain the first wake-up word.
[0073] This step can be referred to in the description of the relevant steps in the above embodiments, and will not be repeated here.
[0074] Step S230: Calculate the score for the first wake-up word, and determine whether the first wake-up word adopts a low crosstalk scheme based on the score.
[0075] In some embodiments, the Viterbi algorithm is used to calculate the scores of several second candidate pronunciation sequences in the first speech data, and the second candidate pronunciation sequence with the highest score is selected as the second optimal candidate pronunciation sequence. The first wake-up word is obtained through the second optimal candidate pronunciation sequence. The score of the second optimal candidate pronunciation sequence is compared with the threshold score of its corresponding first wake-up word. If the score of the second optimal candidate pronunciation sequence exceeds the threshold score of its corresponding first wake-up word, the first wake-up operation is performed; if the score of the second optimal candidate pronunciation sequence is lower than the threshold score of its corresponding first wake-up word, crosstalk is performed on the first wake-up word, and a confidence judgment is made on the first wake-up word. Subsequent steps can be referred to the relevant descriptions above, and will not be elaborated further here.
[0076] Please see Figure 7 , Figure 7 This is a schematic diagram of a framework of an embodiment of the voice wake-up device provided in this application. The voice wake-up device 700 includes: an acquisition module 710, a speech recognition module 720, and a decoding module 730. The acquisition module 710 is used to acquire first speech data; the speech recognition module 720 is used to perform speech recognition on the first speech data to obtain a first wake-up word represented by the first speech data; the decoding module 730 is used to perform a first wake-up operation according to the first wake-up word; and to perform a preset penalty operation on the first wake-up word within a preset time after the first wake-up operation, the preset penalty operation being used to reduce the probability of recognizing second speech data acquired within the preset time as the first wake-up word.
[0077] In some embodiments, the decoding module 730 performs a preset penalty operation on the first wake-up word within a preset time after the first wake-up operation, specifically including: in response to acquiring the second voice data within the preset time, performing voice recognition on the second voice data using the penalty coefficient of the first wake-up word to obtain the second wake-up word represented by the second voice data, wherein the penalty coefficient of the first wake-up word is used to reduce the probability that the second voice data is recognized as the first wake-up word; and performing a second wake-up operation according to the second wake-up word.
[0078] In some embodiments, the penalty coefficient of the first wake word in the decoding module 730 decreases over time.
[0079] In some embodiments, the speech recognition module 720 performs speech recognition on the second speech data using a penalty coefficient of a first wake-up word to obtain a second wake-up word represented by the second speech data. Specifically, this includes: obtaining the acoustic state posterior probability of the second speech data; and decoding based on the acoustic state posterior probability of the second speech data and the penalty coefficient of the first wake-up word to obtain the second wake-up word. The penalty coefficient is used to reduce the state transition probability of at least some of the pronunciation states of the second speech data during the decoding process.
[0080] In some embodiments, the decoding module 730 performs decoding based on the acoustic state posterior probability of the second speech data and the penalty coefficient of the first wake-up word to obtain the second wake-up word. Specifically, this includes: determining the state transition probability between adjacent pronunciation states in a plurality of first candidate pronunciation sequences corresponding to the second speech data based on the acoustic state posterior probability of the second speech data; selecting at least one set of sub-pronunciation sequences from the plurality of first candidate pronunciation sequences, and using the penalty coefficient to reduce the state transition probability between adjacent pronunciation states in each set of sub-pronunciation sequences, wherein the sub-pronunciation sequence consists of at least two consecutive pronunciation states in the first candidate pronunciation sequences, and the pronunciation states in the sub-pronunciation sequence other than the beginning and end are the pronunciation states included in the first wake-up word.
[0081] In some embodiments, the decoding module 730 performs a penalty coefficient to reduce the state transition probability between adjacent pronunciation states in each group of sub-pronunciation sequences. Specifically, for each group of adjacent pronunciation states in each group of sub-pronunciation sequences, the state transition probability of the adjacent pronunciation state is multiplied by the penalty coefficient corresponding to the adjacent pronunciation state to obtain the current state transition probability between adjacent pronunciation states. The penalty coefficient is less than 1, and the penalty coefficients corresponding to different groups of adjacent pronunciation states may be the same or different.
[0082] In some embodiments, before performing the second wake-up operation according to the second wake-up word, the decoding module 730 further includes: obtaining the score of the first optimal candidate pronunciation sequence; and performing the second wake-up operation according to the second wake-up word in response to the score of the first optimal candidate pronunciation sequence exceeding the threshold score of the first optimal candidate pronunciation sequence.
[0083] In some embodiments, before the decoding module 730 performs a preset penalty operation on the first wake-up word within a preset time after the first wake-up operation, the method further includes: in response to the application of a low crosstalk scheme to the first wake-up word, determining whether the speech recognition confidence of the first wake-up word exceeds a confidence threshold; and in response to the speech recognition confidence of the first wake-up word not exceeding the confidence threshold, performing a preset penalty operation on the first wake-up word within a preset time after the first wake-up operation.
[0084] In some embodiments, the speech recognition module 720 performs speech recognition on the first speech data to obtain the first wake-up word represented by the first speech data, including: obtaining the acoustic state posterior probability of the first speech data; and decoding based on the acoustic state posterior probability of the first speech data to obtain the first wake-up word.
[0085] In some embodiments, the speech recognition module 720 performs the acquisition of the acoustic state posterior probability of the first speech data, specifically including: extracting the spectral features of the first speech data; using a speech recognition model to predict the acoustic state of the spectral features, and obtaining the acoustic state posterior probability of the first speech data;
[0086] In some embodiments, the decoding method in the decoding module 730 can be Viterbi decoding, which decodes based on the acoustic state posterior probability of the first speech data to obtain the first wake-up word, including: determining the state transition probability between each adjacent pronunciation state in a plurality of second candidate pronunciation sequences corresponding to the first speech data based on the acoustic state posterior probability of the first speech data; selecting a second optimal candidate pronunciation sequence based on the state transition probability between each adjacent pronunciation state in a plurality of second candidate pronunciation sequences, and obtaining the first wake-up word based on the second optimal candidate pronunciation sequence.
[0087] In some embodiments, before performing the first wake-up operation according to the first wake-up word, the decoding module 730 further includes: obtaining the score of the second optimal candidate pronunciation sequence; and performing the first wake-up operation according to the first wake-up word in response to the score of the second optimal candidate pronunciation sequence exceeding the threshold score of the second optimal candidate pronunciation sequence.
[0088] Those skilled in the art will understand that, in the above-described method of the specific implementation, the order in which each step is written does not imply a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.
[0089] Please see Figure 8 , Figure 8 This is a schematic diagram of a framework of an embodiment of the electronic device 80 of this application. The electronic device 80 includes a memory 81 and a processor 82 coupled to each other. The processor 82 is used to execute program instructions stored in the memory 81 to implement the steps of any of the above-described voice wake-up method embodiments. In a specific implementation scenario, the electronic device 80 may include, but is not limited to, a microcomputer, a server, etc. In addition, the electronic device 80 may also include mobile devices such as laptops and tablets, which are not limited here.
[0090] Specifically, processor 82 controls itself and memory 81 to implement the steps of any of the above-described voice wake-up method embodiments. Processor 82 can also be referred to as a CPU (Central Processing Unit). Processor 82 may be an integrated circuit chip with signal processing capabilities. Processor 82 can also be a general-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. A general-purpose processor can be a microprocessor or any conventional processor. Furthermore, processor 82 can be implemented using integrated circuit chips.
[0091] Please see Figure 9 , Figure 9 This is a schematic diagram of a framework of an embodiment of the computer-readable storage medium 90 of this application. The computer-readable storage medium 90 stores program instructions 901 that can be executed by a processor. The program instructions 901 are used to implement the steps in any of the above-described voice wake-up method embodiments.
[0092] In some embodiments, the functions or modules of the apparatus provided in this disclosure can be used to perform the methods described in the above method embodiments. The specific implementation can be referred to the description of the above method embodiments, and for the sake of brevity, it will not be repeated here.
[0093] The description of the various embodiments above tends to emphasize the differences between the various embodiments. The similarities or similarities between them can be referred to, and for the sake of brevity, they will not be repeated here.
[0094] In the several embodiments provided in this application, it should be understood that the disclosed methods and apparatus can be implemented in other ways. For example, the apparatus implementations described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection of devices or units may be electrical, mechanical, or other forms.
[0095] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0096] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods of various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
Claims
1. A voice wake-up method, characterized in that, include: Acquire the first voice data; Perform speech recognition on the first speech data to obtain the first wake-up word represented by the first speech data; Perform the first wake-up operation according to the first wake-up word; as well as A preset penalty operation is performed on the first wake-up word within a preset time after the first wake-up operation. The preset penalty operation is used to reduce the probability of recognizing the second voice data acquired within the preset time as the first wake-up word. The method further includes, before performing a preset penalty operation on the first wake-up word within a preset time after the first wake-up operation, the method further includes: in response to the first wake-up word applying a low crosstalk scheme, determining whether the speech recognition confidence of the first wake-up word exceeds a confidence threshold; and in response to the first wake-up word's speech recognition confidence not exceeding the confidence threshold, performing the preset penalty operation on the first wake-up word within a preset time after the first wake-up operation.
2. The method according to claim 1, characterized in that, The preset penalty operation on the first wake-up word within a preset time after the first wake-up operation includes: In response to acquiring second voice data within the preset time, the second voice data is used for speech recognition using the penalty coefficient of the first wake-up word to obtain the second wake-up word represented by the second voice data, wherein the penalty coefficient of the first wake-up word is used to reduce the probability that the second voice data is recognized as the first wake-up word; Perform the second wake-up operation according to the second wake-up word.
3. The method according to claim 2, characterized in that, The penalty coefficient of the first wake word decreases over time.
4. The method according to any one of claims 2-3, characterized in that, The step of performing speech recognition on the second speech data using the penalty coefficient of the first wake-up word to obtain the second wake-up word represented by the second speech data includes: Obtain the acoustic state posterior probability of the second speech data; The second wake-up word is obtained by decoding based on the acoustic state posterior probability of the second speech data and the penalty coefficient of the first wake-up word, wherein the penalty coefficient is used to reduce the state transition probability of at least some pronunciation states of the second speech data during the decoding process.
5. The method according to claim 4, characterized in that, The second wake-up word is obtained by decoding the acoustic state posterior probability based on the second speech data and the penalty coefficient of the first wake-up word, including: Based on the acoustic state posterior probability of the second speech data, determine the state transition probability between each adjacent pronunciation state in a plurality of first candidate pronunciation sequences corresponding to the second speech data; At least one set of sub-pronunciation sequences is selected from the plurality of first candidate pronunciation sequences, and the state transition probability between adjacent pronunciation states in each set of sub-pronunciation sequences is reduced by using the penalty coefficient. The sub-pronunciation sequence consists of at least two consecutive pronunciation states in the first candidate pronunciation sequence, and the pronunciation states in the sub-pronunciation sequence other than the beginning and end are the pronunciation states contained in the first wake word.
6. The method according to claim 5, characterized in that, The step of using the penalty coefficient to reduce the state transition probability between adjacent pronunciation states in each group of sub-pronunciation sequences includes: For each group of adjacent pronunciation states in each sub-pronunciation sequence, the state transition probability of the adjacent pronunciation state is multiplied by the penalty coefficient corresponding to the adjacent pronunciation state to obtain the current state transition probability between the adjacent pronunciation states, wherein the penalty coefficient is less than 1, and the penalty coefficients corresponding to the adjacent pronunciation states in different groups may be the same or different. And / or, prior to the second wake-up operation performed according to the second wake-up word, the method further includes: Obtain the score of the first optimal candidate pronunciation sequence; In response to the score of the first optimal candidate pronunciation sequence exceeding the threshold score of the first optimal candidate pronunciation sequence, the second wake-up operation according to the second wake-up word is performed.
7. The method according to claim 1, characterized in that, The step of performing speech recognition on the first speech data to obtain the first wake-up word represented by the first speech data includes: Obtain the acoustic state posterior probability of the first speech data; The first wake-up word is obtained by decoding based on the acoustic state posterior probability of the first speech data.
8. The method according to claim 7, characterized in that, The acquisition of the acoustic state posterior probability of the first speech data includes: Extract the spectral features of the first speech data; The acoustic state prediction of the first speech data is obtained by using a speech recognition model to perform acoustic state prediction on the spectral features. The decoding method based on the acoustic state posterior probability of the first speech data is Viterbi decoding. The decoding based on the acoustic state posterior probability of the first speech data to obtain the first wake-up word includes: Based on the acoustic state posterior probability of the first speech data, determine the state transition probability between each adjacent pronunciation state in several second candidate pronunciation sequences corresponding to the first speech data. Based on the state transition probabilities between adjacent pronunciation states in the plurality of second candidate pronunciation sequences, a second optimal candidate pronunciation sequence is selected, and the first wake-up word is obtained based on the second optimal candidate pronunciation sequence; Before performing the first wake-up operation according to the first wake-up word, the method further includes: Obtain the score of the second optimal candidate pronunciation sequence; In response to the score of the second optimal candidate pronunciation sequence exceeding the threshold score of the second optimal candidate pronunciation sequence, the first wake-up operation according to the first wake-up word is performed.
9. A voice wake-up device, characterized in that, include: The acquisition module is used to acquire the first voice data; The speech recognition module is used to perform speech recognition on the first speech data to obtain the first wake-up word represented by the first speech data; The decoding module is used to perform a first wake-up operation according to the first wake-up word; as well as In response to the first wake-up word, a low crosstalk scheme is applied to determine whether the speech recognition confidence of the first wake-up word exceeds the confidence threshold. In response to the fact that the speech recognition confidence of the first wake-up word does not exceed the confidence threshold, the preset penalty operation on the first wake-up word is performed within a preset time after the first wake-up operation; A preset penalty operation is performed on the first wake-up word within a preset time after the first wake-up operation. The preset penalty operation is used to reduce the probability of recognizing the second voice data acquired within the preset time as the first wake-up word.
10. An electronic device, characterized in that, The device includes a memory and a processor coupled to each other, the processor being configured to execute program instructions stored in the memory to implement the voice wake-up method according to any one of claims 1 to 8.
11. A computer-readable storage medium having program instructions stored thereon, characterized in that, When the program instructions are executed by the processor, they implement the voice wake-up method according to any one of claims 1 to 8.