Command word recognition method, wake-up word recognition method, electronic equipment and storage medium
By optimizing the command word recognition method in smart glasses, combining speech recognition model and scene matching, and adjusting the matching threshold, the problem of low efficiency and low accuracy of command word recognition in smart glasses is solved, and the user experience is improved.
Patent Information
- Application Number
- CN202510694869.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-07-11
AI Technical Summary
The existing technology has not effectively solved the problem of low efficiency and low accuracy of command word recognition in smart glasses.
By inputting the audio feature data of the audio to be identified into the speech recognition model, the candidate string and probability are obtained, the command word set of the current scene is matched, and the threshold is adjusted according to the matching score, the matching process is optimized, and the accuracy and efficiency of command word recognition are improved.
It improves the recognition efficiency and accuracy of command words in smart glasses, enhances the user experience, and can recognize less standard voice commands of users.
Smart Images

Figure CN120299458A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of speech recognition technology, and in particular, to methods for identifying command words and wake-up words, electronic devices, and storage media. Background Art
[0002] With the development of intelligent wearable devices (especially smart glasses), the application of audio recognition technology in devices such as intelligent wearable devices has become increasingly common, providing a more convenient human-computer interaction method for users and also improving the user experience of using intelligent wearable devices. In intelligent wearable devices, the audio recognition technology mostly identifies whether an offline command word is hit in the audio. If a certain offline command word is hit, the recognition is ended and the operation corresponding to the offline command word is executed.
[0003] However, in related technologies, there are generally technical problems of low command word recognition efficiency and low accuracy. For this technical problem, no effective solution has been proposed yet. Summary of the Invention
[0004] Embodiments of this application provide methods for identifying command words and wake-up words, electronic devices, and storage media to at least solve the technical problems of low command word recognition efficiency and low accuracy in smart glasses in related technologies.
[0005] According to one aspect of the embodiments of this application, a method for identifying command words is provided, which is applied to a smart glass and includes: inputting audio feature data of an audio to be recognized into a speech recognition model to obtain a plurality of candidate strings and the probability corresponding to each candidate string; obtaining a command word set corresponding to the current scene; matching each of the plurality of candidate strings with the command words in the command word set one by one; determining whether the matching threshold needs to be adjusted in the current situation; determining that the highest matching score of the plurality of candidate strings reaches the matching threshold; and obtaining the command word with the highest matching score in the command word set as the recognition result.
[0006] Optionally, the matching threshold is initially an initial matching threshold. The determining whether the matching threshold needs to be adjusted in the current situation includes: when the highest matching score of the plurality of candidate strings is lower than the initial matching threshold, determining whether the difference between the highest matching score and the initial matching threshold is equal to or less than a preset difference threshold; and when the difference is equal to or less than the preset difference threshold, lowering the matching threshold to the highest matching score of the plurality of candidate strings.
[0007] Optionally, the method for determining the initial matching threshold includes: matching multiple training samples used during the training of the speech recognition model with the command words in the command word set corresponding to the preset scenario one by one, to obtain the highest matching score for each training sample in the multiple training samples; using the highest matching score of each training sample to traverse the alternative threshold range, and determining the balance point of the precision and recall curves corresponding to multiple alternative thresholds as the initial matching threshold.
[0008] Optionally, the method further includes training the speech recognition model, wherein training the speech recognition model includes: inputting the training samples in the current training batch into the speech recognition model to obtain the predicted values of the training samples, where the training samples include: wake-up word samples and command word samples; determining the loss of the wake-up word samples and the loss of the command word samples based on the predicted values of the training samples; determining the target loss through the following formula:
[0009] (1 + β)Loss kws +(1 - β)Loss cw
[0010] where Loss kws is the loss of whether a training sample is a wake-up word, and Loss cw is the loss of whether a training sample is a command word, and β is a specified coefficient greater than 0; using the target loss to adjust the network parameters of the speech recognition model to obtain the trained speech recognition model.
[0011] Optionally, the speech recognition model at least has a wake-up word recognition function and a command word recognition function, and the method further includes: during the training process, adjusting the specified coefficient to become larger based on ensuring the accuracy of wake-up.
[0012] Optionally, the speech recognition model at least has a wake-up word recognition function and a command word recognition function, and the method further includes: during the training process, when the convergence of the command word samples is too poor, adjusting the specified coefficient to become smaller.
[0013] Optionally, inputting the audio feature data of the audio to be recognized into a speech recognition model to obtain multiple candidate strings and the probability corresponding to each candidate string includes: inputting the audio feature data of the audio to be recognized into the speech recognition model to obtain the characters corresponding to each frame and the probability corresponding to each character; performing connectionist temporal classification (CTC) decoding on the characters corresponding to each frame and the probability corresponding to each character to obtain the multiple candidate strings and the probability corresponding to each candidate string.
[0014] According to another aspect of the embodiments of the present application, there is also provided a wake-up word recognition method, which is applied to smart glasses and includes: inputting the audio feature data of the subsequent audio to be recognized into the speech recognition model in the above command word recognition method to obtain a plurality of subsequent candidate strings and the probability corresponding to each subsequent candidate string; determining that the subsequent current scenario requires waking up; matching the plurality of subsequent candidate strings with the wake-up words one by one to obtain wake-up matching scores; determining that at least one wake-up matching score reaches the wake-up word matching threshold.
[0015] According to still another aspect of the embodiments of the present application, there is also provided an electronic device, including: a processor, and a memory storing a program, where the program includes instructions that, when executed by the processor, cause the processor to execute the method in any of the above embodiments.
[0016] According to still another aspect of the embodiments of the present application, there is also provided a non-transitory machine-readable medium storing computer instructions for causing a computer to execute the method in any of the above embodiments.
[0017] According to still another aspect of the embodiments of the present application, there is also provided a computer program product, including a computer program that, when executed by a processor of a computer, is used to cause the computer to execute the method in any of the above embodiments.
[0018] In the embodiments of the present application, the audio feature data of the audio to be recognized is input into a speech recognition model to obtain a plurality of candidate strings and the probability corresponding to each candidate string; the command word set corresponding to the current scenario is obtained; the plurality of candidate strings are matched with the command words in the command word set one by one; it is determined whether the matching threshold needs to be adjusted in the current situation; it is determined that the highest matching score of the plurality of candidate strings reaches the matching threshold; the command word with the highest matching score in the command word set is obtained as the recognition result. That is to say, the embodiments of the present application combine the current scenario of smart glasses, restrict the types of supported command words, shorten the command word matching time, and relax the matching threshold during command word matching, so that even if the voice command output by the user is not very standard, the command word can be hit, improving the user experience, and thus solving the technical problems of low efficiency and low accuracy of command word recognition in smart glasses in the related art, and achieving the technical effects of improving the efficiency and accuracy of command word recognition in smart glasses.
[0019] Details of one or more embodiments of the present application are set forth in the following drawings and description to make other features, objects, and advantages of the present application more concise and understandable. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other embodiments can be obtained based on these drawings.
[0021] Figure 1 It is a schematic diagram of a command word recognition framework provided according to an embodiment of the present application;
[0022] Figure 2 It is a flowchart of a command word recognition method provided according to an embodiment of the present application;
[0023] Figure 3 It is a flowchart of a calculation method for FBANK features provided according to an embodiment of the present application;
[0024] Figure 4 It is a schematic diagram of a CTC decoding method provided according to an embodiment of the present application;
[0025] Figure 5 It is a schematic diagram of another CTC decoding method provided according to an embodiment of the present application;
[0026] Figure 6 It is a flowchart of a wake word recognition method provided according to an embodiment of the present application;
[0027] Figure 7 It is a flowchart of a speech recognition method provided according to an embodiment of the present application;
[0028] Figure 8 It is a structural block diagram of a command word recognition device provided according to an embodiment of the present application;
[0029] Figure 9 It is a structural block diagram of a wake word recognition device provided according to an embodiment of the present application;
[0030] Figure 10 It is a schematic diagram of the structure of the electronic device in this embodiment. Detailed implementation manners
[0031] The following will describe the embodiments of this embodiment in more detail with reference to the accompanying drawings. Although some embodiments of this embodiment are shown in the accompanying drawings, it should be understood that this embodiment can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand this embodiment. It should be understood that the accompanying drawings and embodiments of this embodiment are only for exemplary purposes and are not used to limit the protection scope of this embodiment.
[0032] AsFigure 1 As shown in the figure, the existing command word recognition solutions usually include the following four parts: feature extraction, speech recognition model (such as based on Convolutional Neural Networks (CNN), Long Short-Term Memory (LSTM), Transformer (a sequence language model based on attention mechanism), etc., typically based on Deep Feedforward Sequential Memory Networks (DFSMN)), Connectionist Temporal Classification (CTC) decoding, and command word matching. Among them, the command word matching process usually matches the recognized string with all the command words in the smart glasses command word library one by one. If a match is found, the corresponding operation of the command word is executed; if no match is found, no operation is executed. This command word matching method has a large amount of computation and usually has the technical problem of low command word recognition efficiency. In addition, it cannot recognize strings that are not very standard but may belong to command words output by users. For example, the string "Open WeChat" corresponding to the audio output by the user cannot be matched with "Open WeChat" in the smart glasses command word library, and the string "Listen to music" corresponding to the audio output by the user cannot be matched with "Open music" in the smart glasses command word library, etc. Therefore, there is also the technical problem of low command word recognition accuracy in the related technologies.
[0033] To solve the above technical problems, relevant solutions are provided in the embodiments of the present application, which are described in detail below. As Figure 2 shown in the flowchart of the command word recognition method, the execution process of this method defaults that the smart glasses are in the wake-up state. Specifically, this method includes the following steps:
[0034] Step S202: Input the audio feature data of the audio to be recognized into a speech recognition model to obtain multiple candidate strings and the probability corresponding to each candidate string.
[0035] It should be noted that the above audio to be recognized is the audio output by the user and contains command words or similar command words, such as "play music", "answer the phone", "open WeChat", "please help me open the address book", etc. In addition, during the process of speech recognition, there may be a large part of the time when the user is not speaking, that is, in a mute state. If the audio data is continuously recorded during this period, due to the existence of mute, it will consume the computing resources of speech recognition, and these calculations are meaningless. Therefore, the embodiments of the present application use voice activity detection (VAD) to determine the start point and end point of the audio, exclude the mute part, and then perform sampling determination in a certain audio sampling method (for example, sampling frequency 16khz, bit width 16bit, single channel, etc.).
[0036] The audio features of the above audio to be recognized usually use spectral features. Typical spectral features include Mel Frequency Cepstral Coefficents (MFCC) and Filterbank Features (FBANK). Among them, FBANK is obtained by taking the sum of the squared magnitudes of the power spectra of Mel filters and then taking the logarithm. Compared with MFCC, it omits the calculation of one-step discrete cosine transform. The embodiments of the present application take FBANK as an example and mainly introduce the calculation process of FBANK features, as Figure 3 shown, the calculation process of FBANK features includes:
[0037] S11, pre-emphasize the audio. The purpose of this step is to enhance the high-frequency signal;
[0038] S12, frame segmentation, which separates the audio signal into frames of 10ms each;
[0039] S13, windowing, which is to prevent spectral leakage. Each time, a 25ms signal is used to calculate the features, that is, each time it moves 10ms, and actually a 25ms signal is used, with 15ms of historical overlapping information.
[0040] S14, Fourier transform. After the Fourier transform, the frequency-domain signal, that is, the spectrum, is obtained using the time-domain signal.
[0041] S15, accumulate the frequency domain at a certain time to obtain the spectrogram, mainly including taking the power spectrum and squaring the amplitude;
[0042] S16, map the frequency to the Mel frequency scale through the Mel filter bank;
[0043] S17. Take the logarithm to obtain the FBANK features. In wake-up applications, generally 40 or 80 Mel filter banks are used, that is, each audio frame corresponds to 40 or 80 outputs. The above calculation process can also add a frame splicing process. Since the features of one frame sometimes seem to have insufficient information for predicting a label, especially when the label pronunciation is relatively long. For example, when using pinyin as the label, it is longer than the phoneme time. So a better method is to combine the previous and next frames to infer the current frame. That is, "J past frames + the current frame + K future frames" are used as input data together. For example, set J = 5 and K = 5, that is, 11 frames are used to infer the label of the current frame. If frame splicing is used before, a frame skipping process can also be selected. Because there may be information redundancy, downsampling can be appropriately performed, such as frame skipping sampling to reduce the data volume and the subsequent processing calculation amount. Usually, 1 frame is sampled every 2 frames, or 1 frame is sampled every 3 frames. In practice, it is comprehensively determined whether to skip frames and the frame skipping value in combination with the above J and K values and the training results.
[0044] Step S204, obtain the set of command words corresponding to the current scene.
[0045] It can be understood that the above current scene can correspond to the application currently running on the smart glasses. The current scene can be "answering a call", "music playing" and other scenes. The sets of command words corresponding to different scenes are different. For example, in the scene of "answering a call", the set of command words can be "answer", "hang up", etc. In the scene of "music playing", the set of command words can be "turn up the volume", "turn down the volume", "play the previous song", "play the next song", "stop playing", etc. In other scenes, the set of command words can be "take a photo", etc.
[0046] Step S206, match each of the multiple candidate strings with the command words in the set of command words one by one.
[0047] Optionally, in step S206, the candidate strings can be taken out one by one from the multiple candidate strings to check whether the candidate string is in the above set of command words. In this step, the set of command words for matching is reduced from the set of command words corresponding to all scenes of the smart glasses to the set of command words corresponding to the current scene (which can be one scene), and the types of command words are significantly reduced, thereby reducing the calculation amount.
[0048] Step S208, determine whether the matching threshold needs to be adjusted in the current situation.
[0049] It should be noted that the above current situation is the matching score obtained by matching each of the multiple candidate strings with the command words in the command word set one by one. In the embodiments of the present application, it is determined whether to adjust the matching threshold according to the matching score, including but not limited to: when the highest matching score of the multiple candidate strings is lower than the initial matching threshold, adjusting the matching threshold based on the difference between the highest matching score and the initial matching threshold, or when the highest matching score of the multiple candidate strings is lower than the initial matching threshold, adjusting the matching threshold according to the ratio between the highest matching score and the initial matching threshold, etc. It can be understood that in the embodiments of the present application, the matching threshold is appropriately relaxed based on the highest matching score (or appropriately relaxed based on a certain recognition accuracy), making it easier to hit the command word, that is, even if the user's speech is not very standard, it can still be hit, improving the user experience.
[0050] Optionally, in the embodiments of the present application, the matching score of the candidate string can be determined by calculating the similarity between each of the multiple candidate strings and each command word in the above command word set. Exemplarily, assume that the current scenario is "music playback", the candidate strings include candidate string 1: "turn up the volume" and candidate string 2: "make the sound louder", and the command word set corresponding to this current scenario includes: command word 1: "make the sound a little louder", command word 2: "make the sound a little quieter", command word 3: "play the previous song", command word 4: "play the next song", command word 5: "stop playing". Then, according to the method of calculating similarity, the matching score between candidate string 1 and command word 1 can be determined to be 0.8, the matching score between candidate string 1 and command word 2 can be determined to be 0.4, and the matching scores between candidate string 1 and command words 3, 4, and 5 are 0. Similarly, the matching score of candidate string 2 can be calculated, which will not be elaborated in this example.
[0051] Optionally, in the embodiments of the present application, when the highest matching score of the multiple candidate strings is the same as the initial matching threshold, the matching threshold may not be adjusted either, and this matching threshold is the initial matching threshold.
[0052] Step S210, determine that the highest matching score of the multiple candidate strings reaches the matching threshold.
[0053] Since the above step S206 matches each of the multiple candidate strings with the command words in the command word set one by one, there is a corresponding relationship between the candidate strings and the command words. After determining the candidate string with the highest matching score, the corresponding command word in the command word set can be determined, and then this command word can be used as the recognition result. Therefore, the embodiments of the present application also propose step S212, where this step S212 includes: obtaining the command word with the highest matching score in the command word set as the recognition result.
[0054] Through the above steps S202 - S212, the audio feature data of the audio to be recognized is input into a speech recognition model to obtain multiple candidate strings and the probability corresponding to each candidate string; the set of command words corresponding to the current scene is obtained; the multiple candidate strings are matched one by one with the command words in the set of command words; it is determined whether the matching threshold needs to be adjusted in the current situation; it is determined that the highest matching score of the multiple candidate strings reaches the matching threshold; the command word with the highest matching score in the set of command words is obtained as the recognition result. That is to say, the embodiment of the present application combines the current scene of the smart glasses, restricts the types of supported command words, shortens the command word matching time, and relaxes the matching threshold during command word matching, so that even if the voice command output by the user is not very standard, the command word can still be hit, improving the user experience, and further solving the technical problems of low efficiency and low accuracy of command word recognition in related technologies, achieving the technical effect of improving the efficiency and accuracy of command word recognition of smart glasses.
[0055] The above matching threshold is initially an initial matching threshold. In a possible implementation, determining whether the matching threshold needs to be adjusted in the current situation includes:
[0056] S21, when the highest matching score of the multiple candidate strings is lower than the initial matching threshold, determining whether the difference between the highest matching score and the initial matching threshold is equal to or less than a preset difference threshold;
[0057] S22, when the difference is equal to or less than the preset difference threshold, lowering the matching threshold to the highest matching score of the multiple candidate strings.
[0058] That is, in the above possible implementation, the matching threshold is adjusted based on the difference between the highest matching score and the initial matching threshold. Exemplarily, assume that the above multiple candidate strings are the top 5 strings sorted by probability, and there are 3 command words in the current scene. The above 5 strings can be used to match the 3 command words one by one, and finally a highest matching score is obtained, such as 0.28, but the initial matching threshold is 0.3. Although the highest matching score is lower than 0.3, the difference is not much, which is 0.02. In this case, the initial matching threshold can be lowered to 0.28, so the command word with a matching score of 0.28 in the set of command words can be successfully matched.
[0059] Through the above steps S21 - S22, when the highest matching score of the multiple candidate strings is lower than the initial matching threshold, the initial matching threshold is lowered based on the highest matching score, making it easier to hit the command word, so that even if the user's speech is not very standard, it can still be hit, further improving the user experience.
[0060] The embodiments of the present application also propose an optional determination method for the above initial matching threshold, where the optional determination method includes:
[0061] S31, match each of the multiple training samples used in the training of the speech recognition model with the command words in the command word set corresponding to the preset scenario one by one, and obtain the highest matching score for each training sample in the multiple training samples;
[0062] S32, use the highest matching score of each training sample to traverse the alternative threshold range, and determine the balance point of the precision and recall curves corresponding to the multiple alternative thresholds as the initial matching threshold.
[0063] It should be noted that the above precision is the probability that the sample actually being positive among all samples predicted as positive, and the above accuracy is the percentage of the correctly predicted results in the total samples. In addition, it should be noted that the initial matching thresholds corresponding to different scenarios may be the same or different. For example, for the scenario of "answering a call", the initial matching threshold may be 0.8, for the scenario of "playing music", the initial matching threshold may be 0.7, and for the scenario of "taking a photo", the initial matching threshold may be 0.7.
[0064] It can be understood that each of the above alternative thresholds corresponds to a precision and a recall. The highest matching scores of the multiple training samples traverse the alternative threshold range respectively to draw the corresponding P-R graph (where the abscissa is the recall rate and the ordinate is the precision rate). Based on this P-R graph, a balance point (that is, a point where both the precision rate and the recall rate are relatively high) is determined, and this balance point is determined as the initial matching threshold, such as 0.75.
[0065] Optionally, in the embodiments of the present application, the above alternative threshold range may be 0.01 to 1, and the step size can be 0.01 during the traversal.
[0066] By determining the balance point of the precision and recall curves as the initial matching threshold through the above steps S31 to S32, a threshold with the best performance (precision and recall) of the training samples can be selected from the alternative threshold range, further improving the accuracy of command word recognition.
[0067] To improve the recognition performance of the speech recognition model, the embodiments of the present application also propose a specific implementation method for training the speech recognition model, where training the speech recognition model includes:
[0068] S41, input the training samples in the current training batch into the speech recognition model to obtain the predicted values of the training samples;
[0069] It should be noted that the above training samples may include: wake-up word samples and command word samples. Among them, the wake-up word is used to wake up the smart glasses, and the wake-up word samples can be "Xiaoxi, Xiaoxi", "Hello, Xiaoxi", etc. In addition, it should be noted that the predicted value of the above training sample is the predicted probability.
[0070] S42. Determine the loss of the wake-up word sample and the loss of the command word sample based on the predicted value of the training sample.
[0071] Optionally, in the embodiments of the present application, the above loss may be the cross-entropy loss. Exemplarily, assuming that the predicted probability is P and y is the value of the sample (the command word sample is 1 and the wake-up word sample is 0), then the loss of whether a training sample is a wake-up word is Loss kws =-log(1 - P), and the loss of whether a training sample is a command word is Loss cw =-log(P).
[0072] In addition, considering the problem of unbalanced training samples, the embodiments of the present application also propose to use the focal loss (Focal Loss) to replace the above cross-entropy loss. Among them, the calculation method of the Focal Loss is shown in Formula 1:
[0073] -y(1 - p) γ log(p)-(1 - y)(p) γ log(1 - p) (Formula 1)
[0074] Among them, the range of γ is (0, 10). The larger P is, the smaller the value of the loss function is, indicating that the contribution of the category that is easier to converge to the loss function is smaller. In this way, the training weight result is biased towards the category with fewer samples.
[0075] S43. Determine the target loss through the following Formula 2:
[0076] (1 + β)Loss kws +(1 - β)Loss cw (Formula 2)
[0077] Among them, Loss kws is the loss of whether a training sample is a wake-up word, and Loss cw is the loss of whether a training sample is a command word.
[0078] Optionally, the above β is a hyperparameter, and its value range can be from 0 to 1, which is preset and fixed in a training batch. For example, β is set to 0.3. Additionally, when updating the training parameters, the loss of a batch of data is used for the update. For example, for 32 training samples, 22 of which are command word samples, 32 losses can be obtained. These 32 losses can be added together (or averaged) to obtain the final loss value for this time. It should be noted that when adding them together, the loss of a training sample can be determined using the above formula 2.
[0079] S44. Use the target loss to adjust the network parameters of the speech recognition model to obtain a trained speech recognition model.
[0080] Since the above speech recognition model at least has the functions of wake word recognition and command word recognition, setting β to a fixed value cannot meet the functional requirements of the speech recognition model. For example, the speech recognition model gives priority to ensuring the wake-up function, etc. Therefore, the embodiments of the present application also propose:
[0081] S51. During the training process, based on ensuring the accuracy of wake-up, adjust the specified coefficient to become larger.
[0082] The larger β is, the more the result tends to wake-up. In this way, the convergence effect of wake-up recognition will be preferentially ensured, and the obtained model will also preferentially ensure the accuracy of wake-up.
[0083] Alternatively, use S52 to replace the above S51, where S52 is that during the training process, when the convergence of the command word samples is too poor, adjust the specified coefficient to become smaller, so that the model biases towards the convergence of the command words.
[0084] It should be noted that the performance of the above poor convergence of the command word samples may include: as the training progresses, the training loss of the model does not decrease significantly, and may even fluctuate and cannot be stabilized at a low level, that is, the model fails to effectively learn the features of the command word samples; the recognition accuracy of the command words remains at a low level during the training process and does not even increase with the increase of training iterations; the model frequently makes mistakes when predicting the command words, resulting in a large number of misidentifications or misclassifications, etc.
[0085] In a possible implementation manner, the above inputting the audio feature data of the audio to be recognized into a speech recognition model to obtain a plurality of candidate strings and the probability corresponding to each candidate string includes:
[0086] S61. Input the audio feature data of the audio to be recognized into the speech recognition model to obtain the characters corresponding to each frame and the probability corresponding to each character.
[0087] It can be understood that the characters corresponding to each frame and the probabilities corresponding to each character form a temporal label matrix. For example, for 16-frame FBANK feature data and an acoustic model with 400 characters, the output matrix is [16, 400]. This step realizes the mapping from FBANK to characters based on the acoustic model of the neural network.
[0088] S62. Perform connectionist temporal classification (CTC) decoding on the characters corresponding to each frame and the probabilities corresponding to each character to obtain the multiple candidate strings and the probabilities corresponding to each candidate string.
[0089] It can be understood that step S62 is to find N best decoding paths in the above-mentioned temporal label matrix obtained in S61. Usually, there are greedy algorithms, beam search algorithms, prefix beam search algorithms, etc.
[0090] CTC decoding belongs to streaming decoding (online decoding). Streaming decoding performs inference and decoding on partial audio features and feeds back the decoding results to the user in real time. For example, when the user says "Help me check the train tickets to Shanghai tomorrow", the process of streaming decoding is as follows:
[0091] The 1st CTC decoding: Help me
[0092] The 2nd CTC decoding: Help me check
[0093] The 3rd CTC decoding: Help me check tomorrow
[0094] The 4th CTC decoding: Help me check tomorrow to
[0095] The 5th CTC decoding: Help me check tomorrow to Shanghai
[0096] The 6th CTC decoding: Help me check tomorrow to Shanghai's
[0097] The 7th CTC decoding: Help me check tomorrow to Shanghai's train tickets
[0098] The greedy algorithm retains the character with the highest probability at each step. The algorithm is simple, but the accuracy is reduced. For example, Table 1 below shows the probabilities at 3 moments, that is, 3 frames T1, T2, T3, and the labels are also 3, namely: blank, A, and B:
[0099] Table 1
[0100] T1 T2 T3 Blank 0.5 0.4 0.6 A 0.2 0.3 0.3 B 0.3 0.3 0.1
[0101] If the blank label is selected, all 3 frames are blank. The probability of the entire path (i.e., the product of all label probabilities) is 0.12. As Figure 4 shown.
[0102] If the A label is selected, there are three combined paths: "A--", "--A", and "-A-", where "-" represents a blank label. The probability of label A is the sum of the probabilities of these 3 paths, 0.09 + 0.048 + 0.06 = 0.198, which is greater than the probability of the blank label. Similarly, the probability of label B is calculated, and finally it is known that the probability of label A is the largest, that is, the decoding result of the greedy algorithm is A. As Figure 5 shown.
[0103] The greedy algorithm selects the maximum probability at each moment, but the purpose of CTC decoding is to select the route with the maximum probability, and the two are sometimes not consistent. The beam search algorithm is to address this inconsistency by maintaining N optimal routes at each time, rather than just using the character with the maximum probability at the current moment. Here, N is a hyperparameter. The beam search algorithm has a greater computational cost for decoding, but the result is more accurate.
[0104] Prefix beam search retains N branches with the maximum probability at each step. If it is found that there are the same routes at the time nodes that have been processed before, they are merged, which is equivalent to increasing the diversity. Compared with the beam search algorithm for decoding, the prefix beam search decoding has higher accuracy. Therefore, in the embodiments of the present application, prefix beam search is used to retain a beam of a preset length, for example, 5 beams with the highest scores.
[0105] Corresponding to the application scenario of the command word recognition method provided in the embodiments of the present application, the embodiments of the present application also provide a wake word recognition method. As Figure 6 shown, the method includes:
[0106] S602, input the audio feature data of the subsequent audio to be recognized into the above speech recognition model to obtain multiple subsequent candidate strings and the probability corresponding to each subsequent candidate string.
[0107] It should be noted that the above subsequent audio to be recognized is an audio output by the user that contains a wake word or a similar wake word, for example, "Xiaoxi, Xiaoxi", "Hello, Xiaoxi", etc. The recognition process of the above subsequent audio to be recognized is the same as the audio recognition process in the above command word recognition method, and the audio feature type of the audio to be recognized is also the same as the feature type used in the above command word recognition method. Reference can be made to the previous description and will not be specifically introduced here.
[0108] S604, determine that the subsequent current scenario requires waking up;
[0109] Optionally, in the embodiments of the present application, the above-mentioned situations that require waking up include but are not limited to: if the smart glasses are in a sleep state and receive the audio to be recognized, it is determined that the current scenario requires waking up; if the user needs to perform new voice control on the application currently running on the smart glasses and receives the audio to be recognized, it is determined that the current scenario requires waking up.
[0110] S606. Match each of multiple subsequent candidate strings with the wake-up word one by one to obtain a wake-up match score.
[0111] The above wake-up word can be one, or can be set to multiple according to user needs. To improve the wake-up speed of the smart glasses, generally the wake-up word is set to one. If the wake-up word is one, in this step S606, multiple subsequent candidate characters can be matched with only this wake-up word one by one to obtain multiple wake-up match scores. Among them, the wake-up match score can be determined by the similarity between the subsequent candidate characters and the wake-up word.
[0112] S608. Determine that at least one wake-up match score reaches the wake-up word match threshold.
[0113] It can be understood that in the embodiment of the present application, when at least one wake-up match score is the same as the wake-up word match threshold, it is determined that the wake-up word is matched, the wake-up is successful and the voice assistant is started.
[0114] Through the above steps S602 - S608, the audio feature data of the subsequent audio to be recognized is input into the above voice recognition model to obtain multiple subsequent candidate strings and the probability corresponding to each subsequent candidate string; it is determined that the subsequent current scene is a scene that needs to be woken up; the multiple subsequent candidate strings are matched with the wake-up word one by one to obtain a wake-up match score; it is determined that at least one wake-up match score reaches the wake-up word match threshold. That is to say, in the embodiment of the present application, combined with the fact that the current scene of the smart glasses is a scene that needs to be woken up, the matching object is limited to only the wake-up word, which shortens the matching time of the wake-up word, and when at least one wake-up match score is the same as the wake-up word match threshold, it is considered that the wake-up is successful, thereby solving the technical problems of low recognition efficiency and low accuracy of the wake-up word of the smart glasses in the related art, and achieving the technical effect of improving the recognition efficiency and accuracy of the wake-up word of the smart glasses.
[0115] Next, specific examples are combined to describe the embodiments of the present application.
[0116] As Figure 7 shown, the embodiment of the present application also provides a voice recognition method, including the following steps:
[0117] S702. Input the audio feature data of the audio to be recognized into a voice recognition model to obtain multiple candidate strings and the probability corresponding to each candidate string.
[0118] S704. Scene matching. If the current scene is a scene that needs to be woken up, then execute S706; if the current scene is a command word recognition scene, then execute S708.
[0119] S706, Wake word recognition. Specifically includes: matching the multiple subsequent candidate strings with the wake word one by one to obtain a wake word matching score; determining that at least one wake word matching score reaches the wake word matching threshold;
[0120] S708, Command word recognition. It includes: obtaining the command word set corresponding to the current scenario; matching the multiple candidate strings with the command words in the command word set one by one; determining whether the matching threshold needs to be adjusted in the current situation; determining that the highest matching score of the multiple candidate strings reaches the matching threshold; obtaining the command word with the highest matching score in the command word set as the recognition result.
[0121] Through the above steps S702 - S708, combined with the current scenario of the smart glasses, command word recognition and wake word recognition are performed, thereby solving the technical problems of low efficiency and low accuracy in wake word recognition of smart glasses in the related art, and achieving the technical effect of improving the efficiency and accuracy of wake word recognition of smart glasses.
[0122] Corresponding to the application scenario and method of the method provided in the embodiments of the present application, the embodiments of the present application also provide a command word recognition device. As Figure 8 shown in the structural block diagram of the command word recognition device according to an embodiment of the present application, it includes:
[0123] The first acquisition module 802 is used to input the audio feature data of the audio to be recognized into a speech recognition model to obtain multiple candidate strings and the probability corresponding to each candidate string;
[0124] The second acquisition module 804 is used to obtain the command word set corresponding to the current scenario;
[0125] The first matching module 806 is used to match the multiple candidate strings with the command words in the command word set one by one;
[0126] The first judgment module 808 is used to judge whether the matching threshold needs to be adjusted in the current situation;
[0127] The first determination module 810 is used to determine that the highest matching score of the multiple candidate strings reaches the matching threshold;
[0128] The first recognition module 812 is used to obtain the command word with the highest matching score in the command word set as the recognition result.
[0129] Through Figure 8The device shown inputs the audio feature data of the audio to be recognized into a speech recognition model to obtain multiple candidate strings and the probability corresponding to each candidate string; obtains the set of command words corresponding to the current scene; matches each of the multiple candidate strings with the command words in the set of command words one by one; determines whether the matching threshold needs to be adjusted in the current situation; determines that the highest matching score of the multiple candidate strings reaches the matching threshold; and obtains the command word with the highest matching score in the set of command words as the recognition result. That is to say, the embodiment of the present application combines the current scene of the smart glasses, restricts the types of supported command words, shortens the command word matching time, and relaxes the matching threshold during command word matching, so that even if the voice command output by the user is not very standard, the command word can be hit, improving the user experience. Furthermore, the technical problems of low command word recognition efficiency and low accuracy in the related art are solved, and the technical effects of improving the command word recognition efficiency and accuracy of the smart glasses are achieved.
[0130] In a possible implementation manner, the above-mentioned matching threshold is initially an initial matching threshold. The first determination module 808 includes: a determination unit, configured to determine whether the difference between the highest matching score and the initial matching threshold is equal to or less than a preset difference threshold when the highest matching score of the multiple candidate strings is lower than the initial matching threshold; an adjustment unit, configured to lower the matching threshold to the highest matching score of the multiple candidate strings when the difference is equal to or less than the preset difference threshold.
[0131] The determination method of the above-mentioned initial matching threshold includes: matching each of the multiple training samples used in the training of the speech recognition model with the command words in the set of command words corresponding to the preset scene one by one to obtain the highest matching score of each training sample in the multiple training samples; using the highest matching score of each training sample to traverse the alternative threshold range, and determining the balance point of the precision rate and recall rate curves corresponding to multiple alternative thresholds as the initial matching threshold.
[0132] Optionally, the above-mentioned device further includes a training module, configured to train the speech recognition model. Among them, training the speech recognition model includes: inputting the training samples in the current training batch into the speech recognition model to obtain the predicted values of the training samples, where the training samples include: wake-up word samples, command word samples; determining the loss of the wake-up word samples and the loss of the command word samples based on the predicted values of the training samples;
[0133] Determine the target loss through the following formula:
[0134] (1 + β)Loss kws +(1 - β)Loss cw
[0135] where Losskws The loss for whether a training sample is a wake word, Loss cw The loss for whether a training sample is a command word, where β is a specified coefficient greater than 0; the network parameters of the speech recognition model are adjusted using this objective loss to obtain a trained speech recognition model.
[0136] The above speech recognition model at least has a wake word recognition function and a command word recognition function. Optionally, in the embodiments of the present application, the above training module is further configured to, during the training process, adjust the specified coefficient to become larger based on ensuring the accuracy of waking up.
[0137] The above speech recognition model at least has a wake word recognition function and a command word recognition function. Optionally, in the embodiments of the present application, the above training module is further configured to, during the training process, when the convergence of the command word samples is too poor, adjust the specified coefficient to become smaller.
[0138] The above first acquisition module 802 includes: an acquisition unit, configured to input the audio feature data of the audio to be recognized into the speech recognition model to obtain the characters corresponding to each frame and the probability corresponding to each character; a decoding unit, configured to perform connectionist temporal classification (CTC) decoding on the characters corresponding to each frame and the probability corresponding to each character to obtain the multiple candidate strings and the probability corresponding to each candidate string.
[0139] The functions of the modules in each device in the embodiments of the present application can refer to the corresponding descriptions in the above method and have corresponding beneficial effects, which will not be elaborated here.
[0140] Corresponding to the application scenario and method of the method provided in the embodiments of the present application, the embodiments of the present application also provide a wake word recognition device. As Figure 9 shown in the structural block diagram of the wake word recognition device according to an embodiment of the present application, it includes:
[0141] A third acquisition module 902, configured to input the audio feature data of the subsequent audio to be recognized into the above speech recognition model to obtain multiple subsequent candidate strings and the probability corresponding to each subsequent candidate string;
[0142] A second determination module 904, configured to determine that the subsequent current scenario requires waking up;
[0143] A second matching module 906, configured to match the multiple subsequent candidate strings with the wake words one by one to obtain a wake matching score;
[0144] A second decision module 908, configured to determine that at least one wake matching score reaches the wake word matching threshold.
[0145] By Figure 9The device inputs the audio feature data of the subsequent audio to be recognized into the above speech recognition model to obtain multiple subsequent candidate strings and the probability corresponding to each subsequent candidate string; determines that the subsequent current scenario requires waking up; matches the multiple subsequent candidate strings with the wake-up word one by one to obtain a wake-up matching score; and determines that at least one wake-up matching score reaches the wake-up word matching threshold. That is to say, in the embodiment of the present application, in combination with the current scenario of the smart glasses requiring waking up, the matching object is limited to only the wake-up word, which shortens the matching time of the wake-up word, and considers the wake-up to be successful when at least one wake-up matching score is the same as the wake-up word matching threshold, thereby solving the technical problems of low recognition efficiency and low accuracy of the wake-up word in the related art, and achieving the technical effects of improving the recognition efficiency and accuracy of the wake-up word of the smart glasses.
[0146] For the functions of the modules in each device in the embodiment of the present application, reference may be made to the corresponding descriptions in the above method, and they have the corresponding beneficial effects, which will not be elaborated here.
[0147] The embodiment of the present application also provides a non-transitory machine-readable medium storing a computer program, wherein the computer program is used to cause the computer to execute the method of the embodiment of the present application when executed by a processor of the computer.
[0148] The embodiment of the present application also provides a computer program product, including a computer program, wherein the computer program is used to cause the computer to execute the method of the embodiment of the present application when executed by a processor of the computer.
[0149] Reference Figure 10 , the block diagram of an electronic device that can be used as a server or a client in the embodiment of the present application will now be described. It is an example of a hardware device that can be applied to various aspects of the present application. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are only examples and are not intended to limit the implementation of the present application described and / or claimed herein.
[0150] As Figure 10As shown, the electronic device includes a computing unit 1001, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1002 or a computer program loaded from a storage unit 1008 into a random access memory (RAM) 1003. In the RAM 1003, various programs and data required for the operation of the electronic device can also be stored. The computing unit 1001, the ROM 1002, and the RAM 1003 are connected to each other via a bus 10010. An input / output (I / O) interface 10010 is also connected to the bus 10010.
[0151] Multiple components in the electronic device are connected to the I / O interface 10010, including: an input unit 1006, an output unit 1007, a storage unit 1008, and a communication unit 10010. The input unit 1006 can be any type of device capable of inputting information into the electronic device. The input unit 1006 can receive input digital or character information and generate key signal inputs related to the user settings and / or function controls of the electronic device. The output unit 1007 can be any type of device capable of presenting information and can include, but is not limited to, a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The storage unit 1008 can include, but is not limited to, magnetic disks and optical discs. The communication unit 10010 allows the electronic device to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks and can include, but is not limited to, a modem, a network card, an infrared communication device, a wireless communication transceiver, and / or a chipset, such as a Bluetooth device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.
[0152] The computing unit 1001 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, a CPU, a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 1001 executes the various methods and processes described above. For example, in some embodiments, the method embodiments of the present application can be implemented as a computer program tangibly contained in a machine-readable medium, such as the storage unit 1008. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device via the ROM 1002 and / or the communication unit 10010. In some embodiments, the computing unit 1001 can be configured to execute the above methods in any other appropriate manner (e.g., by means of firmware).
[0153] The computer program for implementing the method of the embodiments of the present application can be written in any combination of one or more programming languages. These computer programs can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the computer programs are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer programs can be executed entirely on the machine, partially on the machine, executed partially on the machine as an independent software package and partially on a remote machine, or entirely on a remote machine or server.
[0154] In the context of the embodiments of the present application, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable signal medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, or infrared system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0155] It should be noted that the term "including" and its variations used in the embodiments of the present application are open-ended, that is, "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The modifications of "one" and "a plurality" mentioned in the embodiments of the present application are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly specified in the context, it should be understood as "one or more".
[0156] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the embodiments of the present application are all information and data that have been authorized by the user or fully authorized by all parties. And the collection, use, and processing of the relevant data need to comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation entrances are provided for the user to choose to authorize or refuse.
[0157] In the method embodiments provided by the embodiments of the present application, the steps described can be executed in different orders and / or executed in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The protection scope of the present application is not limited in this regard.
[0158] The term "embodiment" in this specification means that the specific features, structures or characteristics described in combination with the embodiments may be included in at least one embodiment of the present application. The phrase appears in various positions in the specification does not necessarily mean the same embodiment, nor does it mean being independent or alternative to other embodiments and mutually exclusive. The embodiments in this specification are all described in a related manner, and the same or similar parts between the embodiments are referred to each other. In particular, for the device, equipment, and system embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts refer to the partial description of the method embodiments.
[0159] The above-described embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of patent protection. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. A command word recognition method applied to smart glasses, comprising: Inputting the audio feature data of the audio to be recognized into a speech recognition model to obtain a plurality of candidate strings and the probability corresponding to each candidate string; Obtaining a set of command words corresponding to the current scenario; Matching each of the plurality of candidate strings with the command words in the set of command words one by one; Judging whether the matching threshold needs to be adjusted under the current situation; Determining that the highest matching score of the plurality of candidate strings reaches the matching threshold; Obtaining the command word with the highest matching score in the set of command words as the recognition result.
2. The method according to claim 1, wherein The matching threshold is initially an initial matching threshold. The judging whether the matching threshold needs to be adjusted under the current situation includes: When the highest matching score of the plurality of candidate strings is lower than the initial matching threshold, judging whether the difference between the highest matching score and the initial matching threshold is equal to or less than a preset difference threshold; When the difference is equal to or less than the preset difference threshold, lowering the matching threshold to the highest matching score of the plurality of candidate strings.
3. The method according to claim 2, wherein, The determination method of the initial matching threshold includes: Matching each of the multiple training samples used in the training of the speech recognition model with the command words in the set of command words corresponding to the preset scenario one by one to obtain the highest matching score of each training sample in the multiple training samples; Using the highest matching score of each training sample to traverse the alternative threshold range, and determining the balance point of the precision rate and recall rate curves corresponding to multiple alternative thresholds as the initial matching threshold.
4. The method according to claim 1, wherein, The method further includes training the speech recognition model. Among them, training the speech recognition model includes: Inputting the training samples in the current training batch into the speech recognition model to obtain the predicted values of the training samples. Among them, the training samples include: wake-up word samples, command word samples; Determining the loss of the wake-up word samples and the loss of the command word samples based on the predicted values of the training samples; Determining the target loss through the following formula: (1 + β)Loss kws +(1 - β)Loss cw Among them, Loss kws is the loss of whether a training sample is a wake word, and Loss cw is the loss of whether a training sample is a command word, where β is a specified coefficient greater than 0; Using the target loss to adjust the network parameters of the speech recognition model to obtain the trained speech recognition model.
5. The method according to claim 4, wherein The speech recognition model at least has a wake-up word recognition function and a command word recognition function. The method further includes: During the training process, based on ensuring the accuracy of waking up, adjusting the specified coefficient to become larger.
6. The method according to claim 5, wherein The speech recognition model at least has a wake-up word recognition function and a command word recognition function. The method further includes: During the training process, when the convergence of the command word samples is too poor, adjusting the specified coefficient to become smaller.
7. According to the method described in claim 1, inputting the audio feature data of the audio to be recognized into a speech recognition model to obtain a plurality of candidate strings and the probability corresponding to each candidate string, including: Inputting the audio feature data of the audio to be recognized into the speech recognition model to obtain the characters corresponding to each frame and the probability corresponding to each character; Performing connectionist temporal classification (CTC) decoding on the characters corresponding to each frame and the probability corresponding to each character to obtain the plurality of candidate strings and the probability corresponding to each candidate string.
8. A wake-up word recognition method, applied to smart glasses, comprising: Inputting the audio feature data of the subsequent audio to be recognized into the speech recognition model according to any one of claims 1 to 7, to obtain a plurality of subsequent candidate strings and the probability corresponding to each subsequent candidate string; Judging that the subsequent current scenario requires waking up; Matching each of the plurality of subsequent candidate strings with the wake-up word one by one to obtain a wake-up matching score; Determining that at least one wake-up matching score reaches the wake-up word matching threshold.
9. An electronic device, comprising a memory, a processor and a computer program stored on the memory, wherein the processor implements the method according to any one of claims 1 to 8 when executing the computer program.
10. A computer-readable storage medium, wherein a computer program is stored in the computer-readable storage medium, and the computer program implements the method according to any one of claims 1 to 8 when being executed by a processor.