Voice wake-up method and device, electronic equipment, storage medium and program product
By using hybrid phoneme state tag optimization classification layer in the speech wake-up acoustic model, the problem that the speech wake-up solution in the prior art cannot be applied on low-power devices is solved, and efficient and accurate speech wake-up detection is achieved.
Patent Information
- Application Number
- CN202510593799.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-05-09
AI Technical Summary
Existing voice wake-up solutions cannot be applied on low-power devices, and the accuracy of predicting single phoneme states is low.
The audio data frame is input through an acoustic model, and a phoneme-level state sequence is output, including single-phoneme states and triphone states. Use mixed phoneme state labels to optimize the classification layer, reduce the number of model output states and reduce resource requirements.
It realizes efficient voice wake-up on low-power devices, improves detection accuracy, and expands the application range.
Smart Images

Figure CN120126455A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of acoustic signal processing, and in particular, to a voice wake-up method, device, electronic device, storage medium, and program product. Background Art
[0002] Voice wake-up means that when a user speaks a voice audio including a wake-up word, an electronic device switches from a sleep state to a voice wake-up state, so as to recognize the user's voice in the voice wake-up state to give a specified response. The voice wake-up method is widely used in various voice control products, such as robots, mobile phones, smart home devices, car machines, learning machines, and wearable devices, etc. Therefore, it is necessary to implement the voice wake-up function.
[0003] Currently, through a neural network model, the triphone state corresponding to each frame of audio signal is predicted, or the phone state corresponding to each frame of audio signal is predicted. However, when predicting the triphone state corresponding to each frame of audio signal, the number of output states is large, and most application devices have limited computing resources and storage resources, and most application devices have requirements for low cost. Therefore, the triphone prediction scheme cannot be applied to low-power devices and has limitations; although predicting the phone state can compress the memory of the neural network model, the distinguishability of the output phone state is worse than that of the triphone state, resulting in lower accuracy of voice wake-up detection. In summary, how to reduce the resources required for the voice wake-up scheme while ensuring accurate voice wake-up is an urgent problem to be solved currently. Summary of the Invention
[0004] The present invention provides a voice wake-up method, device, electronic device, storage medium, and program product, which are used to solve the defect that the voice wake-up scheme in the prior art cannot be applied to low-power devices, and to implement a low-power voice wake-up method.
[0005] The present invention provides a voice wake-up method, including: Inputting each audio data frame in the audio data into an acoustic model to obtain a phone-level state sequence output by the acoustic model; the phone-level state sequence includes the phone-level state classification results of each audio data frame in the audio data, and the phone-level state classification results include phone states and triphone states; Determining whether it is a voice wake-up state based on the phone-level state sequence; Among them, the acoustic model is obtained by optimizing the classification layer in the pre-trained model based on the first sample audio data frames and the corresponding hybrid phoneme state labels; the pre-trained model is trained based on the second sample audio data frames and the corresponding triphone state labels; the hybrid phoneme state labels include the triphone states corresponding to the wake-up state and the monophone states corresponding to the non-wake-up state.
[0006] According to a voice wake-up method provided by the present invention, the acoustic model is trained based on the following method: Based on the second sample audio data frames and the corresponding triphone state labels, the initial model is trained to obtain the pre-trained model; Based on the first sample audio data frames and the corresponding hybrid phoneme state labels, the classification layer in the pre-trained model is optimized to obtain the acoustic model.
[0007] According to a voice wake-up method provided by the present invention, before training the initial model based on the second sample audio data frames and the corresponding triphone state labels to obtain the pre-trained model, it further includes: Obtaining the second sample audio data frames and the corresponding triphone state labels; the triphone state labels include M triphone states; Changing the triphone states corresponding to the non-wake-up state in the triphone state labels to monophone states to obtain the hybrid phoneme state labels; the first sample audio data frames are the same as the second sample audio data frames; Among them, the hybrid phoneme state labels include N triphone states corresponding to the wake-up state and L monophone states corresponding to the non-wake-up state; N + L is less than M.
[0008] According to a voice wake-up method provided by the present invention, the obtaining the second sample audio data frames and the corresponding triphone state labels includes: Obtaining the second sample audio data frames in the sample audio data and the corresponding text data of the sample audio data; Based on the forced alignment result of the sample audio data and the text data, determining the triphone state labels corresponding to the second sample audio data frames; Among them, the sample audio data set containing the sample audio data includes: general audio data and wake-up word audio data; the general audio data does not include the audio data corresponding to the preset wake-up word, and the wake-up word audio data includes the audio data corresponding to the preset wake-up word.
[0009] A voice wake-up method provided by the present invention, training an initial model based on a second sample audio data frame and a triphone state label corresponding to the second sample audio data frame to obtain a trained model, includes: Training an initial model based on a second sample audio data frame, a triphone state label corresponding to the second sample audio data frame, and a binary classification result label corresponding to the second sample audio data frame to obtain a trained model; Wherein, the binary classification result label includes a first classification result or a second classification result, the first classification result is used to represent that the second sample audio data frame belongs to the audio data frame corresponding to the end position of the wake-up word, and the second classification result is used to represent that the second sample audio data frame does not belong to the audio data frame corresponding to the end position of the wake-up word.
[0010] A voice wake-up method provided by the present invention, inputting each audio data frame in audio data into an acoustic model to obtain a phoneme-level state sequence output by the acoustic model, includes: Inputting the acoustic features of each audio data frame in the audio data into a feature extraction layer in the acoustic model to obtain a sequence of feature vectors output by the feature extraction layer; Inputting the sequence of feature vectors into a classification layer in the acoustic model to obtain a phoneme-level state sequence output by the classification layer in the acoustic model.
[0011] The present invention also provides a voice wake-up device, including: A phoneme output module, configured to input each audio data frame in audio data into an acoustic model to obtain a phoneme-level state sequence output by the acoustic model; the phoneme-level state sequence includes phoneme-level state classification results of each audio data frame in the audio data, and the phoneme-level state classification results include single-phoneme states and triphone states; A state determination module, configured to determine whether it is a voice wake-up state based on the phoneme-level state sequence; Wherein, the acoustic model is optimized for a classification layer in the trained model based on a first sample audio data frame and a mixed phoneme state label corresponding to the first sample audio data frame; the trained model is trained based on a second sample audio data frame and a triphone state label corresponding to the second sample audio data frame; the mixed phoneme state label includes triphone states corresponding to the wake-up state and single-phoneme states corresponding to the non-wake-up state.
[0012] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, it implements the voice wake-up method as described in any one of the above.
[0013] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the voice wake-up method described in any one of the above is implemented.
[0014] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the voice wake-up method described in any one of the above is implemented.
[0015] For the voice wake-up method, device, electronic device, storage medium and program product provided by the present invention, each audio data frame in the audio data is input into an acoustic model to obtain a phoneme-level state sequence output by the acoustic model. The acoustic model is optimized for the classification layer in the trained model based on the first sample audio data frame and the corresponding mixed phoneme state label of the first sample audio data frame, and the mixed phoneme state label includes the triphone state corresponding to the wake-up state and the monophone state corresponding to the non-wake-up state. Thus, only when the corresponding phoneme is in the wake-up state, it is necessary to label the triphone state, otherwise only the monophone state with fewer states needs to be labeled, thereby reducing the number of states output by the acoustic model, further reducing the resource requirements of the acoustic model, ensuring that the voice wake-up method can be applied to low-power devices, realizing a low-power voice wake-up method, and ultimately enhancing the application scope of the voice wake-up method. At the same time, the trained model is trained based on the second sample audio data frame and the corresponding triphone state label of the second sample audio data frame. Thus, the triphone state label is first used for training in the first training stage to utilize the fine-grained representation of the triphone and improve the sensitivity of the acoustic model to subtle changes in speech. Then, in the second training stage, only the classification layer is optimized, which can prevent the loss of information learned in the first training stage during the optimization process in the second training stage, thereby ensuring that the classification layer outputs fewer states while improving the performance of the acoustic model, and further improving the detection accuracy of voice wake-up. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0017] Figure 1 is one of the flow diagrams of the voice wake-up method provided by the present invention.
[0018] Figure 2 is another flow diagram of the voice wake-up method provided by the present invention.
[0019] Figure 3 is yet another flow diagram of the voice wake-up method provided by the present invention.
[0020] Figure 4 It is the fourth flow schematic diagram of the voice wake-up method provided by the present invention.
[0021] Figure 5 It is the fifth flow schematic diagram of the voice wake-up method provided by the present invention.
[0022] Figure 6 It is the sixth flow schematic diagram of the voice wake-up method provided by the present invention.
[0023] Figure 7 It is the structural schematic diagram of the voice wake-up device provided by the present invention.
[0024] Figure 8 It is the structural schematic diagram of the electronic device provided by the present invention. Specific embodiments
[0025] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without making creative efforts shall fall within the protection scope of the present invention.
[0026] In current voice wake-up solutions, to ensure the accuracy of voice wake-up detection, most use a neural network model to predict the triphone state corresponding to each frame of audio signal. However, when predicting the triphone state corresponding to each frame of audio signal, the number of output states is relatively large. Therefore, the triphone prediction scheme cannot be applied to low-power devices and has limitations.
[0027] In view of the problem that the current voice wake-up solution has high resource requirements, the present invention has conducted research. The initial idea is to directly predict whether each frame of audio signal is woken up using an end-to-end scheme, that is, the prediction result for each frame is 0 or 1, where 0 means not woken up and 1 means woken up; thus, the triphone state of each frame of audio signal is no longer predicted. However, since the end-to-end scheme is a holistic modeling and mainly provides a wake-up search interval, it cannot accurately know the pronunciation state corresponding to each frame like at the frame level. Therefore, the voice wake-up detection accuracy of the end-to-end scheme is not high.
[0028] In view of the problems existing in the above ideas, the present invention continued its research. During the research process, it occurred to the inventors that it is necessary to determine whether each audio data frame belongs to the wake-up state. When in the wake-up state, triphone state labels are used for model training, and when not in the wake-up state, monophone state labels are used for model training. That is, the labeled data of the sample audio data frames is either triphone state labels or monophone state labels. However, this method targets audio data frames, with too large a granularity, resulting in low accuracy of voice wake-up detection.
[0029] In view of the above defects, the present invention conducted further research. During the research process, it occurred to the inventors that first, the triphone state labels of the sample audio data frames are determined, and then the triphone states corresponding to the non-wake-up state in the triphone state labels are changed to monophone states to obtain hybrid phoneme state labels, so as to directly perform model sequences based on the hybrid phoneme state labels. However, since the hybrid phoneme state labels include monophone states, the granularity is still relatively coarse, resulting in low accuracy of voice wake-up detection.
[0030] In view of the above defects, the present invention further conducts research and finally proposes a voice wake-up method. The voice wake-up method first performs model training based on the second sample audio data frames and their corresponding triphone state labels to obtain a trained model, and then optimizes the classification layer in the trained model based on the first sample audio data frames and their corresponding hybrid phoneme state labels to obtain an acoustic model, so as to perform model training based on the sample audio data frames with all triphone state labels, improve the feature extraction performance of the acoustic model. Then, the classification layer in the trained model is optimized based on the hybrid phoneme state labels to obtain an acoustic model, thereby reducing the number of states output by the acoustic model and further reducing the resources required by the acoustic model.
[0031] Next, the voice wake-up method provided by the present invention will be introduced through the following embodiments. The following combines Figures 1-6 Describe the voice wake-up method of the present invention.
[0032] The execution subject of the voice wake-up method provided by the present invention can be an electronic device or a chip (such as a low-power chip), and the electronic device can be a voice recognition device such as a robot, a mobile phone, a smart home device, a car machine, a learning machine, and a wearable device.
[0033] Figure 1 is one of the flow diagrams of the voice wake-up method provided by the present invention. As Figure 1 shown, the voice wake-up method includes the following steps 110 and 120.
[0034] Step 110, input each audio data frame in the audio data into the acoustic model to obtain a phoneme-level state sequence output by the acoustic model.
[0035] Here, the audio data is the audio data to be detected for triggering the voice wake-up state, that is, to determine whether the corresponding text data includes a preset wake-up word. The text data corresponding to this audio data may or may not include the preset wake-up word. The preset wake-up word is a predefined wake-up word, such as "Xiaofei, Xiaofei".
[0036] This audio data can be collected by a voice collection device. For example, the voice collection device is a microphone on a mobile phone. This audio data can only include user voice data, or can also include environmental noise data. Of course, it can also not include user voice data, that is, the audio data is not specifically limited.
[0037] Here, the audio data frame is the frame data split from the audio data. In one embodiment, the audio data is sampled based on a preset sampling rate to obtain each audio data frame; the preset sampling rate can be set according to actual needs. For example, it is 16 kHz.
[0038] Here, the acoustic model is used to predict the phoneme-level state classification results of each audio data frame in the audio data. Specifically, the classification layer in the acoustic model predicts the phoneme-level state classification results of each audio data frame in the audio data.
[0039] Among them, the acoustic model is obtained by optimizing the classification layer in the trained model based on the first sample audio data frame and the corresponding mixed phoneme state label of the first sample audio data frame; the trained model is trained based on the second sample audio data frame and the corresponding triphone state label of the second sample audio data frame; the mixed phoneme state label includes the triphone state corresponding to the wake-up state and the single phoneme state corresponding to the non-wake-up state.
[0040] In other words, the acoustic model is trained based on the following method: based on the second sample audio data frame and the corresponding triphone state label of the second sample audio data frame, the model is trained to obtain the trained model; based on the first sample audio data frame and the corresponding mixed phoneme state label of the first sample audio data frame, the classification layer in the trained model is optimized to obtain the acoustic model.
[0041] Here, the second sample audio data frame is any audio data frame in the sample audio data, and it can also be the frame data split from the sample audio data. Similarly, the first sample audio data frame can be any audio data frame in the sample audio data, and it can also be the frame data split from the sample audio data. Usually, there are multiple sample audio data, that is, a training data set is obtained for model training.
[0042] Specifically, based on multiple second-sample audio data frames and the corresponding triphone state labels for each second-sample audio data frame, model training can be carried out to obtain a trained model; based on multiple first-sample audio data frames and the corresponding hybrid phoneme state labels for each first-sample audio data frame, the classification layer in the trained model can be optimized to obtain an acoustic model.
[0043] Here, the triphone state label is the labeled triphone state. The triphone state is the state based on the triphone representation, that is, the triphone state can also be understood as a triphone. A triphone refers to the variant of a phoneme under the influence of adjacent phonemes before and after. Its structure is usually expressed as the left phoneme - the current phoneme + the right phoneme (such as B - A + C), indicating the pronunciation of the current phoneme A under the joint action of the left phoneme B and the right phoneme C. For the triphone state, it not only focuses on the acoustic features of the current audio data frame itself, but also fully considers the phoneme information of its front and back frames, so as to capture the subtle changes caused by co-articulation between phonemes. It should be understood that the triphone state label includes multiple triphone states. For example, it includes 9004 triphone states.
[0044] Here, the monophone state is the state based on the monophone representation, that is, the monophone state can also be understood as a monophone. For the monophone state, although it is relatively simpler, it is also necessary to accurately identify the characteristics of the independent pronunciation of each audio data frame.
[0045] It should be understood that the hybrid phoneme state label includes multiple triphone states corresponding to the wake-up state and multiple monophone states corresponding to the non-wake-up state, that is, it can be determined whether each phoneme is in the wake-up state or the non-wake-up state. Among them, the phoneme corresponding to the wake-up state refers to any phoneme included in the preset wake-up word, and the phoneme corresponding to the non-wake-up state refers to any other phoneme other than the phoneme corresponding to the wake-up state. Whether a triphone is a phoneme corresponding to the wake-up state can be based on the middle phoneme of the triphone. Based on this, the hybrid phoneme state label only needs to retain the triphone state when the corresponding phoneme is in the wake-up state. Thus, the hybrid phoneme state label is a more simplified form of phoneme combination, which helps to reduce the number of states output by the classification layer, thereby simplifying the model structure and improving the calculation efficiency, that is, reducing the resource requirements of the acoustic model.
[0046] Based on the above, the acoustic model is trained in two training stages. The second sample audio data frames and their corresponding triphone state labels are the training data for the first training stage. The first sample audio data frames and their corresponding hybrid phoneme state labels are the training data for the second training stage. And in the first training stage, training is performed using the triphone state labels, so as to utilize the fine-grained representation of triphones and improve the sensitivity of the acoustic model to subtle changes in speech. Compared with directly training using the hybrid phoneme state labels, the performance of the acoustic model can be improved, thereby improving the accuracy of voice wake-up detection. That is to say, in the first training stage, the model is trained to predict triphone states, and in the second training stage, the model parameters except the classification layer are fixed (frozen), and the model is trained to predict hybrid phoneme states, and only the classification layer is allowed to be trained according to the new hybrid phoneme state labels, ensuring that the acoustic model output in the inference stage includes phoneme-level state classification results of single phoneme states and triphone states.
[0047] It should be noted that in the second training stage, the model parameters except the classification layer are fixed (frozen). That is, after the acoustic model is trained, the branch for triphone state prediction is discarded, and only the branch for hybrid phoneme state prediction is retained, so as to ensure that the acoustic model in the inference stage only outputs hybrid phoneme states (phoneme-level state classification results).
[0048] It should be understood that fixing (freezing) the model parameters except the classification layer in the second training stage, especially those network layers that have learned rich triphone information, can prevent the loss of these important low-level speech features during the fine-tuning process in the second training stage. The advantages of only optimizing the classification layer in the second training stage are as follows.
[0049] First, maintaining low-level features: By freezing the model parameters except the classification layer, the triphone-level details learned by the acoustic model in the first training stage are retained, thus ensuring that the acoustic model does not forget these key information.
[0050] Second, focusing on high-level tasks: Only updating the last classification layer can make the acoustic model more focused on mapping these low-level features to the final voice wake-up state determination task, improving the performance of the acoustic model in the voice wake-up detection task.
[0051] Third, improving training efficiency: Since there is no need to update a large number of parameters, the training speed is significantly accelerated, and at the same time, the demand for computing resources is reduced.
[0052] The acoustic model includes a classification layer, which is used to classify and obtain a phoneme-level state classification result. Further, the acoustic model includes a feature extraction layer and a classification layer connected in sequence. Still further, the acoustic model includes a feature extraction layer, a quantization layer, and a classification layer connected in sequence. The feature extraction layer is used to extract the input features. In one embodiment, the quantization layer is used to reshape or pad the features extracted by the feature extraction layer into a unified size, and normalize the features to accelerate the training process and improve the model performance.
[0053] Wherein, the phoneme-level state sequence includes the phoneme-level state classification results of each audio data frame in the audio data. That is, each audio data frame is respectively input into the acoustic model to obtain the phoneme-level state classification results of each audio data frame output by the acoustic model.
[0054] Wherein, the phoneme-level state classification results include single-phoneme states and triphone states. Since the acoustic model is optimized for the classification layer in the trained model based on the first sample audio data frame and the corresponding mixed phoneme state label of the first sample audio data frame, and the mixed phoneme state label includes the triphone state corresponding to the wake-up state and the single-phoneme state corresponding to the non-wake-up state, the phoneme-level state classification results output by the acoustic model include single-phoneme states and triphone states.
[0055] Exemplarily, the triphone state label includes M triphone states; the mixed phoneme state label includes N triphone states corresponding to the wake-up state and L single-phoneme states corresponding to the non-wake-up state, and N + L is less than M. It should be understood that the triphone state label usually includes more triphone states. For example, M = 9004, while only the wake-up state in the mixed phoneme state label retains the triphone state, and the non-wake-up state directly uses the single-phoneme state with fewer states. For example, N = 30 and L = 83, so N + L is much less than M. The reason why N is small is that the preset wake-up word includes fewer phonemes, and the reason why L is small is that the single-phoneme state itself is relatively few.
[0056] In one embodiment, the triphone state label is obtained by the following method: obtaining the second sample audio data frame in the sample audio data and the text data corresponding to the sample audio data; determining the triphone state label corresponding to the second sample audio data frame based on the forced alignment result of the sample audio data and the text data.
[0057] In one embodiment, the mixed phoneme state label is obtained by changing the triphone states corresponding to non-awakening states in the pre-labeled triphone state labels to monophone states. For example, the mixed phoneme state label is obtained in the following way: obtain the second sample audio data frame and the corresponding triphone state label of the second sample audio data frame, change the triphone states corresponding to non-awakening states in the triphone state label to monophone states, and obtain the mixed phoneme state label.
[0058] Step 120: Determine whether it is a voice wake-up state based on the phoneme-level state sequence.
[0059] It should be noted that if the text data corresponding to the audio data includes a preset wake-up word, it is determined to be a voice wake-up state; if the text data corresponding to the audio data does not include a preset wake-up word, it is determined not to be a voice wake-up state. Of course, it is not directly determined whether the text data includes a preset wake-up word, but the phoneme-level state sequence is analyzed to determine whether it is a voice wake-up state.
[0060] In the voice wake-up state, the user's voice is recognized to give a specified response. Exemplarily, the voice wake-up method provided by the embodiments of the present invention is applied to the front-end module of a speech recognition system to achieve real-time monitoring and recognition of the user's voice. When a preset wake-up word or preset phrase is detected, the subsequent speech recognition module or human-computer interaction module is quickly activated.
[0061] In a specific embodiment, first, the audio data is classified into phoneme-level states by using a pre-constructed acoustic model, and then the frame-level phoneme-level state classification result is input into a decoder to obtain the wake-up state determination result output by the decoder. The wake-up state determination result of the audio data is used to represent whether it is a voice wake-up state. More specifically, the classification probabilities of the phoneme-level state classification result are input into the decoder. For example, if the phoneme-level state classification result includes N + L phoneme states, the classification probabilities of the N + L phoneme states are input into the decoder.
[0062] In one embodiment, the decoder can use the Viterbi decoding algorithm for decoding, that is, the Viterbi decoding algorithm is used to decode the phoneme-level state sequence to determine whether it is a voice wake-up state. Of course, decoding can also be performed in other ways, which are not specifically limited here.
[0063] Among them, the Viterbi decoding algorithm will further analyze and process the predicted phoneme-level state sequence. It integrates these scattered phoneme-level state classification results into a wake-up state determination result with clear semantics according to the preset decoding rules and probability models. In this way, it can accurately identify whether the user has uttered a preset wake-up word, thereby realizing an efficient and accurate voice wake-up function.
[0064] It should be understood that through the acoustic model, a detailed frame-level analysis and prediction of the audio data is performed. That is, the phoneme-level state sequence includes the phoneme-level state classification results of each audio data frame in the audio data, so as to more accurately determine whether it is a voice wake-up state based on the phoneme-level state sequence, that is, to improve the detection accuracy of voice wake-up.
[0065] To facilitate the understanding of the inventive concept of the embodiments of the present invention, an example is given here. Currently, through the neural network model, the triphone state corresponding to each frame of audio signal is predicted, and the number of states of the output triphone state is 9004. If the number of channels of the classification layer is 64, then the size of the classification layer will be 9004 ×64 / 1024 = 562.75K. The parameters of the classification layer will be larger than those of the feature extraction model in the model and cannot run on low-cost devices or chips. Through the neural network model, the monophone state corresponding to each frame of audio signal is predicted, and the number of states of the output monophone state is 83. If the number of channels of the classification layer is 64, then the size of the classification layer is 83 ×64 / 1024 = 5.1875K. Although monophone modeling can greatly compress the model memory, its effect is poor. Based on this, in the embodiments of the present invention, the classification layer is optimized based on the mixed phoneme state label, so that the classification layer outputs the triphone states corresponding to N (for example, 45) wake-up states and the monophone states corresponding to 83 non-wake-up states. Based on this, if N = 45, that is, the number of states output by the classification layer is 128, the classification layer can be compressed from 562.75K to 8K, so that it can run on low-power devices and ensure a sufficient wake-up rate.
[0066] It should be understood that through the two-stage training strategy of the above acoustic model, while ensuring that the acoustic model has a powerful representation ability (guaranteed by the first training stage), an efficient and reliable phoneme-level state classification can be achieved. And by freezing most of the model parameters in the second training stage, not only the problem of forgetting triphone information is avoided, but also the resources required by the acoustic model are reduced, thereby enhancing the generalization ability of the acoustic model in actual application scenarios.
[0067] The voice wake-up method provided by the embodiment of the present invention inputs each audio data frame in the audio data into an acoustic model to obtain a phoneme-level state sequence output by the acoustic model. The acoustic model is obtained by optimizing the classification layer in the trained model based on the first sample audio data frame and the corresponding mixed phoneme state label of the first sample audio data frame. The mixed phoneme state label includes the triphone state corresponding to the wake-up state and the monophone state corresponding to the non-wake-up state. Therefore, only when the corresponding phoneme is in the wake-up state, it is necessary to label the triphone state, otherwise only the monophone state with fewer states needs to be labeled, thereby reducing the number of states output by the acoustic model, further reducing the resource requirements of the acoustic model, ensuring that the voice wake-up method can be applied to low-power devices, realizing a low-power voice wake-up method, and ultimately enhancing the application scope of the voice wake-up method. At the same time, the trained model is trained based on the second sample audio data frame and the corresponding triphone state label of the second sample audio data frame. Therefore, the triphone state label is first used for training in the first training stage to utilize the fine-grained representation of the triphone and improve the sensitivity of the acoustic model to subtle changes in speech. Then, in the second training stage, only the classification layer is optimized, which can prevent the loss of information learned in the first training stage during the optimization process of the second training stage, thereby ensuring that the classification layer outputs a smaller number of states while improving the performance of the acoustic model, and further improving the detection accuracy of voice wake-up.
[0068] Based on any of the above embodiments, another specific embodiment of the voice wake-up method is given below. This embodiment is used to introduce the training method of the acoustic model. The execution subject of the training method of the acoustic model may be the same as or different from the execution subject of the voice wake-up method. For example, the execution subject of the training method of the acoustic model is a training device.
[0069] Figure 2 It is the second flow diagram of the voice wake-up method provided by the present invention. As Figure 2 shown, the acoustic model is trained based on the following steps 210 and 220.
[0070] Step 210: Train an initial model based on the second sample audio data frame and the corresponding triphone state label of the second sample audio data frame to obtain a trained model.
[0071] Here, the second sample audio data frame is any audio data frame in the sample audio data. Usually, there are multiple sample audio data, that is, a training data set is obtained for model training.
[0072] In one embodiment, the sample audio data set of the sample audio data includes: general audio data and wake word audio data; the general audio data does not include the audio data corresponding to the preset wake word, and the wake word audio data includes the audio data corresponding to the preset wake word. That is, a large amount of general audio data that does not include the preset wake word and a large amount of wake word audio data that includes the preset wake word are obtained to be used as the sample audio data set for model training. Based on this, in addition to including the wake word audio data, it also includes general audio data, which helps the acoustic model better learn to distinguish background noise from the preset wake word, thereby improving the detection accuracy of voice wake-up.
[0073] Specifically, the initial model can be trained based on multiple second sample audio data frames and the corresponding triphone state labels of each second sample audio data frame to obtain the trained model.
[0074] Here, the initial model is the initial network model to be trained, and this initial model can be selected according to actual needs and is not specifically limited here. Of course, the initial model can also be a model that has been preliminarily trained, so as to reduce the training amount and improve the model performance.
[0075] Here, the trained model can predict the triphone state corresponding to the audio data frame, that is to say, the first training stage trains the model to predict the triphone state.
[0076] Specifically, the first training stage is to adjust the model parameters of the entire initial model, so that in the first training stage, training is carried out using the triphone state labels, thereby using the fine-grained representation of the triphone to improve the sensitivity of the acoustic model to the subtle changes in speech, and further improving the performance of the acoustic model, thereby improving the detection accuracy of voice wake-up.
[0077] In one embodiment, the triphone state label is obtained in the following manner: obtaining the second sample audio data frame in the sample audio data and the text data corresponding to the sample audio data; determining the triphone state label corresponding to the second sample audio data frame based on the forced alignment result of the sample audio data and the text data.
[0078] Step 220, optimizing the classification layer in the trained model based on the first sample audio data frame and the corresponding hybrid phoneme state label of the first sample audio data frame to obtain the acoustic model.
[0079] Here, the first sample audio data frame can also be any audio data frame in the sample audio data.
[0080] Specifically, the classification layer in the trained model can be optimized based on multiple first sample audio data frames and the corresponding hybrid phoneme state labels of each first sample audio data frame to obtain the acoustic model.
[0081] It should be understood that the mixed phoneme state label includes the triphone states corresponding to multiple wake-up states and the monophone states corresponding to multiple non-wake-up states, that is, each phoneme can determine whether it is in a wake-up state or a non-wake-up state. Based on this, the mixed phoneme state label only needs to retain the triphone state when the corresponding phoneme is in a wake-up state. Thus, the mixed phoneme state label is a more simplified form of phoneme combination, which helps to reduce the number of states output by the classification layer, thereby reducing the resource requirements of the acoustic model.
[0082] Since the model parameters except for the classification layer are fixed (frozen) in the second training stage, and the model is trained to predict the mixed phoneme state, only allowing the classification layer to be trained according to the new mixed phoneme state label, the acoustic model can only output the phoneme-level state classification results including monophone states and triphone states. That is, after the acoustic model is trained, the branch for triphone state prediction is discarded, and only the branch for mixed phoneme state prediction is retained, so as to ensure that the acoustic model in the inference stage only outputs the mixed phoneme state (phoneme-level state classification result).
[0083] It should be understood that fixing (freezing) the model parameters except for the classification layer in the second training stage, especially those network layers that have learned rich triphone information, can prevent the loss of these important low-level speech features during the fine-tuning process in the second training stage.
[0084] In one embodiment, the mixed phoneme state label is obtained by changing the triphone states corresponding to non-wake-up states in the pre-annotated triphone state label to monophone states. For example, the mixed phoneme state label is obtained in the following way: obtaining the second sample audio data frame and the corresponding triphone state label of the second sample audio data frame, and changing the triphone states corresponding to non-wake-up states in the triphone state label to monophone states to obtain the mixed phoneme state label.
[0085] The voice wake-up method provided by the embodiment of the present invention trains an initial model based on the second sample audio data frame and the corresponding triphone state label of the second sample audio data frame to obtain a trained model. Thus, the first training stage is first trained using the triphone state label to utilize the fine-grained representation of the triphone and improve the sensitivity of the acoustic model to subtle changes in speech. However, based on the first sample audio data frame and the corresponding mixed phoneme state label of the first sample audio data frame, the classification layer in the trained model is optimized to obtain an acoustic model. That is, only the classification layer is optimized in the second training stage, which can prevent the loss of information learned in the first training stage during the optimization process of the second training stage, thereby ensuring that the classification layer outputs a smaller number of states while improving the performance of the acoustic model, and further improving the detection accuracy of voice wake-up. At the same time, only the parameters of the entire model need to be adjusted in the first training stage, and only the parameters of the classification layer need to be adjusted in the second training stage, thereby reducing the training calculation amount and improving the training efficiency.
[0086] Based on any of the above embodiments, another specific embodiment of the voice wake-up method is given below. Figure 3 It is the third schematic flowchart of the voice wake-up method provided by the present invention, as Figure 3 shown. The acoustic model is trained based on the following steps 230, 240, 210, and 220.
[0087] Step 230: Obtain the second sample audio data frame and the corresponding triphone state label of the second sample audio data frame.
[0088] In a specific embodiment, the second sample audio data frame in the sample audio data and the corresponding text data of the sample audio data are obtained; based on the forced alignment result of the sample audio data and the text data, the triphone state label corresponding to the second sample audio data frame is determined.
[0089] Among them, the triphone state label includes M triphone states. Among them, M is greater than 1. For example, the triphone state label includes 9004 triphone states.
[0090] Step 240: Change the triphone states corresponding to the non-wake-up states in the triphone state label to monophone states to obtain the mixed phoneme state label.
[0091] Here, the triphone state label is the labeled triphone state, and the triphone state is the state based on the triphone representation, that is, the triphone state can also be understood as the triphone. A triphone refers to a variant of a phoneme under the influence of adjacent phonemes before and after. Its structure is usually expressed as the left phoneme - the current phoneme + the right phoneme (such as B - A + C), indicating the pronunciation of the current phoneme A under the joint action of the left phoneme B and the right phoneme C. Therefore, it is judged whether it is a wake-up state based on the current phoneme A in the triphone.
[0092] It is understandable that some of the phoneme states in the triphone state labels are non-awakening states. Therefore, the triphone states corresponding to these non-awakening states can be changed to monophone states. That is, the triphone states in the non-awakening states are collapsed into monophone states, and the phonemes corresponding to the awakening states retain their original triphone states.
[0093] In a specific embodiment, the middle phoneme (the current phoneme) in the triphone state corresponding to the non-awakening state is retained as the changed monophone state. For example, the triphone structure is represented as the left phoneme - the current phoneme + the right phoneme (such as B - A + C), indicating the pronunciation of the current phoneme A under the joint action of the left phoneme B and the right phoneme C. Therefore, the changed monophone is A.
[0094] Wherein, the first sample audio data frame is the same as the second sample audio data frame.
[0095] Wherein, the mixed phoneme state label includes N triphone states corresponding to the awakening states and L monophone states corresponding to the non-awakening states; N + L is less than M.
[0096] Exemplarily, the triphone state label usually includes a relatively large number of triphone states. For example, M = 9004, while in the mixed phoneme state label, only the awakening states retain the triphone states, and the non-awakening states directly use the monophone states with a smaller number of states. For example, N = 30 and L = 83, so that N + L is much less than M, and then the states are re-sorted to obtain the mixed phoneme state label with N + L state numbers. The reason why N is small is that the preset wake-up word includes relatively few phonemes, and the reason why L is small is that the monophone state itself is relatively few.
[0097] Step 210, based on the second sample audio data frame and the triphone state label corresponding to the second sample audio data frame, train the initial model to obtain a trained model.
[0098] Step 220, based on the first sample audio data frame and the mixed phoneme state label corresponding to the first sample audio data frame, optimize the classification layer in the trained model to obtain the acoustic model.
[0099] The voice wake-up method provided by the embodiment of the present invention first obtains the second sample audio data frame and the corresponding triphone state label of the second sample audio data frame, and then changes the triphone state corresponding to the non-wake-up state in the triphone state label to a monophone state, so as to obtain a mixed phoneme state label without re-labeling the mixed phoneme state label, thereby reducing the label annotation cost; and the number of states of the mixed phoneme state label is less than the number of states of the triphone state label, thereby reducing the number of states output by the acoustic model, and further reducing the resource requirements of the acoustic model, ensuring that the voice wake-up method can be applied to low-power devices, realizing a low-power voice wake-up method, and finally enhancing the application scope of the voice wake-up method.
[0100] Based on any of the above embodiments, another specific embodiment of the voice wake-up method is given below. Figure 4 It is the fourth flowchart of the voice wake-up method provided by the present invention, as Figure 4 shown, the acoustic model is trained based on the following steps 231, 232, 240, 210 and 220.
[0101] Step 231, obtain the second sample audio data frame in the sample audio data, and the text data corresponding to the sample audio data.
[0102] Here, the sample audio data can be split into multiple second sample audio data frames. The sample audio data is usually multiple, that is, a training data set is obtained for model training.
[0103] Among them, the sample audio data set including the sample audio data includes: general audio data and wake-up word audio data; the general audio data does not include the audio data corresponding to the preset wake-up word, and the wake-up word audio data includes the audio data corresponding to the preset wake-up word.
[0104] Here, the sample audio data set includes multiple sample audio data. The preset wake-up word is a pre-set wake-up word, such as "Xiaofei, Xiaofei".
[0105] It should be understood that the general audio data not including the preset wake-up word and the wake-up word audio data including the preset wake-up word are obtained as the sample audio data set for model training. Based on this, in addition to including the wake-up word audio data, it also includes general audio data, which helps the acoustic model better learn to distinguish background noise from the preset wake-up word, thereby improving the detection accuracy of voice wake-up.
[0106] In one embodiment, the general audio data is obtained by collecting a large amount of daily environmental audio, and of course, this audio should not include the preset wake-up word. Further, the acquisition sources of the general audio data may include, but are not limited to: daily conversations, background music, environmental noises (such as wind sounds, traffic noises, etc.), ensuring that the general audio data is as diverse as possible to improve the robustness of the acoustic model in the inference stage, thereby improving the detection accuracy of voice wake-up.
[0107] In one embodiment, the wake-up word audio data is obtained by collecting audio including the preset wake-up word. To further enhance the sample data, audio of different speakers can be collected, or audio of different speaking speeds can be collected, or audio of different volumes can be collected, or audio under different background noise conditions can be collected. Of course, all audio under real-world conditions should be collected as much as possible; based on this, an acoustic model that can accurately recognize the wake-up word under various conditions is trained, that is, the detection accuracy of voice wake-up is improved.
[0108] In one embodiment, the wake-up word audio data can be used to record the voices of different people with a recording device, or use text-to-speech technology to generate diverse wake-up word audio data.
[0109] Further, the wake-up word audio data is high-fidelity audio data, thereby ensuring the robustness of the acoustic model.
[0110] In one embodiment, the sample audio data can be used to record different audio with a recording device.
[0111] Here, the text data can be the transcribed text of the sample audio data.
[0112] Step 232, based on the forced alignment result of the sample audio data and the text data, determine the triphone state label corresponding to the second sample audio data frame.
[0113] Here, the forced alignment result can be understood as the correspondence between each sample audio data frame in the sample audio data and each word in the text data, that is, the forced alignment result provides a direct mapping from the audio data to its corresponding phoneme state label sequence. Specifically, the forced alignment result is obtained by performing forced alignment on the sample audio data and the text data. Forced Alignment (FA) is a technique used to precisely match audio data with the transcribed text, thereby assigning corresponding triphone state labels to each frame of the sample audio data.
[0114] Of course, based on the forced alignment result of the sample audio data and the text data, the mixed phoneme state label corresponding to the first sample audio data frame can also be determined.
[0115] Step 240: Change the triphone states corresponding to the non-awakening states in the triphone state labels to monophone states to obtain the mixed phoneme state labels.
[0116] Step 210: Train the initial model based on the second sample audio data frames and the triphone state labels corresponding to the second sample audio data frames to obtain a trained model.
[0117] Step 220: Optimize the classification layer in the trained model based on the first sample audio data frames and the mixed phoneme state labels corresponding to the first sample audio data frames to obtain the acoustic model.
[0118] The voice wake-up method provided by the embodiments of the present invention obtains the second sample audio data frames in the sample audio data and the text data corresponding to the sample audio data, so as to accurately determine the triphone state labels corresponding to the second sample audio data frames based on the forced alignment results of the sample audio data and the text data, thereby improving the training effect of the acoustic model and ultimately improving the detection accuracy of voice wake-up. At the same time, the sample audio data set containing the sample audio data includes general audio data and wake-up word audio data, and the general audio data does not include the audio data corresponding to the preset wake-up word, and the wake-up word audio data includes the audio data corresponding to the preset wake-up word, thereby helping the acoustic model to better learn to distinguish background noise from the preset wake-up word, and further improving the detection accuracy of voice wake-up.
[0119] Based on any of the above embodiments, another specific embodiment of the voice wake-up method is given below. Figure 5 is the fifth flow chart of the voice wake-up method provided by the present invention, as Figure 5 shown, the acoustic model is trained based on the following steps 211 and 220.
[0120] Step 211: Train the initial model based on the second sample audio data frames, the triphone state labels corresponding to the second sample audio data frames, and the binary classification result labels corresponding to the second sample audio data frames to obtain a trained model.
[0121] Here, the second sample audio data frame is any audio data frame in the sample audio data. There are usually multiple sample audio data, that is, a training data set is obtained for model training.
[0122] Among them, the binary classification result labels include the first classification result or the second classification result. The first classification result is used to represent that the second sample audio data frame belongs to the audio data frame corresponding to the end position of the wake-up word, and the second classification result is used to represent that the second sample audio data frame does not belong to the audio data frame corresponding to the end position of the wake-up word.
[0123] This binary classification is a boundary classification. Considering that the triphone state labels only focus on local information, the binary classification result labels are used to focus on the overall modeling information, so as to better train the feature extraction ability of the acoustic model. That is, the boundary loss function of this boundary classification is used to particularly focus on the end position of the preset wake-up word, that is, to better determine the exact boundary of the preset wake-up word, thereby improving the detection accuracy of voice wake-up.
[0124] Exemplarily, in order to optimize both the local and overall modeling capabilities of the acoustic model, an initial model is trained through a boundary loss function and a triphone state prediction loss function.
[0125] In a specific embodiment, a comprehensive loss is determined through a boundary loss function and a triphone state prediction loss function, and the initial model is trained based on the comprehensive loss.
[0126] In an embodiment, the formula for determining the comprehensive loss is as follows: ; In the formula, represents the time frame, represents the th frame of the comprehensive loss of the second sample audio data frame, represents the triphone state label corresponding to the th frame of the second sample audio data frame, represents the triphone state prediction result corresponding to the th frame of the second sample audio data frame, represents the triphone state prediction loss; is a weight factor used to balance the importance of the two losses; represents the binary classification result label of the th frame of the second sample audio data frame, represents the binary classification prediction result of the th frame of the second sample audio data frame, represents the boundary loss.
[0127] In an embodiment, the triphone state prediction loss function is a cross-entropy loss function. Exemplarily, the triphone state prediction loss function is as follows: ; In the formula, represents the triphone category. For example, if the triphone state label includes 9004 triphones, then the triphone category has a total of 9004; represents the th triphone in the triphone state label corresponding to the th frame of the second sample audio data frame; represents the The th triphone in the triphone state prediction result corresponding to the second sample audio data frame.
[0128] In one embodiment, the boundary loss function is a binary classification loss function. Exemplarily, the boundary loss function is as follows: ; In the formula, represents the binary classification result label of the th second sample audio data frame, represents the binary classification prediction result of the th second sample audio data frame. Further, assume that represents that the second sample audio data frame belongs to the audio data frame corresponding to the end position of the wake word, represents that the second sample audio data frame does not belong to the audio data frame corresponding to the end position of the wake word.
[0129] Specifically, the initial model can be trained based on multiple second sample audio data frames, the triphone state labels corresponding to each second sample audio data frame, and the binary classification result labels corresponding to each second sample audio data frame to obtain a trained model.
[0130] Here, the trained model can predict the triphone state corresponding to the audio data frame, that is, the first training stage trains the model to predict the triphone state. It should be noted that after the acoustic model training is completed, the branch of boundary classification is discarded, and only the branch of mixed phoneme state prediction is retained, so as to ensure that the acoustic model in the inference stage only outputs the mixed phoneme state (phoneme-level state classification result).
[0131] Specifically, the first training stage is to adjust the model parameters of the entire initial model. In the first training stage, the triphone state labels and binary classification result labels are used for training, so as to utilize the fine-grained representation of triphones to improve the sensitivity of the acoustic model to subtle changes in speech, and then improve the performance of the acoustic model, thereby improving the accuracy of speech wake-up detection; and utilize the binary classification result labels to focus on the overall modeling information and assist in training the feature extraction ability of the acoustic model.
[0132] In other words, in the first training stage, all model parameters will be updated, which means that all layers of the network are participating in the learning process. Doing so can ensure that the model can fully learn the complex structures and patterns in the audio data and can effectively capture the key features of the preset wake word, thereby improving the detection accuracy of speech wake-up.
[0133] Step 220: Optimize the classification layer in the trained model based on the first sample audio data frame and the corresponding hybrid phoneme state label to obtain the acoustic model.
[0134] Here, the first sample audio data frame can also be any audio data frame in the sample audio data.
[0135] Specifically, the classification layer in the trained model can be optimized based on multiple first sample audio data frames and the corresponding hybrid phoneme state labels of each first sample audio data frame to obtain the acoustic model.
[0136] Since the model parameters except the classification layer are fixed (frozen) in the second training stage, the trained model predicts the hybrid phoneme state, and only the classification layer is allowed to be trained according to the new hybrid phoneme state label. Therefore, the acoustic model can only output the phoneme-level state classification result including the single phoneme state and the triphone state. That is, after the acoustic model is trained, the branches for triphone state prediction and boundary classification are discarded, and only the branch for hybrid phoneme state prediction is retained, so as to ensure that the acoustic model in the inference stage only outputs the hybrid phoneme state (phoneme-level state classification result).
[0137] In a specific embodiment, the classification layer in the trained model is optimized through the hybrid phoneme state prediction loss function.
[0138] In an embodiment, the hybrid phoneme state prediction loss function is the cross-entropy loss function. Exemplarily, the hybrid phoneme state prediction loss function is as follows: ; In the formula, represents the hybrid phoneme category; represents the th phoneme in the hybrid phoneme state label corresponding to the th frame of the first sample audio data frame; represents the th phoneme in the predicted result of the hybrid phoneme state corresponding to the th frame of the first sample audio data frame.
[0139] It should be understood that the embodiments of the present invention combine multi-level acoustic feature analysis with a fine classification strategy, improve the performance of the acoustic model, and thus improve the detection accuracy of voice wake-up.
[0140] The voice wake-up method provided by the embodiment of the present invention trains an initial model based on the second sample audio data frame, the three-phoneme state label corresponding to the second sample audio data frame, and the binary classification result label corresponding to the second sample audio data frame, so as to not only focus on local information through the three-phoneme state label, but also focus on overall modeling information through the binary classification result label, thereby better training the feature extraction ability of the acoustic model, and further improving the detection accuracy of voice wake-up.
[0141] Based on any of the above embodiments, another specific embodiment of the voice wake-up method is given below. Figure 6 It is the sixth flowchart of the voice wake-up method provided by the present invention, as Figure 6 shown, the voice wake-up method includes: step 111, step 112 and step 120.
[0142] In step 111, the acoustic features of each audio data frame in the audio data are input into the feature extraction layer in the acoustic model to obtain a sequence of feature vectors output by the feature extraction layer.
[0143] Here, the acoustic features may include but are not limited to: FBANK (Filter Bank Energies) features, mel cepstral coefficient features, intonation features, audio strength features, and so on.
[0144] In one embodiment, the acoustic feature is the fb40 feature, that is, a filter bank composed of 40 filters is used to capture different frequency components of the audio signal.
[0145] Here, the audio data frame is the frame data split from the audio data. In one embodiment, the audio data is sampled based on a preset sampling rate to obtain each audio data frame; the preset sampling rate can be set according to actual needs, for example, it is 16 kHz.
[0146] Exemplarily, the audio data is sampled based on a preset sampling rate to obtain each audio data frame; the short-time Fourier transform (STFT) is performed on each audio data frame to obtain a spectrogram; a set of triangular overlapping band-pass filters is applied to the spectrogram to obtain the filter bank energy corresponding to each time frame; the energy value of the filter bank energy corresponding to each time frame is calculated to obtain the acoustic feature.
[0147] Here, the feature extraction layer is used to extract the feature vectors of the acoustic features. The sequence of feature vectors includes the feature vectors of each acoustic feature.
[0148] In step 112, the sequence of feature vectors is input into the classification layer in the acoustic model to obtain a phoneme-level state sequence output by the classification layer in the acoustic model.
[0149] Here, the classification layer is used to predict a phoneme-level state sequence based on the feature vectors of each acoustic feature.
[0150] Wherein, the phoneme-level state sequence includes the phoneme-level state classification results of each audio data frame in the audio data. That is, each audio data frame is respectively input into the acoustic model to obtain the phoneme-level state classification results of each audio data frame output by the acoustic model.
[0151] In the voice wake-up method provided by the embodiment of the present invention, the acoustic features of each audio data frame in the audio data are input into the feature extraction layer in the acoustic model to obtain a sequence of feature vectors output by the feature extraction layer, so as to input the sequence of feature vectors into the classification layer in the acoustic model to obtain the phoneme-level state sequence output by the classification layer in the acoustic model, thereby first extracting acoustic features and then extracting their feature vectors, so as to improve the classification accuracy of the classification layer and further improve the detection accuracy of voice wake-up.
[0152] Through the above embodiments, the present invention realizes a voice wake-up method with extremely low power consumption. It mainly collapses triphones into hybrid phonemes instead of single phonemes, and predicts the distribution state of the hybrid phonemes and sends them to the decoder to obtain the final wake-up state. At the same time, a two-stage training method is introduced to improve the robustness of the voice wake-up method. Specifically, the two-stage training method not only ensures that the model can accurately recognize specific wake-up words, but also enhances its flexibility and adaptability when facing various wake-up words. This method significantly improves the overall performance and robustness of the system by finely adjusting the model parameters and the design of the loss function. That is, using hybrid phoneme state labels for modeling greatly reduces the model size while ensuring the effect, enabling the voice wake-up method to run on low-computing-power devices.
[0153] The voice wake-up device provided by the present invention will be described below. The voice wake-up device described below can be correspondingly referred to the voice wake-up method described above.
[0154] Figure 7 is a schematic structural diagram of the voice wake-up device provided by the present invention, as Figure 7 shown, the voice wake-up device includes a phoneme output module 710 and a state determination module 720.
[0155] The phoneme output module 710 is configured to input each audio data frame in the audio data into the acoustic model to obtain the phoneme-level state sequence output by the acoustic model; the phoneme-level state sequence includes the phoneme-level state classification results of each audio data frame in the audio data, and the phoneme-level state classification results include single-phoneme states and triphone states.
[0156] A state determination module 720, configured to determine whether it is a voice wake-up state based on the phoneme-level state sequence.
[0157] Wherein, the acoustic model is obtained by optimizing the classification layer in the trained model based on the first sample audio data frame and the mixed phoneme state label corresponding to the first sample audio data frame; the trained model is trained based on the second sample audio data frame and the triphone state label corresponding to the second sample audio data frame; the mixed phoneme state label includes the triphone state corresponding to the wake-up state and the single phoneme state corresponding to the non-wake-up state.
[0158] The voice wake-up device provided in the embodiment of the present invention inputs each audio data frame in the audio data into the acoustic model to obtain a phoneme-level state sequence output by the acoustic model. The acoustic model is obtained by optimizing the classification layer in the trained model based on the first sample audio data frame and the mixed phoneme state label corresponding to the first sample audio data frame, and the mixed phoneme state label includes the triphone state corresponding to the wake-up state and the single phoneme state corresponding to the non-wake-up state. Therefore, only when the corresponding phoneme is in the wake-up state, it is necessary to label the triphone state, otherwise only the single phoneme state with fewer states needs to be labeled, thereby reducing the number of states output by the acoustic model, further reducing the resource requirements of the acoustic model, ensuring that the voice wake-up method can be applied to low-power devices, realizing a low-power voice wake-up method, and ultimately enhancing the application range of the voice wake-up method; at the same time, the trained model is trained based on the second sample audio data frame and the triphone state label corresponding to the second sample audio data frame. Therefore, the triphone state label is first used for training in the first training stage to utilize the fine-grained representation of the triphone and improve the sensitivity of the acoustic model to subtle changes in speech. Then, only the classification layer is optimized in the second training stage, which can prevent the loss of information learned in the first training stage during the optimization process of the second training stage, thereby ensuring that the classification layer outputs fewer states while improving the performance of the acoustic model, and further improving the detection accuracy of voice wake-up.
[0159] Based on any of the above embodiments, the acoustic model is obtained by training based on the following model training module, and the model training module is used for: Training an initial model based on the second sample audio data frame and the triphone state label corresponding to the second sample audio data frame to obtain a trained model; Optimizing the classification layer in the trained model based on the first sample audio data frame and the mixed phoneme state label corresponding to the first sample audio data frame to obtain the acoustic model.
[0160] Based on any of the above embodiments, the model training module is further used for: Obtain the second sample audio data frame and the corresponding triphone state label of the second sample audio data frame; the triphone state label includes M triphone states; Change the triphone states corresponding to the non-awakening states in the triphone state label to monophone states to obtain the mixed phoneme state label; the first sample audio data frame is the same as the second sample audio data frame; Wherein, the mixed phoneme state label includes N triphone states corresponding to awakening states and L monophone states corresponding to non-awakening states; N + L is less than M.
[0161] Based on any of the above embodiments, the model training module is further configured to: Obtain the second sample audio data frame in the sample audio data and the text data corresponding to the sample audio data; Based on the forced alignment result of the sample audio data and the text data, determine the triphone state label corresponding to the second sample audio data frame; Wherein, the sample audio data set containing the sample audio data includes: general audio data and wake-up word audio data; the general audio data does not include audio data corresponding to a preset wake-up word, and the wake-up word audio data includes audio data corresponding to a preset wake-up word.
[0162] Based on any of the above embodiments, the model training module is further configured to: Based on the second sample audio data frame, the corresponding triphone state label of the second sample audio data frame, and the binary classification result label corresponding to the second sample audio data frame, train the initial model to obtain a trained model; Wherein, the binary classification result label includes a first classification result or a second classification result, the first classification result is used to represent that the second sample audio data frame belongs to the audio data frame corresponding to the end position of the wake-up word, and the second classification result is used to represent that the second sample audio data frame does not belong to the audio data frame corresponding to the end position of the wake-up word.
[0163] Based on any of the above embodiments, the phoneme output module 710 is further configured to: Based on the second sample audio data frame, the corresponding triphone state label of the second sample audio data frame, and the binary classification result label corresponding to the second sample audio data frame, train the initial model to obtain a trained model; Wherein, the binary classification result label includes a first classification result or a second classification result, the first classification result is used to represent that the second sample audio data frame belongs to the audio data frame corresponding to the end position of the wake-up word, and the second classification result is used to represent that the second sample audio data frame does not belong to the audio data frame corresponding to the end position of the wake-up word.
[0164] Figure 8 An entity structure schematic diagram of an electronic device is exemplified, as Figure 8 shown. The electronic device may include: a processor 810, a communications interface 820, a memory 830, and a communication bus 840. Among them, the processor 810, the communications interface 820, and the memory 830 complete mutual communication through the communication bus 840. The processor 810 may call logical instructions in the memory 830 to execute a voice wake-up method, which includes: inputting each audio data frame in the audio data into an acoustic model to obtain a phoneme-level state sequence output by the acoustic model; the phoneme-level state sequence includes phoneme-level state classification results of each audio data frame in the audio data, and the phoneme-level state classification results include single-phoneme states and tri-phoneme states; determining whether it is a voice wake-up state based on the phoneme-level state sequence; where the acoustic model is obtained by optimizing a classification layer in a pre-trained model based on first sample audio data frames and mixed phoneme state labels corresponding to the first sample audio data frames; the pre-trained model is trained based on second sample audio data frames and tri-phoneme state labels corresponding to the second sample audio data frames; the mixed phoneme state labels include tri-phoneme states corresponding to the wake-up state and single-phoneme states corresponding to the non-wake-up state.
[0165] In addition, when the logical instructions in the above-mentioned memory 830 can be implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0166] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the voice wake-up method provided by the above-mentioned various methods. The method includes: inputting each audio data frame in the audio data into an acoustic model to obtain a phoneme-level state sequence output by the acoustic model; the phoneme-level state sequence includes the phoneme-level state classification results of each audio data frame in the audio data, and the phoneme-level state classification results include single-phoneme states and triphone states; based on the phoneme-level state sequence, determine whether it is a voice wake-up state; wherein, the acoustic model is obtained by optimizing the classification layer in the trained model based on the first sample audio data frames and the corresponding mixed phoneme state labels of the first sample audio data frames; the trained model is trained based on the second sample audio data frames and the corresponding triphone state labels of the second sample audio data frames; the mixed phoneme state labels include triphone states corresponding to the wake-up state and single-phoneme states corresponding to the non-wake-up state.
[0167] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is implemented to execute the voice wake-up method provided by the above-mentioned various methods. The method includes: inputting each audio data frame in the audio data into an acoustic model to obtain a phoneme-level state sequence output by the acoustic model; the phoneme-level state sequence includes the phoneme-level state classification results of each audio data frame in the audio data, and the phoneme-level state classification results include single-phoneme states and triphone states; based on the phoneme-level state sequence, determine whether it is a voice wake-up state; wherein, the acoustic model is obtained by optimizing the classification layer in the trained model based on the first sample audio data frames and the corresponding mixed phoneme state labels of the first sample audio data frames; the trained model is trained based on the second sample audio data frames and the corresponding triphone state labels of the second sample audio data frames; the mixed phoneme state labels include triphone states corresponding to the wake-up state and single-phoneme states corresponding to the non-wake-up state.
[0168] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative efforts.
[0169] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0170] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A voice wake-up method, characterized in that: include: Input each audio data frame in the audio data into the acoustic model to obtain a phoneme-level state sequence output by the acoustic model; the phoneme-level state sequence includes a phoneme-level state classification result of each audio data frame in the audio data, and the phoneme-level state classification result includes a monophone state and a triphone state; Based on the phoneme-level state sequence, determining whether it is a voice wake-up state; Among them, the acoustic model is obtained by optimizing the classification layer in the trained model based on the first sample audio data frame and the mixed phoneme state label corresponding to the first sample audio data frame; the trained model is trained based on the second sample audio data frame and the triphone state label corresponding to the second sample audio data frame; the mixed phoneme state label includes a triphone state corresponding to the awakening state, and a monophone state corresponding to the non-awakening state.
2. The voice wake-up method according to claim 1, characterized in that: The acoustic model is trained based on the following method: Based on the second sample audio data frame and the triphone state label corresponding to the second sample audio data frame, training the initial model to obtain a trained model; Based on the first sample audio data frame and the mixed phoneme state label corresponding to the first sample audio data frame, the classification layer in the trained model is optimized to obtain the acoustic model.
3. The voice wake-up method according to claim 2, characterized in that: Before training the initial model to obtain the trained model based on the second sample audio data frame and the triphone state label corresponding to the second sample audio data frame, the method further includes: Acquire the second sample audio data frame and a triphone state label corresponding to the second sample audio data frame; the triphone state label includes M triphone states; The triphone state corresponding to the non-awakening state in the triphone state label is changed to a monophone state to obtain the mixed phoneme state label; the first sample audio data frame is the same as the second sample audio data frame; The mixed phoneme state label includes three phoneme states corresponding to N awakening states and single phoneme states corresponding to L non-awakening states; N+L is less than M.
4. The voice wake-up method according to claim 3, characterized in that: The obtaining the second sample audio data frame and the triphone state label corresponding to the second sample audio data frame includes: Acquire a second sample audio data frame in the sample audio data, and text data corresponding to the sample audio data; Determining a triphone state label corresponding to the second sample audio data frame based on a forced alignment result between the sample audio data and the text data; Among them, the sample audio data set containing the sample audio data includes: general audio data and wake-up word audio data; the general audio data does not include audio data corresponding to a preset wake-up word, and the wake-up word audio data includes audio data corresponding to a preset wake-up word.
5. The voice wake-up method according to claim 2, characterized in that: The step of training the initial model based on the second sample audio data frame and the triphone state label corresponding to the second sample audio data frame to obtain the trained model includes: Based on the second sample audio data frame and the triphone state label corresponding to the second sample audio data frame, and the binary classification result label corresponding to the second sample audio data frame, training the initial model to obtain a trained model; Among them, the binary classification result label includes a first classification result or a second classification result, the first classification result is used to characterize that the second sample audio data frame belongs to the audio data frame corresponding to the end position of the wake-up word, and the second classification result is used to characterize that the second sample audio data frame does not belong to the audio data frame corresponding to the end position of the wake-up word.
6. The voice wake-up method according to claim 1, characterized in that: The step of inputting each audio data frame in the audio data into the acoustic model to obtain a phoneme-level state sequence output by the acoustic model comprises: Inputting the acoustic features of each audio data frame in the audio data into a feature extraction layer in the acoustic model to obtain a feature vector sequence output by the feature extraction layer; The feature vector sequence is input into the classification layer in the acoustic model to obtain the phoneme-level state sequence output by the classification layer in the acoustic model.
7. A voice wake-up device, characterized in that: include: A phoneme output module, used to input each audio data frame in the audio data into the acoustic model to obtain a phoneme-level state sequence output by the acoustic model; The phoneme-level state sequence includes a phoneme-level state classification result of each audio data frame in the audio data, and the phoneme-level state classification result includes a monophone state and a triphone state; A state determination module, used to determine whether it is a voice wake-up state based on the phoneme-level state sequence; Among them, the acoustic model is obtained by optimizing the classification layer in the trained model based on the first sample audio data frame and the mixed phoneme state label corresponding to the first sample audio data frame; the trained model is trained based on the second sample audio data frame and the triphone state label corresponding to the second sample audio data frame; the mixed phoneme state label includes a triphone state corresponding to the awakening state, and a monophone state corresponding to the non-awakening state.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the voice wake-up method according to any one of claims 1 to 6 is implemented.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the voice wake-up method according to any one of claims 1 to 6 is implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the voice wake-up method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Acoustic model training method and device, voice awakening method and device and electronic equipment
CN111128134A
Voice recognition method and device thereof
CN112259089A
Equipment awakening method and device and computer readable storage medium
CN116705015A
User-defined wake-up word recognition method and device of intelligent voice equipment and storage medium
CN118675503A
Wakeword detection using multi-word model
US11308939B1