Voice wake-up method, device, electronic device, storage medium, and program product
Through the mixed phoneme state label optimization classification layer and two-stage training strategy of the acoustic model, the problems of high resource demand and low detection accuracy of voice wake-up schemes on low-power devices are solved, and the effective application of low-power voice wake-up method is realized.
Patent Information
- Application Number
- CN202510593799.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-05-09
AI Technical Summary
The existing voice wake-up solution cannot be effectively applied on low-power devices, and there is a problem of high resource requirements and low detection accuracy.
The acoustic model is adopted to optimize the classification layer by mixing phoneme state labels, combined with a two-stage training strategy, and initial training is used to use the triphone state labels, and the classification layer is optimized to reduce the number of states to ensure that the voice wake-up method runs on low-power devices.
The low-power voice wake-up method is realized, which enhances the accuracy and application range of voice wake-up detection and reduces resource requirements.
Smart Images

Figure CN120126455B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of acoustic signal processing, and in particular, to a voice wake-up method, device, electronic device, storage medium, and program product. Background Art
[0002] Voice wake-up means that when a user speaks a voice audio including a wake-up word, an electronic device switches from a sleep state to a voice wake-up state, so as to recognize the user's voice in the voice wake-up state to give a specified response. Voice wake-up methods are widely used in various voice control products, such as robots, mobile phones, smart home devices, car machines, learning machines, and wearable devices, etc. Therefore, it is necessary to implement the voice wake-up function.
[0003] Currently, through a neural network model, the triphone state corresponding to each frame of audio signal is predicted, or the phone state corresponding to each frame of audio signal is predicted. However, when predicting the triphone state corresponding to each frame of audio signal, the number of output states is relatively large. Most application devices have limited computing resources and storage resources, and most application devices have requirements for low cost. Therefore, the triphone prediction scheme cannot be applied to low-power devices and has limitations. Although predicting the phone state can compress the memory of the neural network model, the distinguishability of the output phone state is worse than that of the triphone state, resulting in lower accuracy of voice wake-up detection. In summary, how to reduce the resources required by the voice wake-up scheme while ensuring accurate voice wake-up is an urgent problem to be solved currently. Summary of the Invention
[0004] The present invention provides a voice wake-up method, device, electronic device, storage medium, and program product to solve the defect that the voice wake-up scheme in the prior art cannot be applied to low-power devices, and to implement a low-power voice wake-up method.
[0005] The present invention provides a voice wake-up method, including:
[0006] Inputting each audio data frame in the audio data into an acoustic model to obtain a phone-level state sequence output by the acoustic model; the phone-level state sequence includes the phone-level state classification results of each audio data frame in the audio data, and the phone-level state classification results include phone states and triphone states;
[0007] Determining whether it is a voice wake-up state based on the phone-level state sequence;
[0008] Among them, the acoustic model is obtained by optimizing the classification layer in the trained model based on the first sample audio data frame and the corresponding hybrid phoneme state label of the first sample audio data frame; the trained model is trained based on the second sample audio data frame and the corresponding triphone state label of the second sample audio data frame; the hybrid phoneme state label includes the triphone state corresponding to the wake state and the monophone state corresponding to the non-wake state.
[0009] According to a voice wake-up method provided by the present invention, the acoustic model is trained based on the following method:
[0010] Based on the second sample audio data frame and the corresponding triphone state label of the second sample audio data frame, train the initial model to obtain a trained model;
[0011] Based on the first sample audio data frame and the corresponding hybrid phoneme state label of the first sample audio data frame, optimize the classification layer in the trained model to obtain the acoustic model.
[0012] According to a voice wake-up method provided by the present invention, before training the initial model based on the second sample audio data frame and the corresponding triphone state label of the second sample audio data frame to obtain a trained model, it further includes:
[0013] Obtain the second sample audio data frame and the corresponding triphone state label of the second sample audio data frame; the triphone state label includes M triphone states;
[0014] Change the triphone state corresponding to the non-wake state in the triphone state label to a monophone state to obtain the hybrid phoneme state label; the first sample audio data frame is the same as the second sample audio data frame;
[0015] Among them, the hybrid phoneme state label includes N triphone states corresponding to the wake state and L monophone states corresponding to the non-wake state; N + L is less than M.
[0016] According to a voice wake-up method provided by the present invention, the obtaining of the second sample audio data frame and the corresponding triphone state label of the second sample audio data frame includes:
[0017] Obtain the second sample audio data frame in the sample audio data and the corresponding text data of the sample audio data;
[0018] Based on the forced alignment result of the sample audio data and the text data, determine the triphone state label corresponding to the second sample audio data frame;
[0019] Among them, the sample audio data set containing the sample audio data includes: general audio data and wake-up word audio data; the general audio data does not include the audio data corresponding to the preset wake-up word, and the wake-up word audio data includes the audio data corresponding to the preset wake-up word.
[0020] According to a voice wake-up method provided by the present invention, training the initial model based on the second sample audio data frame and the corresponding triphone state label of the second sample audio data frame to obtain a trained model includes:
[0021] Training the initial model based on the second sample audio data frame, the corresponding triphone state label of the second sample audio data frame, and the binary classification result label corresponding to the second sample audio data frame to obtain a trained model;
[0022] Among them, the binary classification result label includes a first classification result or a second classification result. The first classification result is used to represent that the second sample audio data frame belongs to the audio data frame corresponding to the end position of the wake-up word, and the second classification result is used to represent that the second sample audio data frame does not belong to the audio data frame corresponding to the end position of the wake-up word.
[0023] According to a voice wake-up method provided by the present invention, inputting each audio data frame in the audio data into an acoustic model to obtain a phoneme-level state sequence output by the acoustic model includes:
[0024] Inputting the acoustic features of each audio data frame in the audio data into the feature extraction layer in the acoustic model to obtain a sequence of feature vectors output by the feature extraction layer;
[0025] Inputting the sequence of feature vectors into the classification layer in the acoustic model to obtain a phoneme-level state sequence output by the classification layer in the acoustic model.
[0026] The present invention also provides a voice wake-up device, including:
[0027] A phoneme output module, configured to input each audio data frame in the audio data into an acoustic model to obtain a phoneme-level state sequence output by the acoustic model; the phoneme-level state sequence includes the phoneme-level state classification results of each audio data frame in the audio data, and the phoneme-level state classification results include single-phoneme states and triphone states;
[0028] A state determination module, configured to determine whether it is a voice wake-up state based on the phoneme-level state sequence;
[0029] Among them, the acoustic model is obtained by optimizing the classification layer in the trained model based on the first sample audio data frame and the corresponding hybrid phoneme state label of the first sample audio data frame; the trained model is trained based on the second sample audio data frame and the corresponding triphone state label of the second sample audio data frame; the hybrid phoneme state label includes the triphone state corresponding to the wake-up state and the monophone state corresponding to the non-wake-up state.
[0030] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the voice wake-up method as described in any one of the above is implemented.
[0031] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the voice wake-up method as described in any one of the above is implemented.
[0032] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the voice wake-up method as described in any one of the above is implemented.
[0033] For the voice wake-up method, device, electronic device, storage medium, and program product provided by the present invention, each audio data frame in the audio data is input into the acoustic model to obtain a phoneme-level state sequence output by the acoustic model. The acoustic model is obtained by optimizing the classification layer in the trained model based on the first sample audio data frame and the corresponding hybrid phoneme state label of the first sample audio data frame, and the hybrid phoneme state label includes the triphone state corresponding to the wake-up state and the monophone state corresponding to the non-wake-up state. Therefore, only when the corresponding phoneme is in the wake-up state is it necessary to label the triphone state, otherwise only the monophone state with fewer states needs to be labeled, thereby reducing the number of states output by the acoustic model, and further reducing the resource requirements of the acoustic model, ensuring that the voice wake-up method can be applied to low-power devices, realizing a low-power voice wake-up method, and ultimately enhancing the application scope of the voice wake-up method. At the same time, the trained model is trained based on the second sample audio data frame and the corresponding triphone state label of the second sample audio data frame. Therefore, the triphone state label is first used for training in the first training stage to utilize the fine-grained representation of the triphone and improve the sensitivity of the acoustic model to subtle changes in speech. Then, in the second training stage, only the classification layer is optimized, which can prevent the loss of information learned in the first training stage during the optimization process of the second training stage, thereby ensuring that the classification layer outputs fewer states while improving the performance of the acoustic model, and further improving the detection accuracy of voice wake-up. Description of the Drawings
[0034] To more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0035] Figure 1 It is one of the schematic flowcharts of the voice wake-up method provided by the present invention.
[0036] Figure 2 It is the second of the schematic flowcharts of the voice wake-up method provided by the present invention.
[0037] Figure 3 It is the third of the schematic flowcharts of the voice wake-up method provided by the present invention.
[0038] Figure 4 It is the fourth of the schematic flowcharts of the voice wake-up method provided by the present invention.
[0039] Figure 5 It is the fifth of the schematic flowcharts of the voice wake-up method provided by the present invention.
[0040] Figure 6 It is the sixth of the schematic flowcharts of the voice wake-up method provided by the present invention.
[0041] Figure 7 It is the schematic structural diagram of the voice wake-up device provided by the present invention.
[0042] Figure 8 It is the schematic structural diagram of the electronic device provided by the present invention. Detailed implementation manners
[0043] To make the purpose, technical solutions and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention in conjunction with the drawings in the present invention. Obviously, the described embodiments are some embodiments of the present invention, rather than all embodiments. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0044] In the current voice wake-up scheme, to ensure the accuracy of voice wake-up detection, most use a neural network model to predict the triphone state corresponding to each frame of audio signal. However, when predicting the triphone state corresponding to each frame of audio signal, the number of output states is relatively large. Therefore, the scheme of predicting triphones cannot be applied to low-power devices and has limitations.
[0045] In view of the problem that the current voice wake-up solutions have high resource requirements, the present invention has conducted research. The initial idea was to directly predict whether each audio signal frame is a wake-up using an end-to-end solution, that is, the prediction result for each frame is 0 or 1, where 0 means not wake-up and 1 represents wake-up; thus, the triphone state of each audio signal frame is no longer predicted. However, since the end-to-end solution is a holistic modeling, it mainly provides a wake-up search interval and cannot accurately know the pronunciation state corresponding to each frame like at the frame level. Therefore, the accuracy of voice wake-up detection by the end-to-end solution is not high.
[0046] In view of the problems existing in the above idea, the present invention continued to conduct research. During the research process, it was thought that it is necessary to determine whether each audio data frame belongs to the wake-up state. When in the wake-up state, the triphone state label is used for model training, and when in the non-wake-up state, the monophone state label is used for model training, that is, the annotation data of the sample audio data frame is either a triphone state label or a monophone state label. However, this method targets audio data frames, and the granularity is too large, resulting in low accuracy of voice wake-up detection.
[0047] In view of the above defects, the present invention conducted research again. During the research process, it was thought that first, the triphone state label of the sample audio data frame is determined, and then the triphone state corresponding to the non-wake-up state in the triphone state label is changed to the monophone state to obtain a mixed phoneme state label, and thus the model sequence is directly based on the mixed phoneme state label. However, since the mixed phoneme state label includes the monophone state, the granularity is still relatively coarse, resulting in low accuracy of voice wake-up detection.
[0048] In view of the above defects, the present invention further conducts research and finally proposes a voice wake-up method. The voice wake-up method first performs model training based on the second sample audio data frame and its corresponding triphone state label to obtain a trained model, and then optimizes the classification layer in the trained model based on the first sample audio data frame and its corresponding mixed phoneme state label to obtain an acoustic model, so as to perform model training based on the sample audio data frames with all triphone state labels to improve the feature extraction performance of the acoustic model. Then, the classification layer in the trained model is optimized based on the mixed phoneme state label to obtain an acoustic model, thereby reducing the number of states output by the acoustic model and further reducing the resources required by the acoustic model.
[0049] Next, the voice wake-up method provided by the present invention will be introduced through the following embodiments. The following will be combined with Figures 1 - 6 Describe the voice wake-up method of the present invention.
[0050] The execution subject of the voice wake-up method provided by the present invention can be an electronic device or a chip (such as a low-power chip), and the electronic device can be a voice recognition device such as a robot, a mobile phone, a smart home device, a car machine, a learning machine, and a wearable device.
[0051] Figure 1 is one of the schematic flowcharts of the voice wake-up method provided by the present invention. As Figure 1 shown, the voice wake-up method includes the following steps 110 and 120.
[0052] Step 110: Input each audio data frame in the audio data into an acoustic model to obtain a phoneme-level state sequence output by the acoustic model.
[0053] Here, the audio data is the audio data to be detected for whether it triggers the voice wake-up state, that is, to determine whether the corresponding text data includes a preset wake-up word. The text data corresponding to the audio data may or may not include the preset wake-up word. The preset wake-up word is a pre-defined wake-up word, such as "Xiaofei, Xiaofei".
[0054] The audio data can be collected by a voice collection device. For example, the voice collection device is a microphone on a mobile phone. The audio data may only include user voice data, may also include environmental noise data, and of course may not include user voice data, that is, the audio data is not specifically limited.
[0055] Here, the audio data frame is the frame data split from the audio data. In one embodiment, the audio data is sampled based on a preset sampling rate to obtain each audio data frame; the preset sampling rate can be set according to actual needs. For example, it is 16 kHz.
[0056] Here, the acoustic model is used to predict the phoneme-level state classification results of each audio data frame in the audio data. Specifically, the classification layer in the acoustic model predicts the phoneme-level state classification results of each audio data frame in the audio data.
[0057] Among them, the acoustic model is obtained by optimizing the classification layer in the trained model based on the first sample audio data frame and the mixed phoneme state label corresponding to the first sample audio data frame; the trained model is trained based on the second sample audio data frame and the triphone state label corresponding to the second sample audio data frame; the mixed phoneme state label includes the triphone state corresponding to the wake-up state and the monophone state corresponding to the non-wake-up state.
[0058] In other words, the acoustic model is trained based on the following method: the trained model is obtained by training the model based on the second sample audio data frame and the triphone state label corresponding to the second sample audio data frame; the acoustic model is obtained by optimizing the classification layer in the trained model based on the first sample audio data frame and the mixed phoneme state label corresponding to the first sample audio data frame.
[0059] Here, the second sample audio data frame is any audio data frame in the sample audio data, and it can also be frame data split from the sample audio data. Similarly, the first sample audio data frame can be any audio data frame in the sample audio data, and it can also be frame data split from the sample audio data. Usually, there are multiple sample audio data, that is, a training data set is obtained for model training.
[0060] Specifically, model training can be performed based on multiple second sample audio data frames and the corresponding triphone state labels of each second sample audio data frame to obtain a trained model; based on multiple first sample audio data frames and the corresponding hybrid phoneme state labels of each first sample audio data frame, the classification layer in the trained model can be optimized to obtain an acoustic model.
[0061] Here, the triphone state label is the labeled triphone state, and the triphone state is the state based on the triphone representation, that is, the triphone state can also be understood as a triphone. A triphone refers to the variant of a phoneme under the influence of adjacent phonemes before and after. Its structure is usually expressed as the left phoneme - the current phoneme + the right phoneme (such as B - A + C), indicating the pronunciation of the current phoneme A under the joint action of the left phoneme B and the right phoneme C. For the triphone state, it not only focuses on the acoustic features of the current audio data frame itself, but also fully considers the phoneme information of its front and back frames, so as to capture the subtle changes caused by co-articulation between phonemes. It should be understood that the triphone state label includes multiple triphone states. For example, it includes 9004 triphone states.
[0062] Here, the monophone state is the state based on the monophone representation, that is, the monophone state can also be understood as a monophone. For the monophone state, although it is relatively simple, it is also necessary to accurately identify the characteristics of the independent pronunciation of each audio data frame.
[0063] It should be understood that the hybrid phoneme state label includes multiple triphone states corresponding to the wake-up state and multiple monophone states corresponding to the non-wake-up state, that is, each phoneme can determine whether it is in the wake-up state or the non-wake-up state. Among them, the phoneme corresponding to the wake-up state refers to any phoneme included in the preset wake-up word, and the phoneme corresponding to the non-wake-up state refers to any other phoneme other than the phoneme corresponding to the wake-up state. Whether a triphone is a phoneme corresponding to the wake-up state can be based on the middle phoneme of the triphone. Based on this, the hybrid phoneme state label only needs to retain the triphone state when the corresponding phoneme is in the wake-up state. Therefore, the hybrid phoneme state label is a more simplified form of phoneme combination, which helps to reduce the number of states output by the classification layer, thereby simplifying the model structure and improving the calculation efficiency, that is, reducing the resource requirements of the acoustic model.
[0064] Based on the above, the acoustic model is trained in two training stages. The second sample audio data frames and their corresponding triphone state labels are the training data for the first training stage. The first sample audio data frames and their corresponding mixed phoneme state labels are the training data for the second training stage. And in the first training stage, training is carried out using the triphone state labels, so as to utilize the fine-grained representation of triphones and improve the sensitivity of the acoustic model to subtle changes in speech. Compared with directly training using the mixed phoneme state labels, the performance of the acoustic model can be improved, thereby improving the accuracy of speech wake-up detection. That is to say, in the first training stage, the model is trained to predict triphone states, and in the second training stage, the model parameters except the classification layer are fixed (frozen), and the model is trained to predict mixed phoneme states, and only the classification layer is allowed to be trained according to the new mixed phoneme state labels, ensuring that the acoustic model outputs phoneme-level state classification results including single-phoneme states and triphone states in the inference stage.
[0065] It should be noted that in the second training stage, the model parameters except the classification layer are fixed (frozen). That is, after the acoustic model is trained, the branch for triphone state prediction is discarded, and only the branch for mixed phoneme state prediction is retained, so as to ensure that the acoustic model in the inference stage only outputs the mixed phoneme states (phoneme-level state classification results).
[0066] It should be understood that fixing (freezing) the model parameters except the classification layer in the second training stage, especially those network layers that have learned rich triphone information, can prevent the loss of these important low-level speech features during the fine-tuning process in the second training stage. The benefits of only optimizing the classification layer in the second training stage are as follows.
[0067] First, maintaining low-level features: By freezing the model parameters except the classification layer, the triphone-level details learned by the acoustic model in the first training stage are retained, thus ensuring that the acoustic model does not forget these key information.
[0068] Second, focusing on high-level tasks: Only updating the last classification layer can make the acoustic model more focused on mapping these low-level features to the final speech wake-up state determination task, improving the performance of the acoustic model in the speech wake-up detection task.
[0069] Third, improving training efficiency: Since there is no need to update a large number of parameters, the training speed is significantly accelerated, and at the same time, the demand for computing resources is reduced.
[0070] The acoustic model includes a classification layer, which is used to obtain a phoneme-level state classification result through classification. Further, the acoustic model includes a feature extraction layer and a classification layer connected in sequence. Further still, the acoustic model includes a feature extraction layer, a quantization layer, and a classification layer connected in sequence. The feature extraction layer is used to extract the input features. In one embodiment, the quantization layer is used to reshape or pad the features extracted by the feature extraction layer into a unified size, and normalize the features, so as to accelerate the training process and improve the model performance.
[0071] Wherein, the phoneme-level state sequence includes the phoneme-level state classification results of each audio data frame in the audio data. That is, each audio data frame is respectively input into the acoustic model to obtain the phoneme-level state classification results of each audio data frame output by the acoustic model.
[0072] Wherein, the phoneme-level state classification results include single-phoneme states and triphone states. Since the acoustic model is optimized for the classification layer in the trained model based on the first sample audio data frame and the corresponding mixed phoneme state label of the first sample audio data frame, and the mixed phoneme state label includes the triphone state corresponding to the wake-up state and the single-phoneme state corresponding to the non-wake-up state, the phoneme-level state classification results output by the acoustic model include single-phoneme states and triphone states.
[0073] Exemplarily, the triphone state label includes M triphone states; the mixed phoneme state label includes N triphone states corresponding to the wake-up state and L single-phoneme states corresponding to the non-wake-up state, and N + L is less than M. It should be understood that the triphone state label usually includes more triphone states. For example, M = 9004, while only the wake-up state in the mixed phoneme state label retains the triphone state, and the non-wake-up state directly uses the single-phoneme state with fewer states. For example, N = 30 and L = 83, so N + L is much less than M. The reason why N is small is that the preset wake-up word includes fewer phonemes, and the reason why L is small is that the single-phoneme state itself is relatively few.
[0074] In one embodiment, the triphone state label is obtained in the following manner: obtaining a second sample audio data frame in the sample audio data and the text data corresponding to the sample audio data; determining the triphone state label corresponding to the second sample audio data frame based on the forced alignment result of the sample audio data and the text data.
[0075] In one embodiment, the mixed phoneme state label is obtained by changing the triphone states corresponding to non-awakening states in the pre-labeled triphone state labels to monophone states. For example, the mixed phoneme state label is obtained in the following manner: obtaining the second sample audio data frame and the corresponding triphone state label of the second sample audio data frame, and changing the triphone states corresponding to non-awakening states in the triphone state label to monophone states to obtain the mixed phoneme state label.
[0076] Step 120: Determine whether it is a voice wake-up state based on the phoneme-level state sequence.
[0077] It should be noted that if the text data corresponding to the audio data includes a preset wake-up word, it is determined to be a voice wake-up state; if the text data corresponding to the audio data does not include a preset wake-up word, it is determined not to be a voice wake-up state. Of course, it is not directly determined whether the text data includes a preset wake-up word, but the phoneme-level state sequence is analyzed to determine whether it is a voice wake-up state.
[0078] In the voice wake-up state, the user's voice is recognized to give a specified response. Exemplarily, the voice wake-up method provided in the embodiments of the present invention is applied to the front-end module of a speech recognition system to achieve real-time monitoring and recognition of the user's voice. When a preset wake-up word or preset phrase is detected, the subsequent speech recognition module or human-computer interaction module is quickly activated.
[0079] In a specific embodiment, first, the audio data is classified into phoneme-level states by using a pre-constructed acoustic model, and then the frame-level phoneme-level state classification result is input into a decoder to obtain the wake-up state determination result output by the decoder. The wake-up state determination result of the audio data is used to represent whether it is a voice wake-up state. More specifically, the classification probabilities of the phoneme-level state classification result are input into the decoder. For example, if the phoneme-level state classification result includes N + L phoneme states, the classification probabilities of the N + L phoneme states are input into the decoder.
[0080] In one embodiment, the decoder can perform decoding using the Viterbi decoding algorithm, that is, using the Viterbi decoding algorithm to decode the phoneme-level state sequence to determine whether it is a voice wake-up state. Of course, decoding can also be performed in other ways, which are not specifically limited here.
[0081] Among them, the Viterbi decoding algorithm will further analyze and process the predicted phoneme-level state sequence. It integrates these scattered phoneme-level state classification results into a wake-up state determination result with clear semantics according to the preset decoding rules and probability models. In this way, it can accurately identify whether the user has uttered a preset wake-up word, thereby realizing an efficient and accurate voice wake-up function.
[0082] It should be understood that through the acoustic model, a detailed frame-level analysis and prediction of the audio data is performed, that is, the phoneme-level state sequence includes the phoneme-level state classification results of each audio data frame in the audio data, so as to more accurately determine whether it is a voice wake-up state based on the phoneme-level state sequence, that is, to improve the detection accuracy of voice wake-up.
[0083] For the convenience of understanding the inventive concept of the embodiments of the present invention, an example is given here. Currently, through the neural network model, the triphone state corresponding to each frame of audio signal is predicted, and the number of states of the output triphone state is 9004. If the number of channels of the classification layer is 64, then the size of the classification layer will be 9004 64 / 1024 = 562.75K. The parameters of the classification layer will be larger than those of the feature extraction model in the model and cannot be run on low-cost devices or chips; through the neural network model, the monophone state corresponding to each frame of audio signal is predicted, and the number of states of the output monophone state is 83. If the number of channels of the classification layer is 64, then the size of the classification layer is 83 64 / 1024 = 5.1875K. Although monophone modeling can greatly compress the model memory, its effect is poor. Based on this, in the embodiments of the present invention, the classification layer is optimized based on the hybrid phoneme state label, so that the classification layer outputs the triphone states corresponding to N (for example, 45) wake-up states and the monophone states corresponding to 83 non-wake-up states. Based on this, if N = 45, that is, the number of states output by the classification layer is 128, the classification layer can be compressed from 562.75K to 8K, so that it can be run on low-power devices and ensure a sufficient wake-up rate.
[0084] It should be understood that through the two-stage training strategy of the above acoustic model, while ensuring that the acoustic model has a strong representation ability (guaranteed by the first training stage), an efficient and reliable phoneme-level state classification can be achieved. And by freezing most of the model parameters in the second training stage, not only the problem of forgetting triphone information is avoided, but also the resources required by the acoustic model are reduced, thereby enhancing the generalization ability of the acoustic model in actual application scenarios.
[0085] The voice wake-up method provided by the embodiment of the present invention inputs each audio data frame in the audio data into an acoustic model to obtain a phoneme-level state sequence output by the acoustic model. The acoustic model is obtained by optimizing the classification layer in the trained model based on the first sample audio data frame and the corresponding mixed phoneme state label of the first sample audio data frame. The mixed phoneme state label includes the triphone state corresponding to the wake-up state and the monophone state corresponding to the non-wake-up state. Therefore, only when the corresponding phoneme is in the wake-up state is it necessary to label the triphone state, otherwise only the monophone state with fewer states needs to be labeled, thereby reducing the number of states output by the acoustic model, further reducing the resource requirements of the acoustic model, ensuring that the voice wake-up method can be applied to low-power devices, realizing a low-power voice wake-up method, and ultimately enhancing the application scope of the voice wake-up method. At the same time, the trained model is trained based on the second sample audio data frame and the corresponding triphone state label of the second sample audio data frame. Therefore, the triphone state label is first used for training in the first training stage to utilize the fine-grained representation of the triphone and improve the sensitivity of the acoustic model to subtle changes in speech. Then, only the classification layer is optimized in the second training stage, which can prevent the loss of information learned in the first training stage during the optimization process in the second training stage, thereby ensuring that the classification layer outputs fewer states while improving the performance of the acoustic model, and further improving the detection accuracy of voice wake-up.
[0086] Based on any of the above embodiments, another specific embodiment of the voice wake-up method is given below. This embodiment is used to introduce the training method of the acoustic model. The execution subject of the training method of the acoustic model can be the same as or different from the execution subject of the voice wake-up method. For example, the execution subject of the training method of the acoustic model is a training-end device.
[0087] Figure 2 It is the second flow diagram of the voice wake-up method provided by the present invention. As Figure 2 shown, the acoustic model is trained based on the following steps 210 and 220.
[0088] Step 210: Train an initial model based on the second sample audio data frame and the corresponding triphone state label of the second sample audio data frame to obtain a trained model.
[0089] Here, the second sample audio data frame is any audio data frame in the sample audio data. Usually, there are multiple sample audio data, that is, a training data set is obtained for model training.
[0090] In one embodiment, the sample audio data set of the sample audio data includes: general audio data and wake word audio data; the general audio data does not include the audio data corresponding to the preset wake word, and the wake word audio data includes the audio data corresponding to the preset wake word. That is, a large amount of general audio data that does not include the preset wake word and a large amount of wake word audio data that includes the preset wake word are obtained to be used as the sample audio data set for model training. Based on this, in addition to including the wake word audio data, it also includes general audio data, which helps the acoustic model better learn to distinguish background noise from the preset wake word, thereby improving the detection accuracy of voice wake-up.
[0091] Specifically, the initial model can be trained based on multiple second sample audio data frames and the corresponding triphone state labels of each second sample audio data frame to obtain the trained model.
[0092] Here, the initial model is an initial network model to be trained, and this initial model can be selected according to actual needs and is not specifically limited here. Of course, the initial model can also be a preliminarily trained model, so as to reduce the training amount and improve the model performance.
[0093] Here, the trained model can predict the triphone state corresponding to the audio data frame, that is to say, the first training stage trains the model to predict the triphone state.
[0094] Specifically, in the first training stage, the model parameters of the entire initial model are adjusted, so that in the first training stage, the triphone state labels are used for training, thereby using the fine-grained representation of the triphone to improve the sensitivity of the acoustic model to the subtle changes in speech, and further improving the performance of the acoustic model, thereby improving the detection accuracy of voice wake-up.
[0095] In one embodiment, the triphone state label is obtained in the following manner: obtain the second sample audio data frame in the sample audio data and the text data corresponding to the sample audio data; based on the forced alignment result of the sample audio data and the text data, determine the triphone state label corresponding to the second sample audio data frame.
[0096] Step 220, optimize the classification layer in the trained model based on the first sample audio data frame and the corresponding mixed phoneme state label of the first sample audio data frame to obtain the acoustic model.
[0097] Here, the first sample audio data frame can also be any audio data frame in the sample audio data.
[0098] Specifically, the classification layer in the trained model can be optimized based on multiple first sample audio data frames and the corresponding mixed phoneme state labels of each first sample audio data frame to obtain the acoustic model.
[0099] It should be understood that the mixed phoneme state label includes the triphone states corresponding to multiple wake-up states and the monophone states corresponding to multiple non-wake-up states, that is, each phoneme can determine whether it is in a wake-up state or a non-wake-up state. Based on this, the mixed phoneme state label only needs to retain the triphone state when the corresponding phoneme is in a wake-up state. Thus, the mixed phoneme state label is a more simplified form of phoneme combination, which helps to reduce the number of states output by the classification layer, thereby reducing the resource requirements of the acoustic model.
[0100] Since the model parameters except for the classification layer are fixed (frozen) in the second training stage, and the model is trained to predict the mixed phoneme state, only allowing the classification layer to be trained according to the new mixed phoneme state label, the acoustic model can only output the phoneme-level state classification results including monophone states and triphone states. That is, after the acoustic model is trained, the branch for triphone state prediction is discarded, and only the branch for mixed phoneme state prediction is retained, so as to ensure that the acoustic model in the inference stage only outputs the mixed phoneme state (phoneme-level state classification result).
[0101] It should be understood that fixing (freezing) the model parameters except for the classification layer in the second training stage, especially those network layers that have learned rich triphone information, can prevent the loss of these important low-level speech features during the fine-tuning process in the second training stage.
[0102] In one embodiment, the mixed phoneme state label is obtained by changing the triphone states corresponding to non-wake-up states in the pre-annotated triphone state label to monophone states. For example, the mixed phoneme state label is obtained in the following way: obtain the second sample audio data frame and the corresponding triphone state label of the second sample audio data frame, and change the triphone states corresponding to non-wake-up states in the triphone state label to monophone states to obtain the mixed phoneme state label.
[0103] The voice wake-up method provided by the embodiment of the present invention trains an initial model based on the second sample audio data frame and the corresponding triphone state label of the second sample audio data frame, so as to first perform the first training stage training using the triphone state label to utilize the fine-grained representation of the triphone and improve the sensitivity of the acoustic model to subtle changes in speech. However, based on the first sample audio data frame and the corresponding mixed phoneme state label of the first sample audio data frame, the classification layer in the trained model is optimized to obtain an acoustic model, that is, only the classification layer is optimized in the second training stage, which can prevent the loss of information learned in the first training stage during the optimization process of the second training stage, thereby ensuring that the classification layer outputs a smaller number of states while improving the performance of the acoustic model, and further improving the detection accuracy of voice wake-up; at the same time, only the parameters of the entire model need to be adjusted in the first training stage, and only the parameters of the classification layer need to be adjusted in the second training stage, so as to reduce the training calculation amount and improve the training efficiency.
[0104] Based on any of the above embodiments, another specific embodiment of the voice wake-up method is given below. Figure 3 It is the third flowchart of the voice wake-up method provided by the present invention, as Figure 3 shown. The acoustic model is trained based on the following steps 230, 240, 210, and 220.
[0105] Step 230, obtain the second sample audio data frame and the corresponding triphone state label of the second sample audio data frame.
[0106] In a specific embodiment, obtain the second sample audio data frame in the sample audio data and the corresponding text data of the sample audio data; based on the forced alignment result of the sample audio data and the text data, determine the triphone state label corresponding to the second sample audio data frame.
[0107] Among them, the triphone state label includes M triphone states. Among them, M is greater than 1. For example, the triphone state label includes 9004 triphone states.
[0108] Step 240, change the triphone states corresponding to the non-wake-up states in the triphone state label to monophone states to obtain the mixed phoneme state label.
[0109] Here, the triphone state label is the labeled triphone state, and the triphone state is the state based on the triphone representation, that is, the triphone state can also be understood as the triphone. A triphone refers to the variant of a phoneme under the influence of adjacent phonemes before and after. Its structure is usually represented as the left phoneme - the current phoneme + the right phoneme (such as B - A + C), indicating the pronunciation of the current phoneme A under the joint action of the left phoneme B and the right phoneme C. Therefore, it is judged whether it is a wake-up state based on the current phoneme A in the triphone.
[0110] It is understandable that some of the phoneme states in the triphone state labels are non-awakened states. Therefore, the triphone states corresponding to these non-awakened states can be changed to monophone states. That is, the triphone states in the non-awakened states are collapsed to monophone states, and the phonemes corresponding to the awakened states retain their original triphone states.
[0111] In a specific embodiment, the middle phoneme (the current phoneme) in the triphone state corresponding to the non-awakened state is retained as the changed monophone state. For example, the triphone structure is represented as left phoneme - current phoneme + right phoneme (such as B - A + C), indicating the pronunciation of the current phoneme A under the combined action of the left phoneme B and the right phoneme C. Therefore, the changed monophone is A.
[0112] Wherein, the first sample audio data frame is the same as the second sample audio data frame.
[0113] Wherein, the mixed phoneme state label includes N triphone states corresponding to awakened states and L monophone states corresponding to non-awakened states; N + L is less than M.
[0114] Exemplarily, the triphone state label usually includes a relatively large number of triphone states. For example, M = 9004, while only the awakened states in the mixed phoneme state label retain the triphone states, and the non-awakened states directly use the monophone states with a smaller number of states. For example, N = 30 and L = 83, so that N + L is much less than M, and then the states are re-sorted to obtain the mixed phoneme state label with N + L state numbers. The reason why N is small is that the preset wake-up word includes relatively few phonemes, and the reason why L is small is that the monophone state itself is relatively few.
[0115] Step 210: Train the initial model based on the second sample audio data frame and the triphone state label corresponding to the second sample audio data frame to obtain a trained model.
[0116] Step 220: Optimize the classification layer in the trained model based on the first sample audio data frame and the mixed phoneme state label corresponding to the first sample audio data frame to obtain the acoustic model.
[0117] The voice wake-up method provided by the embodiment of the present invention first obtains the second sample audio data frame and the corresponding triphone state label of the second sample audio data frame, and then changes the triphone state corresponding to the non-wake-up state in the triphone state label to a monophone state, so as to obtain a mixed phoneme state label without re-labeling the mixed phoneme state label, thereby reducing the label annotation cost; and the number of states of the mixed phoneme state label is less than the number of states of the triphone state label, thereby reducing the number of states output by the acoustic model, and further reducing the resource requirements of the acoustic model, ensuring that the voice wake-up method can be applied to low-power devices, realizing a low-power voice wake-up method, and finally enhancing the application scope of the voice wake-up method.
[0118] Based on any of the above embodiments, another specific embodiment of the voice wake-up method is given next. Figure 4 It is the fourth flowchart of the voice wake-up method provided by the present invention. As Figure 4 shown, the acoustic model is trained based on the following steps 231, 232, 240, 210, and 220.
[0119] Step 231, obtain the second sample audio data frame in the sample audio data, and the text data corresponding to the sample audio data.
[0120] Here, the sample audio data can be split into multiple second sample audio data frames. There are usually multiple sample audio data, that is, a training data set is obtained for model training.
[0121] Among them, the sample audio data set containing the sample audio data includes: general audio data and wake-up word audio data; the general audio data does not include the audio data corresponding to the preset wake-up word, and the wake-up word audio data includes the audio data corresponding to the preset wake-up word.
[0122] Here, the sample audio data set includes multiple sample audio data. The preset wake-up word is a pre-set wake-up word, such as "Xiaofei, Xiaofei".
[0123] It should be understood that general audio data that does not include the preset wake-up word and wake-up word audio data that includes the preset wake-up word are obtained as the sample audio data set for model training. Based on this, in addition to including wake-up word audio data, it also includes general audio data, which helps the acoustic model better learn to distinguish background noise from the preset wake-up word, thereby improving the detection accuracy of voice wake-up.
[0124] In one embodiment, the general audio data is obtained by collecting a large amount of daily environmental audio, provided that this audio does not include the preset wake-up word. Further, the acquisition sources of the general audio data may include, but are not limited to: daily conversations, background music, environmental noises (such as wind sounds, traffic noises, etc.), ensuring that the general audio data is as diverse as possible to improve the robustness of the acoustic model during the inference stage, thereby improving the detection accuracy of voice wake-up.
[0125] In one embodiment, the wake-up word audio data is obtained by collecting audio including the preset wake-up word. To further enhance the sample data, audio of different speakers can be collected, or audio of different speaking speeds can be collected, or audio of different volumes can be collected, or audio under different background noise conditions can be collected. Of course, all audio under real-world conditions should be collected as much as possible; based on this, an acoustic model that can accurately recognize the wake-up word under various conditions is trained, that is, the detection accuracy of voice wake-up is improved.
[0126] In one embodiment, the wake-up word audio data can be used to record the voices of different people with a recording device, or text-to-speech technology can be used to generate diverse wake-up word audio data.
[0127] Further, the wake-up word audio data is high-fidelity audio data to ensure the robustness of the acoustic model.
[0128] In one embodiment, the sample audio data can be used to record different audio with a recording device.
[0129] Here, the text data can be the transcribed text of the sample audio data.
[0130] Step 232, based on the forced alignment result of the sample audio data and the text data, determine the triphone state label corresponding to the second sample audio data frame.
[0131] Here, the forced alignment result can be understood as the correspondence between each sample audio data frame in the sample audio data and each word in the text data, that is, the forced alignment result provides a direct mapping from the audio data to its corresponding phoneme state label sequence. Specifically, the forced alignment result is obtained by performing forced alignment on the sample audio data and the text data. Forced Alignment (FA) is a technique used to precisely match audio data with the transcribed text, thereby assigning corresponding triphone state labels to each frame of the sample audio data.
[0132] Of course, the mixed phoneme state label corresponding to the first sample audio data frame can also be determined based on the forced alignment result of the sample audio data and the text data.
[0133] Step 240: Change the triphone states corresponding to the non-awakening states in the triphone state labels to monophone states to obtain the mixed phoneme state labels.
[0134] Step 210: Train an initial model based on the second sample audio data frames and the triphone state labels corresponding to the second sample audio data frames to obtain a trained model.
[0135] Step 220: Optimize the classification layer in the trained model based on the first sample audio data frames and the mixed phoneme state labels corresponding to the first sample audio data frames to obtain the acoustic model.
[0136] The voice wake-up method provided by the embodiments of the present invention obtains the second sample audio data frames in the sample audio data and the text data corresponding to the sample audio data, so as to accurately determine the triphone state labels corresponding to the second sample audio data frames based on the forced alignment result of the sample audio data and the text data, thereby improving the training effect of the acoustic model and ultimately improving the detection accuracy of voice wake-up; at the same time, the sample audio data set containing the sample audio data includes: general audio data and wake-up word audio data, and the general audio data does not include the audio data corresponding to the preset wake-up word, and the wake-up word audio data includes the audio data corresponding to the preset wake-up word, thereby helping the acoustic model to better learn to distinguish background noise from the preset wake-up word, and further improving the detection accuracy of voice wake-up.
[0137] Based on any of the above embodiments, a further specific embodiment of the voice wake-up method is given next. Figure 5 It is the fifth flowchart of the voice wake-up method provided by the present invention. As Figure 5 shown, the acoustic model is trained based on the following steps 211 and 220.
[0138] Step 211: Train an initial model based on the second sample audio data frames, the triphone state labels corresponding to the second sample audio data frames, and the binary classification result labels corresponding to the second sample audio data frames to obtain a trained model.
[0139] Here, the second sample audio data frame is any audio data frame in the sample audio data. Usually, there are multiple sample audio data, that is, a training data set is obtained for model training.
[0140] Among them, the binary classification result labels include the first classification result or the second classification result. The first classification result is used to represent that the second sample audio data frame belongs to the audio data frame corresponding to the end position of the wake-up word, and the second classification result is used to represent that the second sample audio data frame does not belong to the audio data frame corresponding to the end position of the wake-up word.
[0141] This binary classification is a boundary classification. Considering that the triphone state label only focuses on local information, the binary classification result label is used to focus on the overall modeling information, so as to better train the feature extraction ability of the acoustic model. That is, the boundary loss function of this boundary classification is used to particularly focus on the end position of the preset wake-up word, that is, to better determine the exact boundary of the preset wake-up word, thereby improving the detection accuracy of voice wake-up.
[0142] Exemplarily, in order to optimize the local and overall modeling capabilities of the acoustic model simultaneously, the initial model is trained through the boundary loss function and the triphone state prediction loss function.
[0143] In a specific embodiment, the comprehensive loss is determined through the boundary loss function and the triphone state prediction loss function, and the initial model is trained based on the comprehensive loss.
[0144] In an embodiment, the formula for determining the comprehensive loss is as follows:
[0145] ;
[0146] In the formula, represents the time frame, represents the comprehensive loss of the second sample audio data frame of the th frame, represents the triphone state label corresponding to the second sample audio data frame of the th frame, represents the triphone state prediction result corresponding to the second sample audio data frame of the th frame, represents the triphone state prediction loss; is a weight factor used to balance the importance of the two losses; represents the binary classification result label of the second sample audio data frame of the th frame, represents the binary classification prediction result of the second sample audio data frame of the th frame,
[0147] In an embodiment, the triphone state prediction loss function is a cross-entropy loss function. Exemplarily, the triphone state prediction loss function is as follows:
[0148] ;
[0149] In the formula, represents the triphone category. For example, if the triphone state label includes 9004 triphones, then the triphone category has a total of 9004; represents the The th triphone in the triphone state label corresponding to the second sample audio data frame of the th frame;
[0150] In one embodiment, the boundary loss function is a binary classification loss function. Exemplarily, the boundary loss function is as follows:
[0151] ;
[0152] In the formula, represents the binary classification result label of the th frame of the second sample audio data frame, represents the binary classification prediction result of the th frame of the second sample audio data frame. Further, assume that represents that the second sample audio data frame belongs to the audio data frame corresponding to the end position of the wake word, represents that the second sample audio data frame does not belong to the audio data frame corresponding to the end position of the wake word.
[0153] Specifically, the initial model can be trained based on multiple second sample audio data frames, the triphone state labels corresponding to each second sample audio data frame, and the binary classification result labels corresponding to each second sample audio data frame to obtain a trained model.
[0154] Here, the trained model can predict the triphone state corresponding to the audio data frame, that is, the first training stage trains the model to predict the triphone state. It should be noted that after the acoustic model training is completed, the branch of boundary classification is discarded, and only the branch of mixed phoneme state prediction is retained, so as to ensure that the acoustic model in the inference stage only outputs the mixed phoneme state (phoneme-level state classification result).
[0155] Specifically, the first training stage is to adjust the model parameters of the entire initial model. In the first training stage, the triphone state label and the binary classification result label are used for training, so as to utilize the fine-grained representation of the triphone to improve the sensitivity of the acoustic model to the subtle changes of the speech, and then improve the performance of the acoustic model, thereby improving the accuracy of speech wake-up detection; and utilize the binary classification result label to focus on the overall modeling information and assist in training the feature extraction ability of the acoustic model.
[0156] In other words, in the first training stage, all model parameters are updated, which means that all layers of the network participate in the learning process. This ensures that the model can fully learn the complex structures and patterns in the audio data and effectively capture the key features of the preset wake-up word, thereby improving the detection accuracy of voice wake-up.
[0157] Step 220: Optimize the classification layer in the trained model based on the first sample audio data frame and the corresponding hybrid phoneme state label to obtain the acoustic model.
[0158] Here, the first sample audio data frame can also be any audio data frame in the sample audio data.
[0159] Specifically, the classification layer in the trained model can be optimized based on multiple first sample audio data frames and the corresponding hybrid phoneme state labels of each first sample audio data frame to obtain the acoustic model.
[0160] Since the second training stage fixes (freezes) the model parameters except for the classification layer and trains the model to predict the hybrid phoneme state, only allowing the classification layer to be trained according to the new hybrid phoneme state label, the acoustic model can only output the phoneme-level state classification results including single-phoneme state and tri-phoneme state. That is, after the acoustic model is trained, the branches for tri-phoneme state prediction and boundary classification are discarded, and only the branch for hybrid phoneme state prediction is retained, so as to ensure that the acoustic model in the inference stage only outputs the hybrid phoneme state (phoneme-level state classification result).
[0161] In a specific embodiment, the classification layer in the trained model is optimized through a hybrid phoneme state prediction loss function.
[0162] In an embodiment, the hybrid phoneme state prediction loss function is a cross-entropy loss function. Exemplarily, the hybrid phoneme state prediction loss function is as follows:
[0163] ;
[0164] In the formula, represents the hybrid phoneme category; represents the th phoneme in the hybrid phoneme state label corresponding to the th frame of the first sample audio data frame; represents the th phoneme in the hybrid phoneme state prediction result corresponding to the th frame of the first sample audio data frame.
[0165] It should be understood that the embodiments of the present invention combine multi-level acoustic feature analysis with a fine classification strategy, improving the performance of the acoustic model and thus enhancing the detection accuracy of voice wake-up.
[0166] The voice wake-up method provided by the embodiments of the present invention trains an initial model based on the second sample audio data frames, the corresponding triphone state labels of the second sample audio data frames, and the corresponding binary classification result labels of the second sample audio data frames to obtain a trained model. Thus, not only local information is concerned through the triphone state labels, but also overall modeling information is concerned through the binary classification result labels, so as to better train the feature extraction ability of the acoustic model and further improve the detection accuracy of voice wake-up.
[0167] Based on any of the above embodiments, another specific embodiment of the voice wake-up method is given next. Figure 6 It is the sixth flowchart of the voice wake-up method provided by the present invention, as Figure 6 shown. The voice wake-up method includes: step 111, step 112, and step 120.
[0168] In step 111, the acoustic features of each audio data frame in the audio data are input into the feature extraction layer in the acoustic model to obtain a sequence of feature vectors output by the feature extraction layer.
[0169] Here, the acoustic features may include but are not limited to: FBANK (Filter Bank Energies) features, mel cepstral coefficient features, intonation features, audio strength features, and so on.
[0170] In one embodiment, the acoustic feature is the fb40 feature, that is, a filter bank composed of 40 filters is used to capture different frequency components of the audio signal.
[0171] Here, the audio data frame is the frame data split from the audio data. In one embodiment, the audio data is sampled based on a preset sampling rate to obtain each audio data frame; the preset sampling rate can be set according to actual needs, for example, it is 16 kHz.
[0172] Exemplarily, the audio data is sampled based on a preset sampling rate to obtain each audio data frame; each audio data frame is subjected to short-time Fourier transform (STFT, Short Time Fourier Transform) to obtain a spectrogram; a set of triangular overlapping band-pass filters is applied to the spectrogram to obtain the filter bank energy corresponding to each time frame; the energy value of the filter bank energy corresponding to each time frame is calculated to obtain the acoustic feature.
[0173] Here, the feature extraction layer is used to extract the feature vectors of the acoustic features. The sequence of feature vectors includes the feature vectors of each acoustic feature.
[0174] Step 112: Input the feature vector sequence into the classification layer in the acoustic model to obtain the phoneme-level state sequence output by the classification layer in the acoustic model.
[0175] Here, the classification layer is used to predict the phoneme-level state sequence based on the feature vectors of each acoustic feature.
[0176] Among them, the phoneme-level state sequence includes the phoneme-level state classification results of each audio data frame in the audio data. That is, each audio data frame is respectively input into the acoustic model to obtain the phoneme-level state classification results of each audio data frame output by the acoustic model.
[0177] The voice wake-up method provided by the embodiment of the present invention inputs the acoustic features of each audio data frame in the audio data into the feature extraction layer in the acoustic model to obtain the feature vector sequence output by the feature extraction layer, and then inputs the feature vector sequence into the classification layer in the acoustic model to obtain the phoneme-level state sequence output by the classification layer in the acoustic model, so as to first extract the acoustic features and then extract their feature vectors, thereby improving the classification accuracy of the classification layer and further improving the detection accuracy of voice wake-up.
[0178] Through the above embodiments, the present invention realizes a voice wake-up method with extremely low power consumption. It mainly collapses triphones into hybrid phonemes instead of single phonemes, and predicts the distribution state of the hybrid phonemes and sends them to the decoder to obtain the final wake-up state. At the same time, a two-stage training method is introduced to improve the robustness of the voice wake-up method. Specifically, the two-stage training method not only ensures that the model can accurately recognize specific wake-up words, but also enhances its flexibility and adaptability when facing various wake-up words. This method significantly improves the overall performance and robustness of the system by finely adjusting the model parameters and the design of the loss function. That is, by modeling with hybrid phoneme state labels, while ensuring the effect, the model size is greatly reduced, so that the voice wake-up method can run on low-computing-power devices.
[0179] The voice wake-up device provided by the present invention will be described below. The voice wake-up device described below can be mutually corresponding and referred to the voice wake-up method described above.
[0180] Figure 7 is a schematic structural diagram of the voice wake-up device provided by the present invention. As Figure 7 shown, the voice wake-up device includes a phoneme output module 710 and a state determination module 720.
[0181] A phoneme output module 710 is configured to input each audio data frame in the audio data into an acoustic model to obtain a phoneme-level state sequence output by the acoustic model; the phoneme-level state sequence includes the phoneme-level state classification results of each audio data frame in the audio data, and the phoneme-level state classification results include single-phoneme states and tri-phoneme states.
[0182] A state determination module 720 is configured to determine whether it is a voice wake-up state based on the phoneme-level state sequence.
[0183] Wherein, the acoustic model is obtained by optimizing the classification layer in the pre-trained model based on the first sample audio data frame and the mixed phoneme state label corresponding to the first sample audio data frame; the pre-trained model is trained based on the second sample audio data frame and the tri-phoneme state label corresponding to the second sample audio data frame; the mixed phoneme state label includes the tri-phoneme state corresponding to the wake-up state and the single-phoneme state corresponding to the non-wake-up state.
[0184] The voice wake-up device provided by the embodiment of the present invention inputs each audio data frame in the audio data into an acoustic model to obtain a phoneme-level state sequence output by the acoustic model. The acoustic model is obtained by optimizing the classification layer in the pre-trained model based on the first sample audio data frame and the mixed phoneme state label corresponding to the first sample audio data frame, and the mixed phoneme state label includes the tri-phoneme state corresponding to the wake-up state and the single-phoneme state corresponding to the non-wake-up state. Therefore, only when the corresponding phoneme is in the wake-up state, it is necessary to label the tri-phoneme state, otherwise, only the single-phoneme state with fewer states needs to be labeled, thereby reducing the number of states output by the acoustic model, further reducing the resource requirements of the acoustic model, ensuring that the voice wake-up method can be applied to low-power devices, realizing a low-power voice wake-up method, and ultimately enhancing the application scope of the voice wake-up method; at the same time, the pre-trained model is trained based on the second sample audio data frame and the tri-phoneme state label corresponding to the second sample audio data frame. Therefore, the tri-phoneme state label is first used for training in the first training stage to utilize the fine-grained representation of the tri-phoneme, improve the sensitivity of the acoustic model to subtle changes in speech, and then only the classification layer is optimized in the second training stage, which can prevent the loss of information learned in the first training stage during the optimization process in the second training stage, thereby ensuring that the classification layer outputs fewer states while improving the performance of the acoustic model, and further improving the detection accuracy of voice wake-up.
[0185] Based on any of the above embodiments, the acoustic model is trained by the following model training module, and the model training module is configured to:
[0186] Train an initial model based on the second sample audio data frame and the tri-phoneme state label corresponding to the second sample audio data frame to obtain a pre-trained model;
[0187] Based on the first sample audio data frame and the corresponding hybrid phoneme state label of the first sample audio data frame, optimize the classification layer in the trained model to obtain the acoustic model.
[0188] Based on any of the above embodiments, the model training module is further configured to:
[0189] Obtain the second sample audio data frame and the corresponding triphone state label of the second sample audio data frame; the triphone state label includes M triphone states;
[0190] Change the triphone states corresponding to the non-awakening states in the triphone state label to monophone states to obtain the hybrid phoneme state label; the first sample audio data frame is the same as the second sample audio data frame;
[0191] Wherein, the hybrid phoneme state label includes N triphone states corresponding to awakening states and L monophone states corresponding to non-awakening states; N + L is less than M.
[0192] Based on any of the above embodiments, the model training module is further configured to:
[0193] Obtain the second sample audio data frame in the sample audio data and the corresponding text data of the sample audio data;
[0194] Based on the forced alignment result of the sample audio data and the text data, determine the triphone state label corresponding to the second sample audio data frame;
[0195] Wherein, the sample audio data set containing the sample audio data includes: general audio data and wake word audio data; the general audio data does not include audio data corresponding to a preset wake word, and the wake word audio data includes audio data corresponding to a preset wake word.
[0196] Based on any of the above embodiments, the model training module is further configured to:
[0197] Based on the second sample audio data frame, the corresponding triphone state label of the second sample audio data frame, and the binary classification result label corresponding to the second sample audio data frame, train the initial model to obtain a trained model;
[0198] Wherein, the binary classification result label includes a first classification result or a second classification result, the first classification result is used to represent that the second sample audio data frame belongs to the audio data frame corresponding to the end position of the wake word, and the second classification result is used to represent that the second sample audio data frame does not belong to the audio data frame corresponding to the end position of the wake word.
[0199] Based on any of the above embodiments, the phoneme output module 710 is further configured to:
[0200] Train an initial model based on the second sample audio data frame, the triphone state label corresponding to the second sample audio data frame, and the binary classification result label corresponding to the second sample audio data frame to obtain a trained model;
[0201] Wherein, the binary classification result label includes a first classification result or a second classification result. The first classification result is used to represent that the second sample audio data frame belongs to the audio data frame corresponding to the end position of the wake-up word, and the second classification result is used to represent that the second sample audio data frame does not belong to the audio data frame corresponding to the end position of the wake-up word.
[0202] Figure 8 An example of the physical structure diagram of an electronic device is shown as Figure 8 As shown, the electronic device may include: a processor 810, a communication interface 820, a memory 830, and a communication bus 840. Among them, the processor 810, the communication interface 820, and the memory 830 communicate with each other through the communication bus 840. The processor 810 can call the logical instructions in the memory 830 to execute the voice wake-up method, which includes: inputting each audio data frame in the audio data into an acoustic model to obtain a phoneme-level state sequence output by the acoustic model; the phoneme-level state sequence includes the phoneme-level state classification results of each audio data frame in the audio data, and the phoneme-level state classification results include single-phoneme states and triphone states; determining whether it is a voice wake-up state based on the phoneme-level state sequence; wherein, the acoustic model is obtained by optimizing the classification layer in the trained model based on the first sample audio data frame and the mixed phoneme state label corresponding to the first sample audio data frame; the trained model is trained based on the second sample audio data frame and the triphone state label corresponding to the second sample audio data frame; the mixed phoneme state label includes the triphone state corresponding to the wake-up state and the single-phoneme state corresponding to the non-wake-up state.
[0203] In addition, when the logical instructions in the above-mentioned memory 830 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.
[0204] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the voice wake-up method provided by the above-mentioned various methods. The method includes: inputting each audio data frame in the audio data into an acoustic model to obtain a phoneme-level state sequence output by the acoustic model; the phoneme-level state sequence includes the phoneme-level state classification results of each audio data frame in the audio data, and the phoneme-level state classification results include single-phoneme states and tri-phoneme states; based on the phoneme-level state sequence, determine whether it is a voice wake-up state; wherein, the acoustic model is obtained by optimizing the classification layer in the trained model based on the first sample audio data frame and the corresponding mixed phoneme state label of the first sample audio data frame; the trained model is trained based on the second sample audio data frame and the corresponding tri-phoneme state label of the second sample audio data frame; the mixed phoneme state label includes the tri-phoneme state corresponding to the wake-up state and the single-phoneme state corresponding to the non-wake-up state.
[0205] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the voice wake-up method provided by the above-mentioned various methods. The method includes: inputting each audio data frame in the audio data into an acoustic model to obtain a phoneme-level state sequence output by the acoustic model; the phoneme-level state sequence includes the phoneme-level state classification results of each audio data frame in the audio data, and the phoneme-level state classification results include single-phoneme states and tri-phoneme states; based on the phoneme-level state sequence, determine whether it is a voice wake-up state; wherein, the acoustic model is obtained by optimizing the classification layer in the trained model based on the first sample audio data frame and the corresponding mixed phoneme state label of the first sample audio data frame; the trained model is trained based on the second sample audio data frame and the corresponding tri-phoneme state label of the second sample audio data frame; the mixed phoneme state label includes the tri-phoneme state corresponding to the wake-up state and the single-phoneme state corresponding to the non-wake-up state.
[0206] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative labor.
[0207] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0208] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A voice wake-up method, characterized in that, Including: Inputting each audio data frame in the audio data into an acoustic model to obtain a phoneme-level state sequence output by the acoustic model; the phoneme-level state sequence includes the phoneme-level state classification results of each audio data frame in the audio data, and the phoneme-level state classification results include single-phoneme states and triphone states; Determining whether it is a voice wake-up state based on the phoneme-level state sequence; Wherein, the acoustic model is obtained by optimizing the classification layer in the trained model based on the first sample audio data frame and the corresponding mixed phoneme state label of the first sample audio data frame; the trained model is trained based on the second sample audio data frame and the corresponding triphone state label of the second sample audio data frame; the mixed phoneme state label includes the triphone state corresponding to the wake-up state and the single-phoneme state corresponding to the non-wake-up state, and the mixed phoneme state label only needs to retain the triphone state when the corresponding phoneme is in the wake-up state.
2. The voice wake-up method according to claim 1, wherein The acoustic model is trained based on the following method: Training an initial model based on the second sample audio data frame and the corresponding triphone state label of the second sample audio data frame to obtain a trained model; Optimizing the classification layer in the trained model based on the first sample audio data frame and the corresponding mixed phoneme state label of the first sample audio data frame to obtain the acoustic model.
3. The voice wake-up method according to claim 2, wherein Before training the initial model based on the second sample audio data frame and the corresponding triphone state label of the second sample audio data frame to obtain a trained model, it further includes: Obtaining the second sample audio data frame in the sample audio data and the corresponding triphone state label of the second sample audio data frame; the triphone state label includes M triphone states; Changing the triphone state corresponding to the non-wake-up state in the triphone state label to a single-phoneme state to obtain the mixed phoneme state label; the first sample audio data frame is the same as the second sample audio data frame; Wherein, the mixed phoneme state label includes N triphone states corresponding to the wake-up state and L single-phoneme states corresponding to the non-wake-up state; N + L is less than M.
4. The voice wake-up method according to claim 3, wherein The obtaining the second sample audio data frame in the sample audio data and the corresponding triphone state label of the second sample audio data frame includes: Obtaining the second sample audio data frame in the sample audio data and the corresponding text data of the sample audio data; Determining the triphone state label corresponding to the second sample audio data frame based on the forced alignment result of the sample audio data and the text data; Wherein, the sample audio data set containing the sample audio data includes: general audio data and wake-up word audio data; the general audio data does not include the audio data corresponding to the preset wake-up word, and the wake-up word audio data includes the audio data corresponding to the preset wake-up word.
5. The voice wake-up method according to claim 2, wherein The training the initial model based on the second sample audio data frame and the corresponding triphone state label of the second sample audio data frame to obtain a trained model includes: Training the initial model based on the second sample audio data frame, the corresponding triphone state label of the second sample audio data frame, and the binary classification result label corresponding to the second sample audio data frame to obtain a trained model; Wherein, the binary classification result label includes a first classification result or a second classification result, the first classification result is used to represent that the second sample audio data frame belongs to the audio data frame corresponding to the end position of the wake-up word, and the second classification result is used to represent that the second sample audio data frame does not belong to the audio data frame corresponding to the end position of the wake-up word.
6. The voice wake-up method according to claim 1, wherein, The step of inputting each audio data frame in the audio data into the acoustic model to obtain the phoneme-level state sequence output by the acoustic model includes: Inputting the acoustic features of each audio data frame in the audio data into the feature extraction layer in the acoustic model to obtain a sequence of feature vectors output by the feature extraction layer; Inputting the sequence of feature vectors into the classification layer in the acoustic model to obtain the phoneme-level state sequence output by the classification layer in the acoustic model.
7. A voice wake-up device, characterized in that, It includes: A phoneme output module for inputting each audio data frame in the audio data into the acoustic model to obtain the phoneme-level state sequence output by the acoustic model; The phoneme-level state sequence includes the phoneme-level state classification results of each audio data frame in the audio data, and the phoneme-level state classification results include single-phoneme states and triphone states; A state determination module for determining whether it is a voice wake-up state based on the phoneme-level state sequence; Wherein, the acoustic model is obtained by optimizing the classification layer in the trained model based on the first sample audio data frame and the corresponding mixed phoneme state label of the first sample audio data frame; the trained model is trained based on the second sample audio data frame and the corresponding triphone state label of the second sample audio data frame; the mixed phoneme state label includes the triphone state corresponding to the wake-up state and the single-phoneme state corresponding to the non-wake-up state, and the mixed phoneme state label only needs to retain the triphone state when the corresponding phoneme is in the wake-up state.
8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the voice wake-up method according to any one of claims 1 to 6.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the voice wake-up method according to any one of claims 1 to 6.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the voice wake-up method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Equipment awakening method and device and computer readable storage medium
CN116705015A
User-defined wake-up word recognition method and device of intelligent voice equipment and storage medium
CN118675503A