Wake-up model optimization method, wake-up method, device, equipment, medium and product
By optimizing the training method of the voice wake-up model in low-power scenarios, combining multi-level pseudo-label generation strategy and data augmentation technology, the problem of high false wake-up rate is solved, and the robustness and generalization ability of the model are improved.
Patent Information
- Application Number
- CN202510839293.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2045-06-23
AI Technical Summary
In low-power scenarios, the false wake-up rate of the voice wake-up model is high, and the pseudo-labels generated by existing semi-supervised learning have large deviations, resulting in poor performance of the wake-up model.
By obtaining the training subset in the sample dataset, combining the balanced model parameters of the current batch and the previous batch to generate frame-level and word-level wake-up pseudo-labels, perform semi-supervised training, and optimize the wake-up model using consistent regular loss and pseudo-label loss, construct a multi-level pseudo-label generation strategy to improve the quality of the pseudo-label and balance the training data distribution.
It effectively reduces the false wake-up rate in low-power scenarios, and enhances the robustness and generalization ability of the wake-up model in different scenarios.
Smart Images

Figure CN120340466B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of speech processing technology, and in particular to a wake-up model optimization method, a wake-up method, a device, equipment, a medium, and a product. Background Art
[0002] Voice wake-up technology plays a crucial role in the field of intelligent interaction and is widely used in smart homes, mobile devices, and other areas. Therefore, in the application scenarios of low-power voice wake-up technology, how to effectively achieve a low false wake-up rate is an important issue that the industry urgently needs to solve.
[0003] Related technologies use pseudo-labels generated during semi-supervised learning to train wakeup models, optimizing the false wakeup rate. However, in low-power scenarios, limited computing power leads to significant deviations in the pseudo-labels generated during semi-supervised learning. This results in poor performance of the trained wakeup model and a high false wakeup rate. Summary of the Invention
[0004] The present invention provides a wake-up model optimization method, a wake-up method, a device, a equipment, a medium and a product to solve the defects in the prior art.
[0005] The present invention provides a wake-up model optimization method, comprising:
[0006] For the current batch training, obtaining a training subset corresponding to the current batch training in the sample data set;
[0007] Generate a frame-level wakeup pseudo label for each speech frame in each first sample speech data in the training subset based on the weight parameters of the equalization model after training in the current batch and the average model parameters of the equalization model after training in the previous batch, and generate a word-level wakeup pseudo label for each speech word in each first sample speech data based on the frame-level wakeup pseudo label;
[0008] Performing semi-supervised training on the wakeup model trained in the previous batch according to each of the first sample speech data, the frame-level wakeup pseudo-label, the word-level wakeup pseudo-label, each of the second sample speech data in the training subset, and the frame-level wakeup true label of each speech frame and the word-level wakeup true label of each speech word in each of the second sample speech data, to obtain a wakeup model trained in the current batch;
[0009] An optimized wake-up model is determined based on the wake-up models trained after multiple batches.
[0010] According to a wake-up model optimization method provided by the present invention, the second sample speech data includes first enhanced speech sub-data corresponding to the marked sample speech data;
[0011] The first sample speech data includes second enhanced speech sub-data corresponding to the labeled sample speech data, and third enhanced speech sub-data and fourth enhanced speech sub-data corresponding to the unlabeled sample speech data;
[0012] Among them, the content invariance rate between the first enhanced speech sub-data and the labeled sample speech data is greater than a preset value, the content invariance rate between the second enhanced speech sub-data and the labeled sample speech data is less than or equal to the preset value, the content invariance rate between the third enhanced speech sub-data and the unlabeled sample speech data is greater than the preset value, and the content invariance rate between the fourth enhanced speech sub-data and the unlabeled sample speech data is less than or equal to the preset value.
[0013] According to a wake-up model optimization method provided by the present invention, the training step of the wake-up model after the current batch training includes:
[0014] Determining a first prediction loss based on each of the second sample speech data, and a frame-level true wakeup label of each speech frame and a word-level true wakeup label of each speech word in each of the second sample speech data;
[0015] Determining a second prediction loss based on each of the first sample speech data, and a frame-level wakeup pseudo label of each speech frame and a word-level wakeup pseudo label of each speech word in each of the first sample speech data;
[0016] Determining a consistency regularization loss based on each of the first enhanced speech sub-data and each of the third enhanced speech sub-data;
[0017] The wake-up model trained in the previous batch is trained according to the first prediction loss, the second prediction loss, and the consistency regularization loss to obtain the wake-up model trained in the current batch.
[0018] According to a wake-up model optimization method provided by the present invention, the wake-up model trained in the previous batch is trained according to the first prediction loss, the second prediction loss, and the consistency regularization loss to obtain the wake-up model trained in the current batch, including:
[0019] Determining a weighting coefficient of the consistency regularization loss according to the frame-level wake-up pseudo-label of each speech frame in each of the third enhanced speech sub-data;
[0020] performing weighted addition of the second prediction loss and the consistency regularization loss according to the weight coefficient of the consistency regularization loss to obtain an unsupervised loss;
[0021] Determining a target loss based on the unsupervised loss and the first prediction loss;
[0022] The wake-up model trained in the previous batch is trained according to the target loss to obtain the wake-up model trained in the current batch.
[0023] According to a wake-up model optimization method provided by the present invention, determining the weighting coefficient of the consistency regularization loss according to the frame-level wake-up pseudo-label of each speech frame in each of the third enhanced speech sub-data includes:
[0024] Determining a weighting coefficient of the consistency regularization loss according to the total number of target speech frames in the non-silent segments of all the third enhanced speech sub-data in the training subset and the total number of speech frames in all the third enhanced speech sub-data in the training subset;
[0025] The target speech frame is a speech frame whose probability value corresponding to the frame-level wake-up pseudo-label is greater than a probability threshold.
[0026] According to a wake-up model optimization method provided by the present invention, determining the target loss according to the unsupervised loss and the first prediction loss includes:
[0027] Determining a weighting coefficient for the unsupervised loss based on the number of first sample speech data in the training subset corresponding to the current batch training, the number of sample speech data to be labeled in the training subsets corresponding to all historical batch trainings before the current batch training, and the number of sample speech data to be labeled in the sample data set;
[0028] According to the weighting coefficient of the unsupervised loss, the unsupervised loss and the first prediction loss are weightedly added to obtain the target loss.
[0029] According to a wake-up model optimization method provided by the present invention, generating a word-level wake-up pseudo-label for each speech word in each first sample speech data according to the frame-level wake-up pseudo-label includes:
[0030] For each spoken word in each of the first sample speech data, when determining that the spoken word matches the target wake-up word based on the frame-level wake-up pseudo-labels of each speech frame in the spoken word, average the probability values corresponding to the frame-level wake-up pseudo-labels of all speech frames in each character in the spoken word to obtain a score for each character in the spoken word;
[0031] Calculating an average of the scores of all characters in the phonetic word to obtain the score of the phonetic word;
[0032] When the score of the speech word is greater than or equal to a score threshold, marking the word-level wake-up pseudo-label of the speech word as a wake-up word label;
[0033] Each voice word in each of the first sample voice data is traversed to obtain a word-level wake-up pseudo label for each voice word in each of the first sample voice data.
[0034] According to a wake-up model optimization method provided by the present invention, generating a frame-level wake-up pseudo label for each speech frame in each first sample speech data in the training subset based on the weight parameters of the equalization model after training in the current batch and the average model parameters of the equalization model after training in the previous batch includes:
[0035] Performing weighted addition on the weight parameters of the balanced model after the current batch training and the average model parameters of the balanced model after the previous batch training to obtain the average model parameters of the balanced model after the current batch training;
[0036] According to the average model parameters of the equalization model after training of the current batch, a label prediction is performed on each speech frame in each of the first sample speech data to obtain a frame-level wake-up pseudo label of each speech frame in each of the first sample speech data.
[0037] According to a wake-up model optimization method provided by the present invention, the step of determining the weight parameters of the equalization model after the current batch training includes:
[0038] According to each of the second sample voice data and the frame-level wake-up true label of each voice frame in each of the second sample voice data, the weight parameters of the equalization model after the previous batch training are optimized to obtain the weight parameters of the equalization model after the current batch training.
[0039] The present invention also includes a wake-up method, comprising:
[0040] Obtain the target voice data received by the device to be awakened;
[0041] Based on the optimized wake-up model, each speech frame in the target speech data is recognized to obtain a frame-level wake-up recognition result for each speech frame in the target speech data, and each speech word in the target speech data is recognized to obtain a word-level wake-up recognition result for each speech word in the target speech data;
[0042] Determining whether to wake up the device to be awakened according to the word-level wake-up recognition result and the frame-level wake-up recognition result;
[0043] The optimized wake-up model is obtained by optimizing the wake-up model based on any one of the above-mentioned wake-up model optimization methods.
[0044] The present invention also includes a wake-up model optimization device, comprising:
[0045] A first acquisition unit is configured to acquire, for a current batch of training, a training subset corresponding to the current batch of training from a sample data set;
[0046] a processing unit, configured to generate a frame-level wakeup pseudo label for each speech frame in each first sample speech data in the training subset based on the weight parameters of the equalization model after training in the current batch and the average model parameters of the equalization model after training in the previous batch, and generate a word-level wakeup pseudo label for each speech word in each first sample speech data based on the frame-level wakeup pseudo label;
[0047] a training unit, configured to perform semi-supervised training on the wakeup model trained in the previous batch based on the first sample speech data, the frame-level wakeup pseudo-label, the word-level wakeup pseudo-label, the second sample speech data in the training subset, and the frame-level wakeup true label of each speech frame and the word-level wakeup true label of each speech word in each second sample speech data, to obtain the wakeup model trained in the current batch;
[0048] The optimization unit is used to determine an optimized wake-up model based on the wake-up models trained after multiple batches.
[0049] The present invention also includes a wake-up device, comprising:
[0050] A second acquiring unit, configured to acquire target voice data received by the device to be awakened;
[0051] a recognition unit, configured to recognize each speech frame in the target speech data based on the optimized wake-up model, obtain a frame-level wake-up recognition result for each speech frame in the target speech data, and recognize each speech word in the target speech data, obtain a word-level wake-up recognition result for each speech word in the target speech data;
[0052] an awakening unit, configured to determine whether to wake up the device to be awakened according to the word-level awakening recognition result and the frame-level awakening recognition result;
[0053] The optimized wake-up model is obtained by optimizing the wake-up model based on any one of the above-mentioned wake-up model optimization methods.
[0054] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the wake-up model optimization method or the wake-up method described above is implemented.
[0055] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the wake-up model optimization method or the wake-up method as described above is implemented.
[0056] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements any one of the above-mentioned wake-up model optimization methods or wake-up methods.
[0057] The wake-up model optimization method, wake-up method, device, equipment, medium and product provided by the present invention obtain a training subset corresponding to the current batch training in the sample data set, and combine the balanced model weight parameters after the current batch training with the average model parameters of the previous batch to accurately generate the frame-level wake-up pseudo-label of the first sample voice data in the training subset, and then generate the word-level wake-up pseudo-label based on the frame-level wake-up pseudo-label integration. Thus, through parameter smoothing, multi-level pseudo-label generation strategy and multi-scale label generation strategy, the quality of the pseudo-label is improved and the deviation caused by limited computing power is reduced. The training data set of the wake-up model is constructed by combining the first sample voice data and its corresponding frame-level wake-up pseudo-label and word-level wake-up pseudo-label, as well as the second sample voice data and its corresponding frame-level wake-up true label and word-level wake-up true label, further balancing the distribution of the training data of the wake-up model and improving the training data quality of the wake-up model. Thus, the performance of the wake-up model is optimized by the training data set, which can effectively reduce its false wake-up rate in low-power scenarios and enhance its robustness and generalization ability in different wake-up scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0059] Figure 1 This is one of the flow charts of the wake-up model optimization method provided by the present invention.
[0060] Figure 2 It is a structural diagram of the equilibrium model provided by the present invention.
[0061] Figure 3 It is a structural diagram of the wake-up model provided by the present invention.
[0062] Figure 4 This is the second flow chart of the wake-up model optimization method provided by the present invention.
[0063] Figure 5 Schematic diagram of the wake-up model optimization data flow provided by the present invention.
[0064] Figure 6 Schematic diagram of the data augmentation data stream provided by the present invention.
[0065] Figure 7 It is a flowchart of the awakening method provided by the present invention.
[0066] Figure 8 It is a structural diagram of the wake-up model optimization device provided by the present invention.
[0067] Figure 9 It is a structural diagram of the awakening device provided by the present invention.
[0068] Figure 10 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0069] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0070] Voice wake-up technology plays a critical role in the field of intelligent interaction and is widely used in scenarios such as smart homes and mobile devices. In the application scenarios of low-power voice wake-up technology, how to ensure high recall rate and low false wake-up rate is an urgent problem that needs to be solved.
[0071] In the development of voice wake-up technology, relevant personnel have widely used the wake-up system based on the hidden Markov model to realize the wake-up control of the device. The wake-up system specifically receives voice data, outputs phonemes based on the acoustic characteristics of the voice data, and decodes the wake-up results. However, it has problems such as simplified state assumptions and poor adaptability to complex environments. It performs poorly in low signal-to-noise ratio and heavy accent tasks.
[0072] With the introduction of deep learning models, researchers have implemented speech wakeup recognition through supervised training of deep learning models, such as deep neural networks and their derivatives, convolutional neural networks and recurrent neural networks. This approach improves the accuracy and robustness of wakeup recognition by enabling the trained deep learning models to learn complex speech feature representations. However, this approach requires precise labeling of all possible positive and negative examples to guarantee the performance of the trained deep learning models, and thus the performance of wakeup recognition. However, achieving such comprehensive data coverage in real-world scenarios not only requires significant effort from professionals for data collection and labeling, but also the nearly infinite variety of possible false wakeup speech (negative examples) in real-world environments, making it difficult to fully collect all negative examples. Consequently, training high-performance deep learning models in such scenarios is difficult, resulting in poor wakeup recognition performance.
[0073] Given the open-set nature of false awakenings and the high cost of supervised learning annotation, researchers have proposed using semi-supervised learning to train wakeup models. This approach aims to leverage unlabeled speech data to improve performance and robustness, focusing on false awakening frequency and wakeup rate metrics, and optimizing the false awakening rate through the wakeup model trained accordingly. However, in low-power scenarios, limited computing power leads to significant deviations in the pseudo-labels generated by the wakeup model during semi-supervised learning, severely interfering with the training accuracy and stability of the wakeup model. This results in poor performance of the trained wakeup model and a still-high false awakening rate.
[0074] In this regard, in order to solve the problem of high false awakening rate of the awakening model in low power consumption scenarios, the present application provides a wake-up model optimization method.
[0075] Figure 1 This is one of the flow charts of the wake-up model optimization method provided by the present invention; Figure 1 As shown, the method includes step 110 , step 120 , step 130 and step 140 .
[0076] Step 110: For the current batch training, obtain a training subset corresponding to the current batch training in the sample data set;
[0077] Step 120: Generate a frame-level wakeup pseudo label for each speech frame in each first sample speech data in the training subset based on the weight parameters of the equalization model after training the current batch and the average model parameters of the equalization model after training the previous batch, and generate a word-level wakeup pseudo label for each speech word in each first sample speech data based on the frame-level wakeup pseudo label;
[0078] Step 130: Perform semi-supervised training on the wakeup model trained in the previous batch based on the first sample speech data, the frame-level wakeup pseudo-label, the word-level wakeup pseudo-label, the second sample speech data in the training subset, and the frame-level wakeup true label of each speech frame and the word-level wakeup true label of each speech word in each second sample speech data to obtain the wakeup model trained in the current batch;
[0079] Step 140 : Determine an optimized wake-up model based on the wake-up models trained in multiple batches.
[0080] The balancing model (hereinafter referred to as the frame-level category balancing model) is primarily used to address the problem of frame-level labels being biased towards the in-set state due to data imbalance during wakeup model training. Specifically, the balancing model generates frame-level wakeup pseudo-labels for the sample speech data to be labeled. Based on the frame-level pseudo-labels generated by the balancing model, word-level pseudo-labels can be further generated. By introducing a large number of frame-level and word-level pseudo-labels, the training data distribution of the wakeup model can be effectively balanced, thereby effectively alleviating the impact of uneven training data distribution on wakeup model training and improving the performance and generalization of the wakeup model.
[0081] The equalization model includes at least a frame-level wakeup recognition module with a frame-level wakeup recognition function for speech. In addition to the frame-level wakeup recognition module, the wakeup model also includes a word-level wakeup recognition module with a word-level wakeup recognition function for speech. The frame-level wakeup recognition module can be a model prepared for frame-level wakeup recognition after parameter initialization, or a pre-trained model with a frame-level wakeup recognition function; the word-level wakeup recognition module can also be a model prepared for word-level wakeup recognition after parameter initialization, or a pre-trained model with a word-level wakeup recognition function. This embodiment does not specifically limit this.
[0082] Figure 2 It is a structural diagram of the equilibrium model provided by the present invention; Figure 3 It is a structural diagram of the wake-up model provided by the present invention.
[0083] Here, the frame-level wake-up recognition module and the word-level wake-up recognition module can be constructed based on various deep learning models. For example, the frame-level wake-up recognition module is constructed based on the encoder and the first decoder, and the word-level wake-up recognition module is constructed based on the second decoder. Figure 2 and Figure 3As shown, the equalization model is constructed based on a sequentially connected encoder and first decoder, while the wakeup model is constructed based on an encoder, a first decoder, and a second decoder. The encoder is used to encode speech features; the first decoder is used to perform frame-level wakeup recognition based on speech features; and the second decoder is used to perform word-level wakeup recognition based on speech features.
[0084] To simplify the description, Figure 2 The equilibrium model shown, and Figure 3 Taking the wake-up model shown as an example, the method provided by this embodiment is described in detail.
[0085] Optionally, when the wake-up model needs to be optimized to reduce its false wake-up rate in a low-power scenario, the following steps may be performed to iteratively train the wake-up model:
[0086] First, corpus data is collected, which specifically includes supervised corpus data and unsupervised corpus data; the supervised corpus data is the corpus data required for supervised learning, and the unsupervised corpus data is the corpus data required for unsupervised learning.
[0087] Supervised corpus data specifically includes wake-up word data and non-wake-up word data. Wake-up word data refers to corpus data containing wake-up words, focusing on commonly used wake-up words and comprehensively covering pronunciations across different genders, user groups (e.g., seniors, adults, and children), speaking speeds, accents, and background noise levels, ensuring rich speech features in wake-up scenarios. For example, wake-up word pronunciations from children in various emotional states, such as excitement, calmness, and fatigue, can be collected to generate different wake-up word data sets. Non-wake-up word data refers to corpus data that does not contain wake-up words and must comprehensively cover a variety of pronunciations in the target language for the wake-up model, such as across different genders, age groups, speaking speeds, accents, and background noise levels, ensuring rich speech features in non-wake-up scenarios. Unsupervised corpus data focuses on common corpus types such as conversations, interviews, TV dramas, and audio novels, and comprehensively covers diverse themes, styles, and scenarios.
[0088] It should be noted that since the collected supervised corpus needs to be manually labeled before use, while unsupervised data does not require manual labeling, the amount of collected unsupervised corpus data should be greater than the amount of collected supervised corpus data.
[0089] Then, in order to increase the adaptability to the scene, the supervised corpus data and unsupervised corpus data can be replayed and recorded through the target wake-up device in the wake-up scene targeted by the wake-up model to obtain the playback corpus data in the real scene; in addition, simulation tools can be used to perform various data augmentation operations such as noise addition and speech speed change on the supervised corpus data and unsupervised corpus data to obtain enhanced corpus data.
[0090] Then, each supervised corpus data, the corresponding playback corpus data, and the enhanced corpus data are processed into a unified format, and the processed corpus data are labeled with word-level wake-up labels and frame-level wake-up labels to obtain the labeled sample speech data and the frame-level wake-up true labels of each speech frame and the word-level wake-up true labels of each speech word in the labeled sample speech data. Furthermore, each unsupervised corpus data, the corresponding playback corpus data, and the enhanced corpus data are processed into a unified format to obtain the unlabeled sample speech data.
[0091] Then, multi-dimensional data augmentation is performed on each labeled sample speech data and each unlabeled sample speech data to obtain sample speech data to be labeled and labeled sample speech data, as well as the frame-level true wakeup label of each speech frame and the word-level true wakeup label of each speech word in the labeled sample speech data. The sample speech data to be labeled includes sample speech data obtained by performing a strong augmentation operation on each labeled sample speech data (e.g., an augmentation operation in which the content invariance rate before and after augmentation is less than or equal to a preset value), and sample speech data obtained by performing a strong augmentation operation and / or a weak augmentation operation on each unlabeled sample speech data (e.g., an augmentation operation in which the content invariance rate before and after augmentation is greater than a preset value); the labeled sample speech data includes sample speech data obtained by performing a weak augmentation operation on each labeled sample speech data, and the frame-level true wakeup label of each speech frame and the word-level true wakeup label of each speech word in the labeled sample speech data are determined based on the corresponding frame-level true wakeup label of each speech frame and the word-level true wakeup label of each speech word in the labeled sample speech data.
[0092] Then, a sample data set is constructed based on the sample speech data to be labeled and the labeled sample speech data, as well as the frame-level wake-up true label of each speech frame and the word-level wake-up true label of each speech word in each labeled sample speech data.
[0093] After obtaining the sample data set, the balancing model and the wake-up model can be simultaneously optimized and trained based on the sample data set to obtain an optimized wake-up model. The specific training steps are as follows:
[0094] For the current batch training, first obtain the sample speech data to be labeled required for the current batch training as the first sample speech data in the sample data set, and obtain the labeled sample speech data required for the current batch training as the second sample speech data to form a training subset corresponding to the current batch training.
[0095] After obtaining the training subset corresponding to the current batch training, the weight parameters of the equalization model after the current batch training and the average model parameters of the equalization model after the previous batch training can be loaded, and the weight parameters of the equalization model after the current batch training and the average model parameters of the equalization model after the previous batch training can be combined to perform frame-level voice wake-up recognition on each voice frame in each first sample voice data to generate a frame-level wake-up pseudo label for each voice frame in each first sample voice data.
[0096] It should be noted that in the frame-level speech wake-up recognition process, the weight parameters of the equalization model after training the current batch and the average model parameters after training the previous batch can be fused to obtain the fused model parameters, and then the fused model parameters are used to perform wake-up recognition on the speech frames in each first sample speech data to generate frame-level wake-up pseudo labels for the speech frames in each first sample speech data; or, the speech frames in each first sample speech data are wake-up recognized based on the weight parameters of the equalization model after training the current batch and the average model parameters of the equalization model after training the previous batch to obtain their respective label prediction results, and then these label prediction results are fused to finally generate frame-level wake-up pseudo labels for the speech frames in each first sample speech data, etc. This embodiment does not make specific restrictions on this. The fusion here can be weighted addition fusion, or dynamic parameter fusion based on the attention mechanism, etc. This embodiment does not make specific restrictions on this.
[0097] After obtaining the frame-level wakeup pseudo-label for each speech frame in each first sample speech data, a word-level wakeup pseudo-label for each speech word in each first sample speech data can be generated based on the frame-level wakeup pseudo-label for each speech frame in each first sample speech data. In this process, the frame-level wakeup pseudo-labels for each speech frame within each word can be counted and integrated to determine the word-level wakeup pseudo-label for the speech word, etc., which is not specifically limited in this embodiment.
[0098] After obtaining multi-scale pseudo labels (i.e., frame-level wakeup pseudo labels and word-level wakeup pseudo labels) for each first speech sample, the wakeup model trained in the previous batch can be semi-supervisedly trained using the first speech sample and its corresponding multi-scale pseudo labels, as well as the second speech sample and its corresponding multi-scale true labels (i.e., frame-level true labels and word-level true labels). This results in a wakeup model trained in the current batch. During the semi-supervised training process, a consistency regularization loss and a pseudo-label loss can be calculated based on each first speech sample and its corresponding multi-scale pseudo-label, while a supervision loss can be calculated based on each second speech sample and its corresponding multi-scale true label. The wakeup model trained in the previous batch is then updated based on the consistency regularization loss, pseudo-label loss, and supervision loss, thereby obtaining the wakeup model trained in the current batch. For example, a fusion operation such as direct addition or weighted addition of the consistency regularization loss, pseudo-label loss, and supervision loss can be performed to obtain a target loss. This target loss can then be used to update the wakeup model trained in the previous batch through backpropagation to obtain the wakeup model trained in the current batch.
[0099] After obtaining the wake-up model after the current batch training, determine whether the current batch training has reached the preset termination condition, such as whether the training rounds corresponding to the current batch training have reached the maximum training rounds, or whether the wake-up model after the current batch training has reached convergence, etc. If the preset termination condition is not reached, return to step 110 and iterate for the next iterative training until it is determined that the model training has reached the preset termination condition, and then stop training.
[0100] After stopping training, the wake-up model with the best model performance (such as the lowest false wake-up rate) is selected from the wake-up models trained in multiple batches as the optimized wake-up model, so that the optimized wake-up model obtained thereby has a low false wake-up rate in a low power consumption scenario.
[0101] The method provided in this embodiment obtains a training subset corresponding to the current batch training from a sample data set, and combines the weight parameters of the balanced model after the current batch training with the average model parameters of the previous batch to accurately generate a frame-level wake-up pseudo-label for the first sample speech data in the training subset. The word-level wake-up pseudo-label is then integrated based on the frame-level wake-up pseudo-label to generate a word-level wake-up pseudo-label. This improves the quality of the pseudo-label and reduces the deviation caused by limited computing power through parameter smoothing, a multi-level pseudo-label generation strategy, and a multi-scale label generation strategy. The training dataset for the wake-up model is constructed by combining the first sample speech data and its corresponding frame-level wake-up pseudo-label and word-level wake-up pseudo-label, as well as the second sample speech data and its corresponding frame-level wake-up true label and word-level wake-up true label. This further balances the distribution of the training data for the wake-up model and improves the quality of the training data for the wake-up model. The performance of the wake-up model is optimized through the training dataset, effectively reducing its false wake-up rate in low-power scenarios and enhancing its robustness and generalization capabilities in different wake-up scenarios.
[0102] In some embodiments, the second sample speech data includes first enhanced speech sub-data corresponding to the labeled sample speech data; the first sample speech data includes second enhanced speech sub-data corresponding to the labeled sample speech data, and third enhanced speech sub-data and fourth enhanced speech sub-data corresponding to the unlabeled sample speech data; wherein, the content invariance rate between the first enhanced speech sub-data and the labeled sample speech data is greater than a preset value, the content invariance rate between the second enhanced speech sub-data and the labeled sample speech data is less than or equal to the preset value, the content invariance rate between the third enhanced speech sub-data and the unlabeled sample speech data is greater than the preset value, and the content invariance rate between the fourth enhanced speech sub-data and the unlabeled sample speech data is less than or equal to the preset value.
[0103] Optionally, the training subset corresponding to each batch of training can be obtained by performing strong augmentation and weak augmentation on the labeled sample speech data and the unlabeled sample speech data, respectively, so as to determine different losses later based on different augmentation strategies and different data sources. The strong augmentation and weak augmentation are obtained by applying different augmentation strengths to the speech data under different augmentation types, and then dividing them according to the degree of their influence on the content of the speech data. For example, the content invariance rate corresponding to the strong augmentation category is less than or equal to the preset value, while the content invariance rate corresponding to the weak augmentation category is greater than the preset value. The preset value here can be set according to actual needs, such as 95%, etc., and this embodiment does not make specific restrictions on this. The augmentation types here include but are not limited to noise addition, time domain masking, and frequency domain masking, and this embodiment does not make specific restrictions on this.
[0104] The strong augmentation and weak augmentation techniques mentioned in this embodiment are described below with reference to specific augmentation strength classification rules.
[0105] Table 1 Distribution of the impact of different data augmentation methods on content
[0106]
[0107] As shown in Table 1, when the speech data is augmented by adding noise with a signal-to-noise ratio of 20dB, 10dB, and 5dB, the corresponding content invariance rates are all greater than 95%, so the augmentation category to which it belongs is weak augmentation; when the speech data is augmented by adding noise with a signal-to-noise ratio of 0dB and -5dB, the corresponding content invariance rates are all less than 95%, so the augmentation category to which it belongs is strong augmentation; when the speech data is augmented by a time domain mask with a mask (also called a mask) duration of 40ms, the corresponding content invariance rates are all less than 95%, so the augmentation category to which it belongs is strong augmentation. If the corresponding content invariance rate is greater than 95%, the augmentation category to which it belongs is weak augmentation; when the speech data is augmented using a time domain mask with a mask duration of 80ms and 160ms, since the corresponding content invariance rate is less than or equal to 95%, the augmentation category to which it belongs is strong augmentation; when the speech data is augmented using a frequency domain mask with a mask range of 0.5k, 1k, and 2k, since the corresponding content invariance rates are all greater than 95%, the augmentation category to which it belongs is weak augmentation.
[0108] In summary, the specific steps for obtaining the training subset corresponding to the current batch training include:
[0109] The labeled sample speech data is weakly augmented to obtain the first enhanced speech sub-data. Specifically, the data can be obtained by using at least one of the multiple augmentation methods corresponding to the weak augmentation categories shown in Table 1, which will not be described in detail here.
[0110] The labeled sample speech data is strongly augmented to obtain the second enhanced speech sub-data. Specifically, the second enhanced speech sub-data can be obtained by at least one augmentation method among the multiple augmentation methods corresponding to the strong augmentation categories shown in Table 1, which will not be described in detail here.
[0111] The unlabeled sample speech data is weakly augmented to obtain the third enhanced speech sub-data. Specifically, the data can be obtained by at least one of the multiple augmentation methods corresponding to the weak augmentation categories shown in Table 1, which will not be described in detail here.
[0112] The unlabeled sample speech data is strongly augmented to obtain the fourth enhanced speech sub-data, which can be obtained by implementing at least one augmentation method among the multiple augmentation methods corresponding to the strong augmentation categories shown in Table 1, which will not be repeated here.
[0113] Then, the second sample speech data is determined based on the first enhanced speech sub-data, and the first sample speech data is determined based on the second enhanced speech sub-data, the third enhanced speech sub-data and the fourth enhanced speech sub-data. Thus, a large number of negative samples and positive samples for supervised learning are provided for the wake-up model training through the weak augmented samples of the labeled sample speech data. A large number of negative samples and positive samples for unsupervised learning are provided for the wake-up model training through the strong augmented samples and weak augmented samples of the unlabeled sample speech data and the strong augmented samples of the labeled sample speech data. Accordingly, the robustness and generalization ability of the wake-up model in different scenarios are improved, its performance is optimized, and its false awakening rate is reduced.
[0114] Figure 4 This is the second flow chart of the wake-up model optimization method provided by the present invention; Figure 4 As shown, the method includes step 410 , step 420 , step 430 and step 440 .
[0115] Step 410: For the current batch training, obtain a training subset corresponding to the current batch training in the sample data set;
[0116] Step 420: Generate a frame-level wakeup pseudo label for each speech frame in each first sample speech data in the training subset based on the weight parameters of the equalization model after training the current batch and the average model parameters of the equalization model after training the previous batch, and generate a word-level wakeup pseudo label for each speech word in each first sample speech data based on the frame-level wakeup pseudo label;
[0117] Step 430: Determine a first prediction loss based on each second sample speech data, and the frame-level wake-up true label of each speech frame and the word-level wake-up true label of each speech word in each second sample speech data; determine a second prediction loss based on each first sample speech data, and the frame-level wake-up pseudo label of each speech frame and the word-level wake-up pseudo label of each speech word in each first sample speech data; determine a consistency regularization loss based on each first enhanced speech sub-data and each third enhanced speech sub-data; train the wake-up model trained in the previous batch based on the first prediction loss, the second prediction loss, and the consistency regularization loss to obtain the wake-up model trained in the current batch;
[0118] Step 440 : Determine an optimized wake-up model based on the wake-up models trained in multiple batches.
[0119] Optionally, for the current batch training, the sample speech data to be labeled required for the current batch training can be obtained from the sample data set as the first sample speech data, and the labeled sample speech data required for the current batch training can be obtained as the second sample speech data to form a training subset corresponding to the current batch training. The specific implementation steps can be found in step 110 and will not be repeated here.
[0120] After obtaining the training subset corresponding to the current batch training, the weight parameters of the equalization model after the current batch training and the average model parameters of the equalization model after the previous batch training can be loaded, and the weight parameters of the equalization model after the current batch training and the average model parameters of the equalization model after the previous batch training can be combined to perform frame-level voice wake-up recognition on each voice frame in each first sample voice data to generate a frame-level wake-up pseudo-label for each voice frame in each first sample voice data, and based on the frame-level wake-up pseudo-label of each voice frame in each first sample voice data, generate a word-level wake-up pseudo-label for each voice word in each first sample voice data. The specific implementation steps can be found in step 120, which will not be repeated here.
[0121] Figure 5 Schematic diagram of the wake-up model optimization data flow provided by the present invention. Figure 5 As shown, after obtaining the frame-level wake-up pseudo label of each speech frame and the word-level wake-up pseudo label of each speech word in each first sample speech data, the first sample speech data and its corresponding frame-level wake-up pseudo label and word-level wake-up pseudo label, as well as the second sample speech data and its corresponding frame-level wake-up true label and word-level wake-up true label, can be combined to perform semi-supervised training on the wake-up model trained in the previous batch to obtain the wake-up model trained in the current batch. The specific implementation steps include:
[0122] Each second sample speech data is input into the wakeup model trained in the previous batch. The wakeup model trained in the previous batch performs frame-level and word-level wakeup prediction on each second sample speech data to obtain a frame-level wakeup prediction label for each speech frame and a word-level wakeup prediction label for each speech word in each second sample speech data. A frame-level supervision loss is determined based on the error between the frame-level wakeup prediction label and the true frame-level wakeup label for each speech frame in each second sample speech data. A word-level supervision loss is also determined based on the error between the word-level wakeup prediction label and the true word-level wakeup label for each speech word in each second sample speech data. The first prediction loss (also called supervision loss) is then calculated by combining the frame-level supervision loss and the word-level supervision loss. The loss functions used for the frame-level supervision loss and the word-level supervision loss can be set according to actual needs, such as the maximum pooling loss function (also called MaxpoolingLoss) or the cross-entropy loss function.
[0123] Figure 6 This is a schematic diagram of the data augmentation data flow provided by the present invention; Figure 6 As shown in Figure 2, when pseudo label loss and consistency regularization loss are used, the following steps can be performed:
[0124] Each first-sample speech data set is input into the wakeup model trained in the previous batch. The model then performs frame-level and word-level wakeup predictions on each first-sample speech data set to obtain a frame-level wakeup prediction label for each speech frame and a word-level wakeup prediction label for each speech word in each first-sample speech data set. A frame-level unsupervised loss is determined based on the error between the frame-level wakeup prediction label and the frame-level wakeup pseudo-label for each speech frame in each first-sample speech data set. A word-level unsupervised loss is also determined based on the error between the word-level wakeup prediction label and the word-level pseudo-label for each speech word in each second-sample speech data set. The second prediction loss (also known as the pseudo-label loss) is then calculated by combining the frame-level and word-level unsupervised losses. The loss functions used for the word-level and frame-level unsupervised losses can also be configured based on actual needs, such as a maximum pooling loss function or a cross-entropy loss function.
[0125] It's also worth noting that in the field of semi-supervised learning, self-training methods were often used early on, with model predictions as the training objective. However, these methods were prone to error accumulation, particularly in imbalanced tasks. Subsequently, pseudo-label generation methods based on threshold screening emerged. With the integration of deep learning and semi-supervised training, consistency regularization techniques were introduced, encompassing data augmentation and adversarial training strategies. These techniques encourage models to learn invariance to small input perturbations and enhance generalization capabilities. However, existing technologies fail to consider the failure of the consistency assumption, resulting in poor model training results.
[0126] In this regard, this embodiment uses highly reliable weakly augmented sample data to calculate the consistency regularization loss when calculating the consistency regularization loss, so as to alleviate the problem of consistency assumption failure and thus improve the model training effect. The specific implementation steps are as follows: based on each first enhanced speech sub-data and each third enhanced speech sub-data, determine the consistency regularization loss; for example, the original labeled sample speech data and its corresponding first enhanced speech sub-data are input into the model respectively to obtain two sets of label prediction results; similarly, the unlabeled sample speech data and its corresponding third enhanced speech sub-data are input into the model respectively to obtain two sets of label prediction results. For each pair of label prediction results of the original data and its enhanced data, the consistency difference between them is calculated, and all the consistency difference results are summarized to obtain the consistency regularization loss.
[0127] After obtaining the first prediction loss, second prediction loss, and consistency regularization loss based on the above steps, the first prediction loss, second prediction loss, and consistency regularization loss are fused to construct a comprehensive target loss. The target loss is then used to perform backpropagation training on the wake-up model trained in the previous batch to obtain the wake-up model trained in the current batch. The fusion can be performed by operations such as addition or weighted addition, which is not specifically limited in this embodiment.
[0128] After obtaining the wake-up model after the current batch training, determine whether the current batch training has reached the preset termination condition, such as whether the training rounds corresponding to the current batch training have reached the maximum training rounds, or whether the wake-up model after the current batch training has reached convergence, etc. If the preset termination condition is not reached, return to step 410 and perform the next iterative training until it is determined that the model training has reached the preset termination condition, and then stop training.
[0129] After stopping training, a trained wake-up model with the best model performance (such as the lowest false wake-up rate) is selected from the wake-up models trained in multiple batches as the optimized wake-up model, so that the optimized wake-up model has a low false wake-up rate in a low power consumption scenario.
[0130] The method provided in this embodiment accurately generates multi-scale pseudo labels by combining multi-batch model parameters. It also constructs a comprehensive target loss based on the supervision loss of labeled data, the pseudo-label loss of unlabeled data, and the consistency regularization loss of weakly augmented data to iteratively optimize the wake-up model. This not only improves the training data quality of the wake-up model, effectively reduces the false wake-up rate of the optimized wake-up model in low-power scenarios, enhances the generalization ability of the optimized wake-up model, enables it to maintain stable performance in complex speech environments, but also alleviates the problem of consistency assumption failure, thereby improving the robustness of the optimized wake-up model.
[0131] In one possible implementation, in step 430, the wake-up model trained in the previous batch is trained according to the first prediction loss, the second prediction loss, and the consistency regularization loss, and the steps of obtaining the wake-up model trained in the current batch specifically include: step 431, step 432, step 433, and step 434.
[0132] Step 431: determining a weighting coefficient of the consistency regularization loss according to the frame-level wakeup pseudo-label of each speech frame in each of the third enhanced speech sub-data;
[0133] Step 432: performing weighted addition of the second prediction loss and the consistency regularization loss according to the weight coefficient of the consistency regularization loss to obtain an unsupervised loss;
[0134] Step 433: determining a target loss based on the unsupervised loss and the first predicted loss;
[0135] Step 434 : Train the wake-up model trained in the previous batch according to the target loss to obtain the wake-up model trained in the current batch.
[0136] like Figure 6 As shown, after obtaining the first prediction loss, the second prediction loss, and the consistency regularization loss, the following steps can be performed to train the wake-up model after the previous batch training:
[0137] First, the confidence level of the frame-level wakeup pseudo-label for each speech frame in each third enhanced speech sub-data is analyzed. The weighting coefficient of the consistency regularization loss is determined based on the pseudo-label confidence analysis results. For example, if the pseudo-label reliability is high, the weight of the consistency regularization loss can be appropriately increased to make it play a greater role in the training process; conversely, if the pseudo-label reliability is low, its weight can be appropriately reduced.
[0138] In a possible implementation, the step of determining the weighted coefficient of the consistency regularization loss based on the frame-level wakeup pseudo-label of each speech frame in each of the third enhanced speech sub-data specifically includes:
[0139] Determining a weighting coefficient of the consistency regularization loss according to the total number of target speech frames in the non-silent segments of all the third enhanced speech sub-data in the training subset and the total number of speech frames in all the third enhanced speech sub-data in the training subset;
[0140] The target speech frame is a speech frame whose probability value corresponding to the frame-level wake-up pseudo-label is greater than a probability threshold.
[0141] Optionally, when determining the weighted coefficient of the consistency regularization loss, the weighted coefficient of the consistency regularization loss can be dynamically adjusted based on the probability value of the pseudo-label to prevent the awakening model from learning incorrect data features. This allows the awakening model to better adapt to different data characteristics during training, avoids the problem of data augmentation changing the audio content, and enhances the generalization ability and stability of the awakening model. The specific determination steps include:
[0142] In the non-silent segments of each third enhanced speech sub-data, a target speech frame having a probability value corresponding to a frame-level wake-up pseudo-label greater than a probability threshold is determined, and the total number of target speech frames included in the non-silent segments of all third enhanced speech sub-data in the training subset is determined to obtain a first total number of frames; and the total number of speech frames included in all third enhanced speech sub-data in the training subset is determined to obtain a second total number of frames; then, a weighting coefficient of the consistency regularization loss is determined by combining the first total number of frames and the second total number of frames. For example, the weighting coefficient of the consistency regularization loss is determined by calculating the ratio between the first total number of frames and the second total number of frames. The specific calculation formula is as follows:
[0143] ;
[0144] in, is the weighted coefficient of consistency regularization loss; For the non-silent segment The probability value corresponding to the frame-level wakeup pseudo label of the speech frame; is a set of numbers of speech frames in non-silent segments of all third enhanced speech sub-data in the training subset; is the total number of all speech frames contained in all third enhanced speech sub-data in the training subset; is a probability threshold, which can be set according to actual needs, such as 0.8 or 0.9, etc. This embodiment does not make any specific limitation on this.
[0145] After obtaining the weighted coefficient of the consistency regularization loss through the above steps, the second prediction loss and the consistency regularization loss can be weighted added according to the weighted coefficient of the consistency regularization loss to obtain the unsupervised loss , the specific calculation formula is as follows:
[0146] ;
[0147] in, is the frame-level unsupervised loss in the second prediction loss, is the word-level unsupervised loss in the second prediction loss, is the consistency regularization loss; is the weighted coefficient of the consistency regularization loss.
[0148] After obtaining the unsupervised loss, the unsupervised loss and the first prediction loss can be fused to obtain the target loss. The fusion here can be direct addition or weighted addition.
[0149] In a possible implementation, the step of determining the target loss based on the unsupervised loss and the first predicted loss specifically includes:
[0150] Determining a weighting coefficient for the unsupervised loss based on the number of first sample speech data in the training subset corresponding to the current batch training, the number of sample speech data to be labeled in the training subsets corresponding to all historical batch trainings before the current batch training, and the number of sample speech data to be labeled in the sample data set;
[0151] According to the weighting coefficient of the unsupervised loss, the unsupervised loss and the first prediction loss are weightedly added to obtain the target loss.
[0152] Optionally, when calculating the target loss, the weighting coefficient of the unsupervised loss can be determined by combining the number of the first sample speech data in the training subset, the number of sample speech data to be labeled in the training subset corresponding to all historical batch trainings, and the number of sample speech data to be labeled in the sample data set. For example, the sum of the number of the first sample speech data in the training subset and the number of sample speech data to be labeled in the training subset corresponding to all historical batch trainings can be calculated to obtain the total number of sample speech data to be labeled traversed up to the current batch training, and then the ratio of the total number to the total number of all sample speech data to be labeled in the sample data set can be calculated to obtain the weighting coefficient of the unsupervised loss. The specific calculation formula is as follows:
[0153] ;
[0154] Where, is the weighted coefficient of unsupervised loss; is the sum of the number of all sample speech data to be labeled in the sample dataset; The total number of sample speech data to be labeled that have been traversed by the end of the current batch training.
[0155] After obtaining the weighted coefficient of the unsupervised loss, the unsupervised loss and the first prediction loss can be weighted and added according to the weighted coefficient to obtain the target loss. The specific calculation formula is as follows:
[0156] ;
[0157] in, is the first prediction loss (also called supervision loss); is the unsupervised loss. It is the weighting coefficient of the unsupervised loss, which alleviates the problem of low credibility of pseudo-labels of sample speech data expected to be labeled at the beginning of balanced training through gradual changes, thereby improving the accuracy, robustness and stability of wake-up recognition of the optimized wake-up model trained based on it.
[0158] After obtaining the target loss, the wake-up model trained in the previous batch can be back-propagated according to the target loss to obtain the wake-up model trained in the current batch.
[0159] The method provided in this embodiment dynamically determines the weighted coefficient of the consistency regularization loss by analyzing the pseudo-label confidence, and dynamically adjusts the weighted coefficient of the unsupervised loss based on the traversed data. In this way, the target loss is adaptively constructed through dynamic weights, and the wake-up model is trained accordingly. This effectively alleviates the problem of consistency assumption failure, improves the adaptability of the wake-up model to different data characteristics, reduces the false wake-up rate, and enhances the generalization ability, robustness, and stability of the wake-up model in low-power scenarios.
[0160] In one possible implementation, the step of generating a word-level wake-up pseudo-label for each speech word in each first sample speech data based on the frame-level wake-up pseudo-label specifically includes: for each speech word in each first sample speech data, when determining that the speech word matches the target wake-up word based on the frame-level wake-up pseudo-label of each speech frame in the speech word, the probability values corresponding to the frame-level wake-up pseudo-labels of all speech frames in each word in the speech word are averaged to obtain the score of each word in the speech word; the scores of all words in the speech word are averaged to obtain the score of the speech word; when the score of the speech word is greater than or equal to the score threshold, the word-level wake-up pseudo-label of the speech word is marked as the wake-up word label; and each speech word in each first sample speech data is traversed to obtain the word-level wake-up pseudo-label of each speech word in each first sample speech data.
[0161] Optionally, when determining a word-level wakeup pseudo-label for each speech word in each first sample speech data, the following steps may be specifically performed:
[0162] Based on the frame-level wake-up pseudo-label of each speech frame in the speech word, it is determined whether the speech word matches the target wake-up word. If not, the word-level wake-up pseudo-label of the speech word is directly marked as a non-wake-up word label; if they match, the probability values corresponding to the frame-level wake-up pseudo-labels of all speech frames of each word in the speech word are first averaged to obtain the score of each word in the speech word; then the scores of all words in the speech word are averaged to obtain the score of the speech word; then, the score of the speech word is compared with the score threshold to determine whether to mark the speech word with a wake-up word label; if the score of the speech word is greater than or equal to the score threshold, it is determined to be marked as a wake-up word label (such as 1); if the score of the speech word is less than the score threshold, the speech word is not marked with a wake-up word label, but is marked as a non-wake-up word label (such as 0), thereby ensuring that only speech words with high scores and meeting the wake-up conditions are marked as wake-up word labels, thereby improving the quality of pseudo-labels.
[0163] Among them, the score of the phonetic word The calculation formula is as follows:
[0164] ;
[0165] Where, is the number of characters in the phonetic word; The first The number of frames of the word; The first The first word The probability value corresponding to the frame-level wakeup pseudo label of the speech frame; Indicates the wake-up word label determined based on the frame-level wake-up pseudo-label of each speech frame in the speech word With target wake word match.
[0166] The word-level wakeup pseudo-label of the spoken word The calculation formula is as follows:
[0167] ;
[0168] Where, is the score threshold, which can be set according to actual needs, such as 0.8 or 0.9, etc. This implementation does not make a specific limitation on this.
[0169] Following the above word-level wakeup pseudo-label generation process, each word in each first sample speech data is traversed to obtain a word-level wakeup pseudo-label for each word in each first sample speech data. Thus, through step-by-step matching and threshold comparison, the word-level wakeup pseudo-label for each word in each first sample speech data is accurately identified and labeled. This effectively ensures the accuracy and effectiveness of the pseudo-labels at different scales, improves the quality of the wakeup model training data, and thus improves the performance of the trained wakeup model and reduces the false wakeup rate.
[0170] In one possible implementation, the step of generating a frame-level wake-up pseudo-label for each speech frame in each first sample speech data in the training subset based on the weight parameters of the equalization model after training the current batch and the average model parameters of the equalization model after training the previous batch specifically includes: weighted addition of the weight parameters of the equalization model after training the current batch and the average model parameters of the equalization model after training the previous batch to obtain the average model parameters of the equalization model after training the current batch; and performing label prediction on each speech frame in each first sample speech data based on the average model parameters of the equalization model after training the current batch to obtain a frame-level wake-up pseudo-label for each speech frame in each first sample speech data.
[0171] Optionally, when generating frame-level wakeup pseudo labels, a smoothing averaging technique may be applied to the model parameters of the equalization model to generate frame-level wakeup pseudo labels. The specific implementation steps are as follows:
[0172] For the current batch training, the weight parameters of the balanced model after the current batch training can be weighted and added to the average model parameters of the balanced model after the previous batch training to obtain the average model parameters of the balanced model after the current batch training. The specific calculation can be obtained by the following formula:
[0173] ;
[0174] in, For the The average model parameters of the balanced model after batch training; For the The average model parameters of the balanced model after batch training; is the weighting coefficient; For the Weight parameters of the balanced model after batch training.
[0175] Here It can be pre-set according to the training objectives and actual application scenarios of the balancing model; or it can be dynamically determined according to the performance of the balancing model after the current batch training (such as the false awakening rate). If the performance of the balancing model after the current batch training is lower than that of the balancing model after the previous batch training, the value can be increased. To enhance the weight of the model parameters of the balanced model after historical batch training, if the performance of the balanced model after the current batch training is improved compared with the performance of the balanced model after the historical batch training, you can reduce The value of is used to enhance the weight of the model parameters of the equalization model after the current batch training, thereby improving the reliability of the word-level wake-up pseudo-label generation of the current batch training; or, the importance of the weight parameters of the equalization model after the current batch training and the average model parameters of the equalization model after the previous batch training are automatically learned and output through the attention mechanism, etc. This embodiment does not specifically limit this.
[0176] After obtaining the average model parameters of the equalization model after the current batch training, the labels of each speech frame in each first sample voice data can be predicted based on the average model parameters of the equalization model after the current batch training, so as to obtain the frame-level wake-up pseudo-label of each speech frame in each first sample voice data. For example, each first sample voice data can be input into a model constructed based on the average model parameters of the equalization model after the current batch training, and each speech frame in each first sample voice data can be feature-encoded based on the model and then feature-decoded to predict the frame-level wake-up pseudo-label of each speech frame in each first sample voice data, etc. This embodiment does not specifically limit this.
[0177] The method provided in this embodiment uses smoothing averaging technology to generate more reliable average model parameters by weighted fusion of the parameters of the equalization model of the current batch and the previous batch, thereby improving the accuracy and stability of frame-level wake-up pseudo-label generation, thereby improving the quality of pseudo-labels required for training the wake-up model and enhancing the performance and adaptability of the wake-up model.
[0178] In one possible implementation, the step of determining the weight parameters of the equalization model after the current batch training includes: optimizing the weight parameters of the equalization model after the previous batch training based on each second sample voice data and the frame-level wake-up true label of each voice frame in each second sample voice data to obtain the weight parameters of the equalization model after the current batch training.
[0179] Optionally, the weight parameters of the balanced model after the current batch training are specifically determined based on the following steps:
[0180] Inputting each second sample voice data into the equalization model trained in the previous batch, and using the equalization model trained in the previous batch and the weight parameters of the equalization model trained in the previous batch to perform label prediction on each voice frame in each second sample voice data, to obtain a frame-level wakeup prediction label for each voice frame in each second sample voice data;
[0181] The equalization frame-level loss is determined based on the deviation between the frame-level wake-up prediction label of each speech frame in each second sample speech data and the frame-level wake-up true label of each speech frame in each second sample speech data; based on the equalization frame-level loss, the weight parameters of the equalization model after training in the previous batch are optimized to obtain the weight parameters of the equalization model after training in the current batch.
[0182] The method provided in this embodiment inputs each second sample voice data into the equalization model trained in the previous batch, uses the weight parameters of the model to perform frame-level wake-up prediction, calculates the deviation between the predicted label and the true label to obtain the equalization frame-level loss, and optimizes the weight parameters of the equalization model accordingly, thereby improving the performance of the model after the current batch training, improving the accuracy of the frame-level wake-up prediction, and thus enhancing the reliability and stability of the wake-up model.
[0183] In some embodiments, the present application also provides a wake-up method. Figure 7 It is a flowchart of the awakening method provided by the present invention; Figure 7 As shown, the method includes: step 710, step 720 and step 730.
[0184] Step 710: Acquire target voice data received by the device to be awakened;
[0185] Step 720: Based on the optimized wake-up model, recognize each speech frame in the target speech data to obtain a frame-level wake-up recognition result for each speech frame in the target speech data, and recognize each speech word in the target speech data to obtain a word-level wake-up recognition result for each speech word in the target speech data.
[0186] Step 730 : Determine whether to wake up the device to be awakened based on the word-level wake-up recognition result and the frame-level wake-up recognition result; wherein the optimized wake-up model is obtained by optimizing the wake-up model based on a wake-up model optimization method.
[0187] The target voice data here is the voice data required for wake-up recognition received by the device to be awakened, which can be specifically collected by the device to be awakened through a microphone or sensor.
[0188] Optionally, after acquiring the target voice data received by the device to be awakened, an optimized awakening model may be loaded. The specific training process of the optimized awakening model can be found in Figure 1 The flow chart shown is not described here in detail.
[0189] After loading into the optimized wake-up model, the optimized wake-up model can be used to recognize each speech frame in the target speech data to obtain a frame-level wake-up recognition result for each speech frame in the target speech data, and to recognize each speech word in the target speech data to generate a word-level wake-up recognition result for each speech word in the target speech data.
[0190] Then, the word-level wake-up recognition result and the frame-level wake-up recognition result are combined to determine whether to wake up the device to be awakened, that is, whether to switch the device to be awakened from the sleep state to the working mode. In this process, the corresponding confidence thresholds can be set for the word-level wake-up recognition result and the frame-level wake-up recognition result respectively. For example, the word-level confidence threshold is set to Tw and the frame-level confidence threshold is set to Tf. When the confidence of the wake-up word in the word-level recognition result is higher than Tw, and at the same time, the confidence of the frame-level recognition result is higher than Tf in several consecutive frames, it is determined that the device is awakened to switch the device to be awakened from the sleep state to the working mode; or, the proportion of the number of frames judged to be awakened in the frame-level recognition result to the total number of frames is determined, and whether the wake-up word appears in the word-level recognition result. If the proportion of the number of awakening frames exceeds a certain threshold and the wake-up word appears in the word-level recognition result, it is determined that the device is awakened to switch the device to be awakened from the sleep state to the working mode, etc. This embodiment does not specifically limit this.
[0191] The method provided in this embodiment generates multi-scale, high-quality pseudo labels through parameter smoothing, a multi-level pseudo-label generation strategy, and a multi-scale label generation strategy. The generated multi-scale, high-quality pseudo labels, as well as the frame-level wake-up true labels and the word-level wake-up true labels are combined to perform semi-supervised training on the wake-up model. This can effectively obtain an optimized wake-up model with a low false wake-up rate in low-power scenarios and high robustness and high generalization ability in different wake-up scenarios. When the optimized wake-up model is used for device wake-up recognition, the false wake-up rate of the device can be effectively reduced.
[0192] The wake-up model optimization device provided by the present invention is described below. The wake-up model optimization device described below and the wake-up model optimization method described above can be referenced to each other.
[0193] Figure 8 Schematic diagram of the structure of the wake-up model optimization device provided by the present invention; Figure 8As shown, the device includes: a first acquisition unit 810 for acquiring a training subset corresponding to the current batch training in the sample data set for the current batch training; a processing unit 820 for generating a frame-level wake-up pseudo-label for each speech frame in each first sample speech data in the training subset according to the weight parameters of the equalization model after the current batch training and the average model parameters of the equalization model after the previous batch training, and generating a word-level wake-up pseudo-label for each speech word in each first sample speech data according to the frame-level wake-up pseudo-label; a training unit 830 for performing semi-supervised training on the wake-up model after the previous batch training according to each first sample speech data, the frame-level wake-up pseudo-label, the word-level wake-up pseudo-label, each second sample speech data in the training subset, and the frame-level wake-up true label of each speech frame and the word-level wake-up true label of each speech word in each second sample speech data to obtain the wake-up model after the current batch training; an optimization unit 840 for determining an optimized wake-up model based on the wake-up models after multiple batch training.
[0194] The device provided in this embodiment combines the weight parameters of the balanced model after training of the current batch with the average model parameters of the previous batch to accurately generate frame-level wake-up pseudo-labels for the first sample voice data in the training subset, and then generates word-level wake-up pseudo-labels based on the frame-level wake-up pseudo-labels. Thus, through parameter smoothing, multi-level pseudo-label generation strategy and multi-scale label generation strategy, the quality of pseudo-labels is improved and the deviation caused by limited computing power is reduced. The training data set of the wake-up model is constructed by combining the first sample voice data and their corresponding frame-level wake-up pseudo-labels and word-level wake-up pseudo-labels, as well as the second sample voice data and their corresponding frame-level wake-up true labels and word-level wake-up true labels, further balancing the distribution of training data for the wake-up model and improving the quality of training data for the wake-up model. Thus, the performance of the wake-up model is optimized by the training data set, which can effectively reduce its false wake-up rate in low-power scenarios and enhance its robustness and generalization ability in different wake-up scenarios.
[0195] Figure 9 Schematic diagram of the structure of the awakening device provided by the present invention; Figure 9 As shown, the device includes: a second acquisition unit 910 for acquiring target voice data received by the device to be awakened; a recognition unit 920 for recognizing each voice frame in the target voice data based on an optimized wake-up model, and obtaining a frame-level wake-up recognition result of each voice frame in the target voice data, and recognizing each voice word in the target voice data, and obtaining a word-level wake-up recognition result of each voice word in the target voice data; a wake-up unit 930 for determining whether to wake up the device to be awakened based on the word-level wake-up recognition result and the frame-level wake-up recognition result; wherein the optimized wake-up model is optimized based on the wake-up model optimization method as described in any of the above items.
[0196] The device provided in this embodiment generates multi-scale, high-quality pseudo labels through parameter smoothing, a multi-level pseudo-label generation strategy, and a multi-scale label generation strategy. The multi-scale, high-quality pseudo labels generated thereby are combined with the frame-level wake-up true labels and the word-level wake-up true labels to perform semi-supervised training on the wake-up model. This can effectively obtain an optimized wake-up model with a low false wake-up rate in low-power scenarios and high robustness and high generalization ability in different wake-up scenarios. When the optimized wake-up model is used for wake-up recognition of the device, the false wake-up rate of the device can be effectively reduced.
[0197] The device provided by the present invention is used to execute the above-mentioned method embodiments. Please refer to the above-mentioned embodiments for the specific processes and detailed contents, which will not be repeated here.
[0198] Figure 10 An example of a physical structure diagram of an electronic device is shown below. Figure 10As shown, the electronic device may include: a processor (processor) 1010, a communication interface (Communications Interface) 1020, a memory (memory) 1030 and a communication bus 1040, wherein the processor 1010, the communication interface 1020, and the memory 1030 communicate with each other through the communication bus 1040. The processor 1010 can call the logic instructions in the memory 1030 to execute the wake-up model optimization method, which includes: for the current batch training, obtaining the training subset corresponding to the current batch training in the sample data set; generating a frame-level wake-up pseudo-label for each speech frame in each first sample voice data in the training subset according to the weight parameters of the equalization model after the current batch training and the average model parameters of the equalization model after the previous batch training, and generating a word-level wake-up pseudo-label for each voice word in each first sample voice data according to the frame-level wake-up pseudo-label; generating a word-level wake-up pseudo-label for each voice word in each first sample voice data according to each first sample voice data, the frame-level wake-up pseudo-label, the word-level wake-up pseudo-label, each second sample voice data in the training subset, and each second sample voice data The frame-level wake-up true label of the speech frame and the word-level wake-up true label of each speech word are used to perform semi-supervised training on the wake-up model trained in the previous batch to obtain the wake-up model trained in the current batch; based on the wake-up models trained in multiple batches, an optimized wake-up model is determined; or, a wake-up method is executed, which includes: obtaining target speech data received by the device to be awakened; based on the optimized wake-up model, each speech frame in the target speech data is recognized to obtain a frame-level wake-up recognition result of each speech frame in the target speech data, and each speech word in the target speech data is recognized to obtain a word-level wake-up recognition result of each speech word in the target speech data; based on the word-level wake-up recognition result and the frame-level wake-up recognition result, it is determined whether to wake up the device to be awakened.
[0199] Furthermore, the logic instructions in the aforementioned memory 1030 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0200] On the other hand, the present invention also provides a computer program product, which includes a computer program, which can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the wake-up model optimization method provided by the above methods, which includes: for the current batch training, obtaining a training subset corresponding to the current batch training in the sample data set; generating a frame-level wake-up pseudo-label for each speech frame in each first sample speech data in the training subset based on the weight parameters of the equalization model after the current batch training and the average model parameters of the equalization model after the previous batch training, and generating a word-level wake-up pseudo-label for each speech word in each first sample speech data based on the frame-level wake-up pseudo-label; generating a word-level wake-up pseudo-label for each speech word in each first sample speech data based on the first sample speech data, the frame-level wake-up pseudo-label, the word-level wake-up pseudo-label, the The wake-up model trained in the previous batch is semi-supervisedly trained based on each second sample voice data in the training subset, as well as the frame-level wake-up true label of each voice frame and the word-level wake-up true label of each voice word in each second sample voice data, to obtain the wake-up model trained in the current batch; an optimized wake-up model is determined based on the wake-up models trained in multiple batches; or, a wake-up method is executed, which includes: obtaining target voice data received by the device to be awakened; based on the optimized wake-up model, each voice frame in the target voice data is recognized to obtain a frame-level wake-up recognition result of each voice frame in the target voice data, and each voice word in the target voice data is recognized to obtain a word-level wake-up recognition result of each voice word in the target voice data; and according to the word-level wake-up recognition result and the frame-level wake-up recognition result, it is determined whether to wake up the device to be awakened.
[0201] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which is implemented by a processor to execute the wake-up model optimization method provided by the above-mentioned methods, the method comprising: for the current batch training, obtaining a training subset corresponding to the current batch training in the sample data set; generating a frame-level wake-up pseudo-label for each speech frame in each first sample speech data in the training subset according to the weight parameters of the equalization model after the current batch training and the average model parameters of the equalization model after the previous batch training, and generating a word-level wake-up pseudo-label for each speech word in each first sample speech data according to the frame-level wake-up pseudo-label; generating a word-level wake-up pseudo-label for each speech word in each first sample speech data according to the first sample speech data, the frame-level wake-up pseudo-label, the word-level wake-up pseudo-label, and the second sample speech data in the training subset according to the weight parameters of the equalization model after the current batch training and the average model parameters of the equalization model after the previous batch training. According to the data, as well as the frame-level wake-up true label of each voice frame and the word-level wake-up true label of each voice word in each second sample voice data, semi-supervised training is performed on the wake-up model trained in the previous batch to obtain the wake-up model trained in the current batch; based on the wake-up models after multiple batches of training, an optimized wake-up model is determined; or, a wake-up method is executed, which includes: obtaining target voice data received by the device to be awakened; based on the optimized wake-up model, each voice frame in the target voice data is recognized to obtain a frame-level wake-up recognition result of each voice frame in the target voice data, and each voice word in the target voice data is recognized to obtain a word-level wake-up recognition result of each voice word in the target voice data; according to the word-level wake-up recognition result and the frame-level wake-up recognition result, it is determined whether to wake up the device to be awakened.
[0202] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.
[0203] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.
[0204] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A wake-up model optimization method, characterized in that: include: For the current batch training, obtaining a training subset corresponding to the current batch training in the sample data set; Generate a frame-level wakeup pseudo label for each speech frame in each first sample speech data in the training subset based on the weight parameters of the equalization model after training in the current batch and the average model parameters of the equalization model after training in the previous batch, and generate a word-level wakeup pseudo label for each speech word in each first sample speech data based on the frame-level wakeup pseudo label; Performing semi-supervised training on the wakeup model trained in the previous batch according to each of the first sample speech data, the frame-level wakeup pseudo-label, the word-level wakeup pseudo-label, each of the second sample speech data in the training subset, and the frame-level wakeup true label of each speech frame and the word-level wakeup true label of each speech word in each of the second sample speech data, to obtain a wakeup model trained in the current batch; An optimized wake-up model is determined based on the wake-up models trained after multiple batches.
2. The wake-up model optimization method according to claim 1, characterized in that: The second sample speech data includes first enhanced speech sub-data corresponding to the marked sample speech data; The first sample speech data includes second enhanced speech sub-data corresponding to the labeled sample speech data, and third enhanced speech sub-data and fourth enhanced speech sub-data corresponding to the unlabeled sample speech data; Among them, the content invariance rate between the first enhanced speech sub-data and the labeled sample speech data is greater than a preset value, the content invariance rate between the second enhanced speech sub-data and the labeled sample speech data is less than or equal to the preset value, the content invariance rate between the third enhanced speech sub-data and the unlabeled sample speech data is greater than the preset value, and the content invariance rate between the fourth enhanced speech sub-data and the unlabeled sample speech data is less than or equal to the preset value.
3. The wake-up model optimization method according to claim 2, characterized in that: The training steps of the wake-up model after the current batch training include: Determining a first prediction loss based on each of the second sample speech data, and a frame-level true wakeup label of each speech frame and a word-level true wakeup label of each speech word in each of the second sample speech data; Determining a second prediction loss based on each of the first sample speech data, and a frame-level wakeup pseudo label of each speech frame and a word-level wakeup pseudo label of each speech word in each of the first sample speech data; Determining a consistency regularization loss based on each of the first enhanced speech sub-data and each of the third enhanced speech sub-data; The wake-up model trained in the previous batch is trained according to the first prediction loss, the second prediction loss, and the consistency regularization loss to obtain the wake-up model trained in the current batch.
4. The wake-up model optimization method according to claim 3, characterized in that: The step of training the wake-up model trained in the previous batch according to the first prediction loss, the second prediction loss, and the consistency regularization loss to obtain the wake-up model trained in the current batch includes: Determining a weighting coefficient of the consistency regularization loss according to the frame-level wake-up pseudo-label of each speech frame in each of the third enhanced speech sub-data; performing weighted addition of the second prediction loss and the consistency regularization loss according to the weight coefficient of the consistency regularization loss to obtain an unsupervised loss; Determining a target loss based on the unsupervised loss and the first prediction loss; The wake-up model trained in the previous batch is trained according to the target loss to obtain the wake-up model trained in the current batch.
5. The wake-up model optimization method according to claim 4, characterized in that: The determining, according to the frame-level wake-up pseudo-label of each speech frame in each of the third enhanced speech sub-data, a weighted coefficient of the consistency regularization loss includes: Determining a weighting coefficient of the consistency regularization loss according to the total number of target speech frames in the non-silent segments of all the third enhanced speech sub-data in the training subset and the total number of speech frames in all the third enhanced speech sub-data in the training subset; The target speech frame is a speech frame whose probability value corresponding to the frame-level wake-up pseudo-label is greater than a probability threshold.
6. The wake-up model optimization method according to claim 4, characterized in that: The determining of a target loss according to the unsupervised loss and the first prediction loss includes: Determining a weighting coefficient for the unsupervised loss based on the number of first sample speech data in the training subset corresponding to the current batch training, the number of sample speech data to be labeled in the training subsets corresponding to all historical batch trainings before the current batch training, and the number of sample speech data to be labeled in the sample data set; According to the weighting coefficient of the unsupervised loss, the unsupervised loss and the first prediction loss are weightedly added to obtain the target loss.
7. The wake-up model optimization method according to any one of claims 1 to 6, characterized in that: Generating a word-level wake-up pseudo-label for each speech word in each of the first sample speech data according to the frame-level wake-up pseudo-label includes: For each spoken word in each of the first sample speech data, when determining that the spoken word matches the target wake-up word based on the frame-level wake-up pseudo-labels of each speech frame in the spoken word, average the probability values corresponding to the frame-level wake-up pseudo-labels of all speech frames in each character in the spoken word to obtain a score for each character in the spoken word; Calculating an average of the scores of all characters in the phonetic word to obtain the score of the phonetic word; When the score of the speech word is greater than or equal to a score threshold, marking the word-level wake-up pseudo-label of the speech word as a wake-up word label; Each voice word in each of the first sample voice data is traversed to obtain a word-level wake-up pseudo label for each voice word in each of the first sample voice data.
8. The wake-up model optimization method according to any one of claims 1 to 6, characterized in that: The generating, based on the weight parameters of the equalization model after training the current batch and the average model parameters of the equalization model after training the previous batch, a frame-level wake-up pseudo label for each speech frame in each first sample speech data in the training subset includes: Performing weighted addition on the weight parameters of the balanced model after the current batch training and the average model parameters of the balanced model after the previous batch training to obtain the average model parameters of the balanced model after the current batch training; Based on the average model parameters of the equalization model after training of the current batch, a label prediction is performed on each speech frame in each of the first sample speech data to obtain a frame-level wake-up pseudo label for each speech frame in each of the first sample speech data.
9. The wake-up model optimization method according to any one of claims 1 to 6, characterized in that: The step of determining the weight parameters of the balanced model after the current batch training includes: According to each of the second sample voice data and the frame-level wake-up true label of each voice frame in each of the second sample voice data, the weight parameters of the equalization model after the previous batch training are optimized to obtain the weight parameters of the equalization model after the current batch training.
10. A wake-up method, characterized in that: include: Obtain the target voice data received by the device to be awakened; Based on the optimized wake-up model, each speech frame in the target speech data is recognized to obtain a frame-level wake-up recognition result for each speech frame in the target speech data, and each speech word in the target speech data is recognized to obtain a word-level wake-up recognition result for each speech word in the target speech data; Determining whether to wake up the device to be awakened according to the word-level wake-up recognition result and the frame-level wake-up recognition result; The optimized wake-up model is obtained by optimizing the wake-up model according to any one of claims 1 to 9.
11. A wake-up model optimization device, characterized in that: include: A first acquisition unit is configured to acquire, for a current batch of training, a training subset corresponding to the current batch of training from a sample data set; a processing unit, configured to generate a frame-level wakeup pseudo label for each speech frame in each first sample speech data in the training subset based on the weight parameters of the equalization model after training in the current batch and the average model parameters of the equalization model after training in the previous batch, and generate a word-level wakeup pseudo label for each speech word in each first sample speech data based on the frame-level wakeup pseudo label; a training unit, configured to perform semi-supervised training on the wakeup model trained in the previous batch based on each of the first sample speech data, the frame-level wakeup pseudo-label, the word-level wakeup pseudo-label, each of the second sample speech data in the training subset, and the frame-level wakeup true label of each speech frame and the word-level wakeup true label of each speech word in each of the second sample speech data, to obtain a wakeup model trained in the current batch; The optimization unit is used to determine an optimized wake-up model based on the wake-up models trained after multiple batches.
12. A wake-up device, characterized in that: include: A second acquiring unit, configured to acquire target voice data received by the device to be awakened; a recognition unit, configured to recognize each speech frame in the target speech data based on the optimized wake-up model, obtain a frame-level wake-up recognition result for each speech frame in the target speech data, and recognize each speech word in the target speech data, obtain a word-level wake-up recognition result for each speech word in the target speech data; an awakening unit, configured to determine whether to wake up the device to be awakened according to the word-level awakening recognition result and the frame-level awakening recognition result; The optimized wake-up model is obtained by optimizing the wake-up model according to any one of claims 1 to 9.
13. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the wake-up model optimization method according to any one of claims 1 to 9 or the wake-up method according to claim 10 is implemented.
14. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the wake-up model optimization method according to any one of claims 1 to 9 or the wake-up method according to claim 10 is implemented.
15. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the wake-up model optimization method according to any one of claims 1 to 9 or the wake-up method according to claim 10 is implemented.
Citation Information
Patent Citations
Training method of voice wake-up model, wake-up word detection method and related equipment
CN113963688A
Pre-training optimization method and device for emotion prediction model, equipment and medium
CN115457982A