Keyword detection model training and reasoning method, electronic equipment and storage medium

Through mask self-distillation training and semi-autoregressive decoding methods, the overfitting problem of keyword detection models is solved, the performance and robustness of the model in complex acoustic scenarios is improved, and the dependence on large-scale data is reduced.

CN120338042APending Publication Date: 2025-07-18AISPEECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510424690.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing keyword detection models are prone to overfitting, resulting in low wake-up rate and high false alarm rate in complex acoustic scenarios, especially in high noise and low signal-to-noise ratio environments.

Method used

Mask self-distillation (MSD) training strategy and semi-autoregression (SAR) decoding method are used to randomly mask the output of the prediction network during the training process, reducing the model's dependence on the predictor, and combining the advantages of autoregression and non-autoregression decoding strategies to generate the final decoding results.

Benefits of technology

It improves the accuracy and robustness of the model in different scenarios, reduces the need for large-scale annotation data, and enhances the adaptability and reliability of the model in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120338042A_ABST
    Figure CN120338042A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a keyword detection model training and reasoning method, electronic equipment and a storage medium, and the method comprises the steps: carrying out the random shielding of a potential representation of a text sequence in a training stage, and generating a shielded version; using the keyword detection model to use the shielded version and the unshielded version together, calculating predicted distribution through the joint network, and calculating distribution of the unshielded version as a teacher signal of self-distillation learning at the same time; in the inference stage, the keyword detection model supports two decoding modes: autoregression decoding and non-autoregression decoding, in the autoregression decoding mode, the output of the prediction network is used to generate activation scores of keywords, in the non-autoregression decoding mode, the output of the prediction network is shielded, and in the non-autoregression decoding mode, the activation scores of the keywords are generated. Generating an activation score using only the encoder and the joint network; and fusing the activation fractions of the two decoding modes to obtain a final decoding result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of task-based dialogue, and particularly relates to a method for training and inferring a keyword detection model, an electronic device, and a storage medium. Background Art

[0002] In related technologies, keyword detection technology is to detect whether a preset keyword appears in a continuous speech stream, and wake up the device when the wake-up word appears. A keyword detection system (KWS, Keyword Spotting) aims to detect predefined keywords in continuous speech, and the keyword-based interaction front-end is crucial for intelligent devices. Due to computational and memory limitations, designing a compact and powerful KWS model has become a challenging research topic. With natural flow capabilities and excellent performance, RNN-T has achieved great success in multiple research fields, such as automatic speech recognition (ASR, Automatic Speech Recognition), speech translation (ST, Speech Translation), and text-to-speech (TTS, Text-to-Speech). Sensor-based KWS systems have also received much attention. Some technologies introduce a bias module into RNN-T to enhance keyword detection. Some technologies achieve high performance of RNN-T KWS by using a large amount of synthetic and real speech data. Some technologies adopt a multi-stage strategy to reduce false alarms. Some technologies (such as TDT-KWS) propose an efficient decoding algorithm for RNN-T KWS, using token and duration transducer variants to speed up decoding.

[0003] The inventors found that since the positive training data of KWS usually only includes keyword sequences, the KWS model tends to overfit the keywords, thereby reducing the distinguishability between positive and negative samples and resulting in a high false alarm rate. Due to the simplicity of the prediction network architecture, this problem is particularly prominent in the RNN-T KWS model. Previous studies have shown that using a hybrid encoder or a multi-stage strategy can effectively alleviate overfitting and reduce false alarms. For example, some technologies propose a multi-stage detection paradigm based on the posterior probability output by the sensor acoustic model, which can gradually reduce false alarms. Some technologies (such as CaTT-KWS) integrate a transducer-based keyword detector, a frame-level force alignment module, and a Transformer-based decoder to create a multi-stage decoding process. Some technologies (such as U2-KWS) use the CTC branch and the decoder branch as the first and second stage models respectively to improve detection reliability. Although hybrid systems and multi-stage strategies can alleviate overfitting, they also bring complexity to the training process and the decoding pipeline. Excessive manual operations and prior knowledge in multiple stages may lead to unsatisfactory performance of keyword detection. Summary of the Invention

[0004] An embodiment of the present invention provides a method for training and inferring a keyword detection model, an electronic device, and a storage medium, which are used to solve at least one of the above technical problems.

[0005] In a first aspect, an embodiment of the present invention provides a method for training and inferring a keyword detection model. The keyword detection model includes an encoder, a prediction network, and a joint network. The method includes: in the training stage, randomly masking the latent representation of a text sequence to generate a version with some information hidden; using the keyword detection model to use this masked version and the unmasked version together to calculate a predicted distribution through the joint network, and the keyword detection model simultaneously calculates the distribution of the unmasked version as a teacher signal for self-distillation learning; in the inference stage, the keyword detection model supports two decoding modes: autoregressive decoding and non-autoregressive decoding. Among them, in the autoregressive decoding mode, the keyword detection model uses the output of the prediction network to generate activation scores for keywords, and in the non-autoregressive decoding mode, the keyword detection model masks the output of the prediction network and only uses the encoder and the joint network to generate activation scores; fusing the activation scores of these two decoding modes to obtain a final decoding result.

[0006] In a second aspect, an embodiment of the present invention further provides a computer program product. The computer program product includes a computer program stored on a non-volatile computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer is enabled to execute the steps of the method for training and inferring a keyword detection model according to any embodiment of the present invention.

[0007] In a third aspect, an embodiment of the present invention further provides an electronic device, which includes: at least one processor, and a memory communicatively connected to the at least one processor. Among them, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the steps of the method described in the first aspect.

[0008] In a fourth aspect, an embodiment of the present invention further provides a storage medium, on which a computer program is stored. The computer program is characterized in that when the computer program is executed by a processor, the steps of the method described in the first aspect are implemented.

[0009] In the method of the embodiment of the present application, by randomly masking the output of the prediction network during the training process, the model is made independent of the predictor, thereby reducing overfitting and improving the generalization ability of the model. In the inference stage, by combining the advantages of autoregressive decoding and non-autoregressive decoding, and by weighted fusion of the activation scores of the two decoding strategies, the performance of the model in different scenarios is improved. This method not only improves the accuracy and robustness of the model, but also reduces the need for large-scale labeled data and improves the data utilization efficiency. In addition, the adaptability and reliability of the model in complex scenarios have also been significantly improved, providing a basis for further optimization and improvement of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0011] Figure 1 Processing flow provided by an embodiment of the present invention; Figure 2 Flowchart of an embodiment of the keyword detection model training and inference method provided by an embodiment of the present invention; Figure 3 Overview of the training and inference of the proposed system provided by an embodiment of the present invention; Figure 4 Selected keywords in LibriKWS-20 provided by an embodiment of the present invention; Figure 5 When applying different mask ratios γmask to the prediction output on LibriKWS20 in an embodiment of the present invention, the macro call results of SAR are different; Figure 6 Recall rates of different decoding strategies when the FAR value is 4 in an embodiment of the present invention; Figure 7 Proposed model results with and without masked self-distillation loss LMSD using AR, NAR, and SAR decoding strategies in all 5 test sets in an embodiment of the present invention; Figure 8 Recall values of the Xiaowen dataset in various low FAs in an embodiment of the present invention. This dataset shows the most serious overfitting problem under strict test conditions; Figure 9 Schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0012] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without making creative efforts fall within the scope of protection of the present invention.

[0013] The inventors found that, on the one hand, the prediction network structure of the model in the related art is relatively simple and easily falls into an overfitting state, resulting in a low wake-up rate at a low number of false wake-ups, which is particularly obvious in complex acoustic scenarios with high noise and low signal-to-noise ratio. On the other hand, the original autoregressive KWS decoding algorithm overly focuses on the keywords used in the training process.

[0014] The above defects are mainly caused by the following reasons: First, due to the simple prediction network architecture, the model is prone to overfitting to the keywords, which is particularly obvious in the keyword detection model based on the transducer structure. Second, since the positive training data of KWS usually only contains the keyword sequence, the KWS model often overfits the keywords, thereby reducing the distinguishability between positive and negative samples, resulting in a high false wake-up rate. Third, the decoding algorithm is not advanced enough and lacks pertinence, and is not specifically optimized for the keyword detection system.

[0015] When practitioners in this industry solve these defects, they may increase the training data and use more data to train a model with better performance and greater robustness; they may also design a model with a larger number of parameters and a larger computational amount, and train a larger model to alleviate the overfitting phenomenon, thereby enhancing the performance in low signal-to-noise ratio scenarios; they may also design a multi-task and multi-level system to enhance the performance through the multi-level system.

[0016] The reason why the practitioners in this industry are not likely to come up with the solution of this application is as follows: First, the masked self-distillation (MSD) training strategy and semi-autoregressive (SAR) decoding method we proposed are relatively novel in the field of keyword detection. These methods randomly mask the output of the predictor, enabling the model to learn not to rely on the predictor during training, thereby reducing overfitting. This training strategy and decoding method are not common in traditional models. In the field of keyword detection, traditional solutions usually focus on increasing model complexity, data augmentation, multi-task learning, etc. These methods can alleviate the overfitting problem to a certain extent, but may not consider the innovative idea of reducing overfitting by masking the predictor output. Second, no previous research team has studied this problem along this line of thought in this complex scenario, and there is very little previous work that can be referred to.

[0017] In the embodiments of this application, we aim to find a simpler and more effective method for Transducer-based models. Recently, a hybrid-autoregressive Transducer has emerged, which converts the Transducer into a non-autoregressive mode. This model learns to process speech without relying on the prediction network by randomly setting its output to zero during training. The embodiments of this application propose a sensor-based KWS self-correction model that adopts a semi-autoregressive decoding strategy. Our core contributions can be summarized as follows: - We propose a masked self-distillation (MSD) training framework for on-device Transducer-based KWS systems. By adopting the MSD strategy, mask learning is significantly enhanced, enabling the model to effectively perform masked NAR decoding.

[0018] - We introduce a semi-autoregressive (SAR) decoding strategy for KWS, which combines the excellent performance of standard keyword-based Transducers with the robust and non-overfitting ability of zero-based Transducer decoding.

[0019] - Experimental results show that our framework can not only maintain strong performance on simpler datasets, but also mitigate overfitting and achieve robust performance in challenging test cases.

[0020] Specifically, in this embodiment, MSD training forces the model to learn without relying on the predictor by randomly masking the output of the prediction network during training. The specific operation is to apply a masking probability to the latent text representation derived from the text sequence, and then use these masked representations to calculate the loss. In addition, a self-distillation loss is introduced to use the output of the unmasked predictor to guide the masked predictor, ensuring that the model can produce accurate results without the assistance of the predictor. This method helps reduce the risk of overfitting by reducing the model's dependence on the prediction network. Further, SAR decoding combines the advantages of autoregressive and non-autoregressive decoding strategies. During the inference process, the model can adopt two modes: autoregressive decoding, which uses the prediction network to generate keyword activation scores; non-autoregressive decoding, which masks the output of the prediction network and only uses the encoder and the joint network for decoding. The final decoding score is a weighted combination of the autoregressive and non-autoregressive decoding scores, enabling the model to utilize the high performance of autoregressive decoding on normal datasets while benefiting from the robustness of non-autoregressive decoding in complex scenarios prone to overfitting.

[0021] Please refer to Figure 1 , which shows the processing flow of the embodiment of the present application.

[0022] During the training phase, the model processes two types of speech samples simultaneously: one is normal unmasked samples, and the other is samples with randomly masked parts of the prediction network output. For each mini-batch of data, we randomly select some samples for masking according to a preset masking probability. Specifically, we randomly mask the latent representation of the text sequence to generate a version with some information hidden. Then, the model uses this masked version together with the normal audio features to calculate the predicted distribution through the joint network. At the same time, the model also calculates the distribution of the unmasked version as the teacher signal for self-distillation learning. In this way, the model not only learns the normal output of the prediction network during training but also learns how to make predictions when part of the information of the prediction network is missing, thus improving the robustness of the model.

[0023] During the inference phase, the model supports two decoding modes: autoregressive decoding and non-autoregressive decoding. In the autoregressive decoding mode, the model uses the output of the prediction network to generate keyword activation scores like a traditional RNN-T model. In the non-autoregressive decoding mode, the model masks the output of the prediction network and only uses the encoder and the joint network to generate activation scores. Finally, we fuse the activation scores of these two decoding modes to obtain the final decoding result. This fusion strategy enables the model to maintain the high performance of autoregressive decoding while also being able to utilize the robustness of non-autoregressive decoding, thus achieving excellent performance in different test scenarios.

[0024] In the embodiments of the present application, by randomly masking the output of the prediction network during the training process, the model is made independent of the predictor, thereby reducing overfitting and improving the generalization ability of the model. In the inference stage, by combining the advantages of autoregressive decoding and non-autoregressive decoding, and by weighted fusion of the activation scores of the two decoding strategies, the performance of the model in different scenarios is improved. This method not only improves the accuracy and robustness of the model, but also reduces the need for large-scale labeled data and improves the data utilization efficiency. In addition, the adaptability and reliability of the model in complex scenarios have also been significantly improved, providing a basis for further optimization and improvement of the model.

[0025] Please refer to Figure 2 , which shows a flowchart of an embodiment of the method for training and inferring a keyword detection model of the present application. Among them, the keyword detection model includes an encoder, a prediction network, and a joint network.

[0026] As Figure 2 shown, in step 201, in the training stage, the latent representation of the text sequence is randomly masked to generate a version with some information hidden; In step 202, the keyword detection model uses this masked version and the unmasked version together to calculate the predicted distribution through the joint network, and the keyword detection model simultaneously calculates the distribution of the unmasked version as the teacher signal for self-distillation learning; In step 203, in the inference stage, the keyword detection model supports two decoding modes: autoregressive decoding and non-autoregressive decoding. Among them, in the autoregressive decoding mode, the keyword detection model uses the output of the prediction network to generate the activation score of the keyword, and in the non-autoregressive decoding mode, the keyword detection model masks the output of the prediction network and only uses the encoder and the joint network to generate the activation score; the activation scores of these two decoding modes are fused to obtain the final decoding result.

[0027] In the method of the embodiments of the present application, by randomly masking the output of the prediction network during the training process, the model is made independent of the predictor, thereby reducing overfitting and improving the generalization ability of the model. In the inference stage, by combining the advantages of autoregressive decoding and non-autoregressive decoding, and by weighted fusion of the activation scores of the two decoding strategies, the performance of the model in different scenarios is improved. This method not only improves the accuracy and robustness of the model, but also reduces the need for large-scale labeled data and improves the data utilization efficiency. In addition, the adaptability and reliability of the model in complex scenarios have also been significantly improved, providing a basis for further optimization and improvement of the model.

[0028] In some optional embodiments, during the training process, the prediction network receives transcription information, while during the inference process, the prediction network receives keyword information.

[0029] In some alternative embodiments, randomly masking the latent representation of the text sequence to generate a version with some information hidden includes: masking the latent representation of the text sequence at the token level according to a preset masking probability to obtain a masked version.

[0030] In some alternative embodiments, the masked version is taught using the logarithm corresponding to the unmasked version.

[0031] In some alternative embodiments, using a self-teacher-student learning paradigm helps ensure that the masked output is substantially the same as the output of the keyword detection model, enabling the keyword detection model to learn how to produce accurate results without relying on the prediction network.

[0032] In some alternative embodiments, the keyword detection model is an RNN-T model.

[0033] It should be noted that the above method steps do not limit the execution order of each step. In fact, some steps may be executed simultaneously or in the reverse order of the step definition, and there is no limitation in this application.

[0034] The solution of the present application will be described below through a specific embodiment to enable those skilled in the art to better understand the solution of the present application.

[0035] RNN-T based keyword spotting (KWS) using autoregressive (AR) decoding has received much attention due to its streaming architecture and excellent performance. However, the simplicity of the RNN-T prediction network may lead to overfitting in complex scenarios, resulting in performance degradation. To this end, the present application proposes a masked self-distillation (MSD) training strategy to avoid the RNN-T's over-reliance on the prediction network, thereby alleviating the overfitting problem. This training can achieve masked non-autoregressive (NAR) decoding, completely masking the output of the RNN-T predictor during the KWS decoding process. In addition, we also propose a semi-autoregressive (SAR) decoding method to combine the advantages of AR and NAR decoding. Experimental results show that the masked self-distillation MSD training can effectively alleviate overfitting, and the SAR decoding method not only retains the excellent performance of AR decoding but also benefits from the overfitting suppression of NAR decoding, ultimately achieving excellent KWS recognition results.

[0036] Working method Transducer-based keyword spotting The transducer consists of an encoder, a prediction network (predictor), and a joint network (jointer). The encoder converts the acoustic features x = [x1,x2,...,x T ∈ R T×F (where each xt is an acoustic feature vector of dimension F and T represents the total number of frames) into a latent representation h audio . The predictor converts the transcriptions y = [y1,y2,...,y U ∈R U×1 (where y u is a label token and U is the total number of tokens) into a high-level text representation h text . The jointer conditions on the high-level acoustic representation h audio and the high-level text representation h text sequences to predict the next token logits (raw unnormalized scores) p token ∈ R T×U×V , where V is the vocabulary size and the size of the special blank. The logits for one step can be expressed as where , and the subscripts t and u represent the current acoustic frame index as t and the text index as u respectively.

[0037] For the decoding strategy of sensor-based KWS, most methods adopt ASR decoding techniques to generate hypotheses and match keywords. These algorithms consider the entire search space rather than focusing on the presence of specific keywords, so they are not ideal for discovering keywords. Related technologies have proposed a streaming decoding algorithm designed specifically for sensor-based KWS. This algorithm fixes the predictor input as the keyword during the inference process and searches within a specific keyword grid. The search can be dynamically initiated at any time step, enabling streaming inference. In the embodiments of this part, we will extend our framework by combining the streaming sensor decoding algorithm.

[0038] Figure 3Shows an overview of the training and inference of the proposed system. Traditional RNN-T training and masked RNN-T training share the same structure and parameters. During training, the prediction network receives transcription information, while during inference, it receives keyword information. Chinese-English comparison is as follows: Prediction Network: Prediction neural network module, Transcription / Keyword: Transcription text / Keyword, Waveform: Audio, Random Maskγ mask : Random mask γ 掩码 ,FullMask: Full mask, Joint Network: Joint network, h text : h 文本 ,h audio : h 音频 ,h mask : h 掩码 ,Sharedweights: Shared parameters, Streaming Decoding: Streaming decoding, Score AR : Score AR ,Score NAR : Score NAR ,Score SAR : Score SAR ,Fusion: Fusion, training: Training, inference: Inference.

[0039] Masked self-distillation training Figure 3 Shows the overall structure of our framework. This section will introduce the training process in detail. In each mini-batch, there are two types of training samples: non-masked and masked utterances. According to the preset masking probability , we perform masking at the token level (where token is the smallest semantic / syntactic unit after splitting text or speech data). Specifically, for the latent textual representation h text derived from the text sequence y, each latent vector The probability of being masked .

[0040] This operation is represented as: Therefore, the operation of the joint network can be expressed as: In addition, the RNN-T loss for the training pair (x, y) can be expressed as The RNN-T loss that masks the predictor output is shown as: where , where U and D represent the sequence length and the latent dimension respectively.

[0041] Our preliminary experiments show that due to the limited scale of the KWS model, the transformer cannot converge well, and it is difficult for the model to learn the masked prediction output well. Therefore, for an utterance (a complete unit of language expression), we propose a Masked Self-Distillation (MSD) strategy to assist NAR training, that is, using the logarithm corresponding to the unmasked input to teach the masked input. We apply the Kullback-Leibler (KL) divergence to two different output distributions. For the entire utterance, the MSD loss can be expressed as: The self-teacher-student (STS) learning paradigm helps to ensure that the masked output is very similar to the standard sensor output, enabling the model to learn how to produce accurate results without relying on the prediction network. This reduces the overfitting of the predictor to the keyword data. In addition, the transformer loss is also applied to the logarithms corresponding to the unmasked and masked inputs. The formula for the total training loss is: where λ mask and λ SD are loss coefficients used to balance the loss contributions.

[0042] Semi-autoregressive decoding inference Previous RNN-T required autoregressive (AR) style inference, while masked training allows the predictor to be omitted at each inference step, making non-autoregressive (NAR) decoding possible. As Figure 3As shown in the inference part, our self-distillation transducer supports two inference modes: autoregressive decoding (AR), where the model operates as a standard transducer and uses a streaming algorithm to infer the keyword activation scores at each time step; and non-autoregressive decoding (NAR), where the predicted network output is pre-masked, leaving only the encoder and joint network for decoding, effectively making the model a stateless acoustic model. In both modes, we employ a keyword lattice and a state-of-the-art (SOTA) search algorithm to compute the keyword scores at each time step. To fully utilize the superior performance of AR decoding and the non-overfitting characteristics of NAR decoding, we further fuse the two activation scores and propose semi-autoregressive decoding (SAR). At time step t, the decoding score can be expressed as: where , , represent the three activation scores of different decoding strategies in frame t, and α is a coefficient that balances the importance of the two decoding strategies.

[0043] Experimental setup Dataset To comprehensively evaluate the performance and robustness of the model, we evaluated the proposed model under different scenarios: 1) Fixed English keywords (EN Fixed); 2) Arbitrary English keywords (EN Arbitrary); 3) Noisy fixed Mandarin keywords (ZH Noisy). The datasets are as follows: - Hey Snips (Snips): Snips is a widely used KWS benchmark that contains positive samples with the keyword "Hey Snips". However, the official dataset lacks the transcripts of negative samples, which are crucial for training the RNN-T model. To fully utilize this data, we compiled all the negative samples in the training set, development set, and test set into a larger 97-hour negative test dataset.

[0044] - LibriSpeech and LibriKWS-20: LibriSpeech is a publicly available high-quality English speech dataset with transcripts. We trained English RNN-T models on the 960-hour training set of LibriSpeech for two purposes. First, these models can serve as seed models for training the Snips KWS model. Second, their arbitrary KWS performance was evaluated on LibriKWS-20. This dataset selected 20 arbitrary keywords from the LibriSpeech test-clean and test-other subsets, Figure 4These keywords are listed.

[0045] Figure 4 Shows the selected keywords in LibriKWS-20. The keywords for the test clean and test-other datasets are the same.

[0046] - AISHELL-2 and MobvoiHotwords: AISHELL-2 is an open-source Mandarin ASR dataset containing approximately 1000 hours of training data. We used AISHELL2 to train the seed RNN-T model for Mandarin KWS experiments. MobvoiHotwords is a Mandarin KWS dataset containing two keywords: "nihao wenwen" (classical Chinese) and "hi xiaowen" (Xiaowen). (Different from Snips and LibriKWS-20, the non-keyword data was collected in a less controlled environment at different distances from the smart speaker. These recordings include background noise with different signal-to-noise ratios (SNRs), such as typical household sounds (e.g., music and TV). Therefore, this type of dataset is more challenging and prone to causing the model on the device to overfit to the noise, resulting in a higher false alarm rate.

[0047] Configuration Feature extraction. 40-dimensional FBank features are extracted using a 25-millisecond window and a 10-millisecond shift. The first five frames and the last five frames are concatenated to create a 440-dimensional encoder input. During training, an online speed perturbation is applied with a ratio randomly selected from {0.9, 1.0, 1.1}. Additionally, to enhance robustness, SpecAugment is also used during training, applying two random masks in the frequency domain and the time domain, f max = 10 and t max = 50.

[0048] Model configuration. The encoder consists of 8 layers of deep feed-forward sequential memory network (DFSMN) blocks, with input, hidden, and output sizes of 440, 768, and 320 respectively. The left and right context frames for each DFSMN layer are set to 20 and 8 respectively. The context size used by the stateless predictor implemented in NeMo is 2, and the embedding dimension is 320. The joint network merges the 320-dimensional hidden vectors of the encoder and decoder into a 256-dimensional hidden representation. The final output of the transducer includes 70 monophones and 1 additional blank token. These monophone symbols are converted using a grapheme-to-phoneme (G2P) tool.

[0049] Training details: We used the AdamW optimizer with an initial learning rate of 1e-3 and betas set to (0.9, 0.999). The learning rate was adjusted using the ReduceLROnPlateau scheduler. The model was trained on 8 NVIDIA V100 GPUs with a mini-batch size of 12,288 frames. We set λ mask = 1 and λ MSD = 0.003 to balance different loss terms.

[0050] Evaluation: We reported the recall rate for all datasets at a specific false alarm rate (FA). For the LibriKWS-20 test-clean set and test-other set containing multiple keywords, we gave the average macro recall rate for all 20 keywords. For the three English test sets (EN fixed set and EN arbitrary set), we set the SAR decoding coefficient to 0.5 as described in Equation (9). For the challenging Mandarin "Wenwen" and "Xiaowen" datasets (ZH Noisy), we set α = 0.3 to make the NAR score dominant in the SAR decoding process.

[0051] Results and Analysis Optimal Masking Rate Figure 5 Shows the different macro call results of SAR when applying different masking ratios γ mask to the predicted output on LibriKWS20. The FA value for all keywords was 4.

[0052] Figure 5 Lists the KWS performance at different masking probabilities γ mask . From the table, we can see that the performance of KWS generally exceeds that of the standard RNN-T, and the performance of KWS is the best when the masking ratio is around 0.35. These results demonstrate the advantage of mask training for the RNN-T-based KWS system and highlight the important role of the predictor. After the optimal value, the performance gradually decreases as the masking ratio increases. In all subsequent experiments, unless otherwise stated, we use the optimal masking ratio γmask = 0.35.

[0053] Comparison of Various Decoding Strategies Figure 6Shows the recall rate of different decoding strategies at a FAR value of 4. For test-clean (test set - simple) and test-other (test set - complex) containing 20 keywords, we report the macro recall rate. The underlined results indicate that the model overfits the noise on the Mandarin KWS dataset. Among them, the Chinese-English comparison is as follows: Decoding: decoding algorithm, Recall: recall rate, #FA = 4: number of false awakenings = 4, EN Fixed: English fixed wake word, EN Arbitrary: English arbitrary wake word, ZH Noisy: Chinese fixed wake word.

[0054] Figure 6 Lists the KWS performance of various decoding strategies in different test scenarios. The decoding strategies for KWS (AR, NAR, SAR) are superior to those for ASR (greedy search and beam search). Although AR decoding performs well on datasets with fixed English keywords and arbitrary English keywords, it performs poorly on the Mandarin noise dataset (underlined results). There may be two reasons for this: 1) The Xiaowen and classical Chinese datasets are too noisy for the model to effectively learn keyword patterns in this environment (insufficient fitting of keywords); and 2) Since the forward data only contains keywords, the prediction network of RNN-T overfits the keywords and misinterprets the noise as keywords. Surprisingly, NAR decoding (without a predictor) performs well on the Xiaowen dataset and the classical Chinese text dataset, verifying the second hypothesis. As Figure 6 shown, the proposed NAR decoding effectively alleviates the overfitting problem in challenging datasets. For standard datasets (Snips, test-clean, test-other), the performance of AR decoding is comparable to that of NAR decoding and even exceeds it, indicating that the information provided by the prediction network is still crucial for achieving good performance. To combine the advantages of AR (high performance) and NAR (suppressing overfitting), SAR decoding achieves the best performance on all datasets, demonstrating its efficiency.

[0055] Importance of masked self-distillation Figure 7 Shows the results of the proposed model with and without the masked self-distillation loss LMSD using AR, NAR, and SAR decoding strategies in all 5 test sets. The results for all keywords are reported with an FA value of 4. "Avg.Imp." represents the average absolute improvement after applying LMSD. Among them, Avg.Imp.: average improvement.

[0056] Figure 7The impact of Masked Self-Distillation Mask (MSD) training was investigated. We comprehensively evaluated this training mechanism across all test datasets and decoding strategies. For both NAR and AR decoding strategies, when the model without self-jitter already performs well, MSD training slightly degrades performance, while when the original model performs poorly, MSD significantly improves performance. Specifically, the average improvement rates of MSD for AR and NAR are 14.47% and 18.03% respectively. The improvement in AR performance indicates that better learning of the masked predictor can help traditional RNN-T models overcome overfitting problems in challenging datasets while maintaining comparable performance in clean test sets. The improvement in NAR highlights the difficulty of learning without the predictor output, while MSD effectively helps the encoder and connector in the RNN-T model to better learn to directly encode audio. The evaluation of SAR decoding confirmed the general advantages of MSD. By fully leveraging the advantages of both decoding strategies, MSD achieved excellent results on all datasets, with an average improvement of 16.64%.

[0057] Effect of suppressing overfitting Figure 8 The recall values of the Xiaowen dataset at various low FAs are shown, which exhibits the most severe overfitting problem under strict test conditions.

[0058] In Figure 8 we reported the recall values at low FAs where the traditional RNN-T based KWS performed the worst on the Xiaowen test set to further analyze the overfitting problem. Compared with NAR and SAR, AR decoding performs significantly worse at very low FAs (#FA ≤ 6), while when the FAs increase, such as when FAs are 12 and 24 (97.72% vs. 98.37% and 99.49% vs. 99.46%), AR decoding finally achieves comparable results. In contrast, NAR decoding performs better under extremely low FA conditions, highlighting the ability of our training paradigm to suppress overfitting when masking the predictor output. Among all decoding strategies, SAR always has the best performance, indicating its strong stability and generalization ability under strict test conditions.

[0059] Limitations. Due to the small size of the KWS model, training the RNN-T model without a prediction network may be unstable. Although MSD training is adopted, it seems that additional attention needs to be paid to adjusting the training coefficients to ensure correct convergence.

[0060] Conclusion In the embodiments of this application, we use the output of the random masking predictor in the RNN-T KWS system to alleviate the overfitting problem in challenging scenarios. To implement masked sensor learning on the device, we propose a masked self-distillation (MSD) training paradigm that restricts the output or non-output of the prediction network. To fully utilize the advantages of these two training methods, we introduce an integrated decoding strategy SAR that combines the excellent performance of AR decoding on normal datasets and the robustness of NAR decoding in cases prone to overfitting. The results on three different KWS datasets show that our MSD training strategy and SAR decoding method significantly improve the performance of the RNN-T-based KWS model.

[0061] In some other embodiments, the embodiments of the present invention further provide a non-volatile computer storage medium storing computer-executable instructions that can execute the keyword detection model training and inference methods in any of the above method embodiments. As an implementation, the non-volatile computer storage medium of the present invention stores computer-executable instructions, and the computer-executable instructions are set as follows: In the training stage, randomly mask the latent representation of the text sequence to generate a version with some information hidden. Use the keyword detection model to calculate the predicted distribution through the joint network using this masked version and the unmasked version together, and the keyword detection model simultaneously calculates the distribution of the unmasked version as the teacher signal for self-distillation learning. In the inference stage, the keyword detection model supports two decoding modes: autoregressive decoding and non-autoregressive decoding. Among them, in the autoregressive decoding mode, the keyword detection model uses the output of the prediction network to generate the activation scores of the keywords. In the non-autoregressive decoding mode, the keyword detection model masks the output of the prediction network and only uses the encoder and the joint network to generate the activation scores; fuse the activation scores of these two decoding modes to obtain the final decoding result.

[0062] A non - volatile computer - readable storage medium may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the keyword detection model training and inference method and system, etc. In addition, the non - volatile computer - readable storage medium may include high - speed random - access memory, and may also include non - volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non - volatile solid - state storage devices. In some embodiments, the non - volatile computer - readable storage medium optionally includes a memory remotely set relative to the processor, and these remote memories can be connected to the keyword detection model training and inference method through a network. Examples of the above - mentioned network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0063] An embodiment of the present invention also provides a computer program product. The computer program product includes a computer program stored on a non - volatile computer - readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer is made to execute any one of the above - mentioned keyword detection model training and inference methods.

[0064] Figure 9 is a schematic structural diagram of an electronic device provided by an embodiment of the present invention, as Figure 9 shown, the device includes: one or more processors 910 and a memory 920. Figure 9 Here, one processor 910 is taken as an example. The device of the keyword detection model training and inference method and system may further include: an input device 930 and an output device 940. The processor 910, the memory 920, the input device 930, and the output device 940 can be connected through a bus or other means. Figure 9 Here, the connection through a bus is taken as an example. The memory 920 is the above - mentioned non - volatile computer - readable storage medium. The processor 910 executes various functional applications and data processing of the server by running non - volatile software programs, instructions, and modules stored in the memory 920, that is, implements the keyword detection model training and inference method in the above - mentioned method embodiments. The input device 930 can receive input digital or character information, and generate key signal inputs related to the user settings and function controls of the keyword detection model training and inference device. The output device 940 may include a display device such as a display screen.

[0065] The above - mentioned product can execute the method provided by the embodiment of the present invention, and has corresponding functional modules and beneficial effects for executing the method. Technical details not described in detail in this embodiment can be referred to the method provided by the embodiment of the present invention.

[0066] As an implementation, the above electronic device is applied to a keyword detection model training and inference device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to: In the training stage, randomly mask the latent representation of the text sequence to generate a version with some information hidden. Use the keyword detection model to calculate the predicted distribution through the joint network using this masked version and the unmasked version together. The keyword detection model simultaneously calculates the distribution of the unmasked version as the teacher signal for self-distillation learning. In the inference stage, the keyword detection model supports two decoding modes: autoregressive decoding and non-autoregressive decoding. Among them, in the autoregressive decoding mode, the keyword detection model uses the output of the prediction network to generate the activation score of the keyword. In the non-autoregressive decoding mode, the keyword detection model masks the output of the prediction network and only uses the encoder and the joint network to generate the activation score; fuse the activation scores of these two decoding modes to obtain the final decoding result.

[0067] The electronic device in the embodiment of the present application exists in various forms, including but not limited to: (1) Mobile communication devices: These devices are characterized by having mobile communication functions and mainly aim to provide voice and data communication. Such terminals include: smart phones, multimedia phones, functional phones, and low-end phones, etc.

[0068] (2) Ultra-mobile personal computer devices: These devices belong to the category of personal computers, have computing and processing functions, and generally also have the characteristic of mobile Internet access. Such terminals include: PDAs, MIDs, and UMPC devices, etc.

[0069] (3) Portable entertainment devices: These devices can display and play multimedia content. Such devices include: audio and video players, handheld game consoles, e-books, and smart toys and portable in-vehicle navigation devices.

[0070] (4) Servers: Devices that provide computing services. The composition of a server includes a processor, a hard disk, a memory, a system bus, etc. Servers are similar to general computer architectures, but due to the need to provide highly reliable services, they have higher requirements in terms of processing power, stability, reliability, security, scalability, and manageability.

[0071] (5) Other electronic devices with data interaction functions.

[0072] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative effort.

[0073] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods of each embodiment or some parts of the embodiments.

[0074] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments or equivalently replace some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for training and inferring a keyword detection model, the keyword detection model including an encoder, a prediction network, and a joint network, the method comprising: In the training phase, randomly mask the latent representation of the text sequence to generate a version with some information hidden; Use the keyword detection model to calculate the predicted distribution through the joint network using this masked version and the unmasked version together, and the keyword detection model simultaneously calculates the distribution of the unmasked version as the teacher signal for self-distillation learning; In the inference phase, the keyword detection model supports two decoding modes: autoregressive decoding and non-autoregressive decoding. Among them, in the autoregressive decoding mode, the keyword detection model uses the output of the prediction network to generate the activation scores of the keywords. In the non-autoregressive decoding mode, the keyword detection model masks the output of the prediction network and only uses the encoder and the joint network to generate the activation scores; Fuse the activation scores of these two decoding modes to obtain the final decoding result.

2. The method according to claim 1, wherein During the training process, the prediction network receives transcription information, while during the inference process, the prediction network receives keyword information.

3. The method according to claim 1, wherein, The randomly masking the latent representation of the text sequence to generate a version with some information hidden includes: According to a preset masking probability, mask the latent representation of the text sequence at the token level to obtain the masked version.

4. The method according to claim 3, wherein Use the logarithm corresponding to the unmasked version to teach the masked version.

5. The method according to claim 1, wherein Using the self-teacher-student learning paradigm helps to ensure that the masked output is basically the same as the output of the keyword detection model, so that the keyword detection model can learn how to produce accurate results without relying on the prediction network.

6. The method according to any one of claims 1-5, wherein The keyword detection model is an RNN-T model.

7. An electronic device, comprising: At least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the steps of the method according to any one of claims 1-6.

8. A storage medium having a computer program stored thereon, characterized in that, The computer program, when executed by a processor, implements the steps of the method according to any one of claims 1-6.