Keyword detection system training evaluation method, electronic device, and storage medium
Patent Information
- Application Number
- CN202510495086.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-19
- Publication Date
- 2026-09-08
- Estimated Expiration
- 2045-04-19
AI Technical Summary
具体来说,CTC通常应用于AED中的声学编码器,以提高收敛性并限制注意力模型产生过于灵活的输出
[0010] Thirdly, embodiments of the present invention also provide an electronic device comprising: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the steps of the method described in the first aspect.
Smart Images

Figure CN120412548B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of task-oriented dialogue technology, and particularly relates to a training and evaluation method for a keyword detection system, an electronic device, and a storage medium. Background Technology
[0002] In related technologies, keyword detection technology detects whether preset keywords appear in a continuous stream of speech, and wakes the device when a wake word appears. Keyword detection (KWS) is emphasized as wake word detection (WWD). Word Detection, designed to detect predefined keywords from continuous audio streams in memory- and computationally-constrained environments, serves as a primary interaction point for intelligent assistants. With the rapid development of multimodal large language models and end-to-end speech foundation models, creating more intelligent and personalized assistants has become increasingly feasible, attracting significant user attention. In this context, designing a powerful and efficient user interface has become increasingly crucial.
[0003] KWS systems can generally be broadly categorized into two types. The first type involves pattern matching or multi-label classification, where the KWS model processes audio blocks (or frames) in streaming mode, determining whether each block is activated by a specific keyword or identifying which keyword it belongs to. This approach is simple and straightforward, employing an end-to-end method. However, as a segmented classification model, this system is highly sensitive to data. While it may perform well in quiet or controlled environments, its performance is inconsistent across different scenarios, and it lacks robustness to noise.
[0004] Another widely adopted approach is to utilize Automatic Speech Recognition (ASR) within a two-stage framework. The ASR standard is used to train the acoustic model, while the decoding algorithm handles the frame-level posterior. Compared to classification-based training, sequence-to-sequence ASR training allows the acoustic model to adapt to different acoustic conditions, ensuring robust performance, especially in complex environments. Furthermore, ASR-based KWS supports arbitrary keyword detection, providing greater flexibility for user-defined scenarios. The main challenge for ASR-based systems lies in designing effective decoding algorithms for KWS tasks. Traditionally, researchers construct systems based on Weighted Finite-State Transducers (WFSTs). The decoding graph of the Transducer integrates keywords and fill paths. However, WFST-based decoding is complex to implement and maintain. Other methods, such as greedy search or prefix beam search for KWS based on Connectionist Temporal Classification (CTC), and greedy or autoregressive beam search for RNNT-based systems, do not explicitly prioritize keywords when exploring the entire search space, often leading to suboptimal results. Therefore, developing efficient algorithms for specific keywords for ASR-based streaming KWS remains a promising research direction.
[0005] RNN-T, also known as Transducer, is the mainstream architecture for ASR and KWS in academia and industry. Its inherent support for streaming inference makes it well-suited for KWS where real-time feedback is critical. Related technologies have proposed TDT-KWS, a Transducer-based system that is faster and more accurate than the original RNN-T-based system. This is achieved through two key modifications: (1) replacing the standard RNN-T with a Token-and-Duration Transducer (TDT) that simultaneously predicts the token and its duration; and (2) designing a streaming decoding algorithm specifically for Transducer-based KWS. The inventors found that while TDT-KWS significantly outperforms the strong Transducer-based baseline, there is still room for improvement. Error accumulation can occur because the posterior prediction of the current frame depends on the history of the prediction sequence. This reliance on previous predictions can hinder the model's ability to distinguish between positive samples and challenging negative samples, leading to performance degradation under complex acoustic conditions. Further research is needed to ensure stable performance under challenging conditions such as noisy environments or arbitrary keyword detection.
[0006] Joint multi-task training of CTC with AED (AutoEncoder Decoder) (CTC-AED) or CTC with Transducer (CTC-Transducer) has proven to outperform single-architecture models and has become a standard paradigm for ASR. Multi-task systems fully leverage the strengths of both branches while mitigating their respective limitations during training. Specifically, CTC is often applied to the acoustic encoder in AED to improve convergence and limit overly flexible outputs from attention-based models. Summary of the Invention
[0007] This invention provides a keyword detection system training and evaluation method, electronic device, and storage medium to at least solve one of the above-mentioned technical problems.
[0008] In a first aspect, embodiments of the present invention provide a training and evaluation method for a keyword detection system, comprising: collecting a fixed keyword dataset and an arbitrary keyword dataset, and data augmentation under noisy conditions; constructing a multi-head frame asynchronous keyword detection model based on a CTC-Transducer joint training framework, wherein during training, a multi-task learning strategy is used to combine the CTC loss and the Transducer loss to optimize the performance of the multi-head frame asynchronous keyword detection model; employing a multi-head frame asynchronous decoding strategy to fuse the decoding results of CTC and Transducer; evaluating the model on the fixed keyword dataset and the arbitrary keyword dataset, and analyzing the performance of different fusion strategies and decoding methods under noisy conditions to verify the effectiveness and robustness of the multi-head frame asynchronous keyword detection model in different scenarios.
[0009] Secondly, embodiments of the present invention also provide a computer program product, the computer program product including a computer program stored on a non-volatile computer-readable storage medium, the computer program including program instructions, which, when executed by a computer, cause the computer to perform the steps of the keyword detection system training and evaluation method of any embodiment of the present invention.
[0010] Thirdly, embodiments of the present invention also provide an electronic device comprising: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the steps of the method described in the first aspect.
[0011] Fourthly, embodiments of the present invention also provide a storage medium storing a computer program thereon, characterized in that the computer program, when executed by a processor, implements the steps of the method described in the first aspect.
[0012] In the method of this application embodiment, the MFA-KWS training framework is based on multi-task learning, combining CTC and Transducer architectures to help reduce error accumulation and improve the model's robustness in complex acoustic environments. Simultaneously, the Transducer architecture can perform more refined modeling of speech signals, capturing the temporal characteristics of speech. During training, the model employs a joint loss function, combining CTC loss and Transducer loss, and optimizing model performance by adjusting their weight coefficients. Furthermore, to further enhance the model's expressive power and adaptability, TDT is introduced, which can predict not only speech units but also their duration, enabling the model to more accurately capture the duration information of keywords. Further, the decoding strategy employs a multi-head-frame asynchronous decoding method, fusing the decoding results of CTC and Transducer. Moreover, regarding the fusion strategy, multiple methods are designed to more effectively combine the advantages of CTC and Transducer, improving decoding accuracy and robustness. Attached Figure Description
[0013] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0014] Figure 1 The processing flow provided in one embodiment of the present invention; Figure 2 A flowchart illustrating an embodiment of a keyword detection system training and evaluation method provided by an embodiment of the present invention; Figure 3 An overview of the MFA-KWS framework provided in one embodiment of the present invention; Figure 4 This invention provides a Transducer-based streaming decoding algorithm according to an embodiment of the present invention. Figure 5 This is an MFS streaming decoding algorithm provided in an embodiment of the present invention; Figure 6 This is the decoding path of an RNN-T KWS system provided in an embodiment of the present invention; Figure 7 The selected keywords in librikws-20 provided in an embodiment of the present invention; Figure 8 This invention provides a performance comparison of different frame skipping rates on the SNIPS dataset with FAR=0.02, as provided in one embodiment of the invention. Figure 9This invention provides a comparison of recall rates of a training system on the SNIPS dataset at different false acceptance rates (FAR) levels, as provided in an embodiment of the present invention. Figure 10 This invention provides a comparison of recall rates for different fusion strategies of multi-head decoding in three test sets when the FA is fixed at 2, as an embodiment of the present invention. Figure 11 A comparison of recall performance of a joint model based on RNN-T and TDT provided in an embodiment of the present invention when FA=2; Figure 12 A comparison of recall rates between a proposed multi-head system and end-to-end KWS systems at different levels on the HEY-SNPS dataset, as provided in an embodiment of the present invention. Figure 13 A comparison of the recall rates of the Mandarin dataset mobvoihotwords provided in an embodiment of the present invention at FAR = 0.5 / h; Figure 14 The wake-up score heatmap of the Transducer and CTC branch at each (t,u) provided in an embodiment of the present invention, and the joint decoding score of MFA; Figure 15 A performance comparison of MFA-KWS with ASR-based CTC and Transducer KWS baselines provided in an embodiment of the present invention; Figure 16 Performance comparison of different decoding strategies provided in an embodiment of the present invention under different noise conditions on Snips, test-clean and test-other datasets with fixed FA=2; Figure 17 This invention provides a comparison of the relative speeds of all decoding algorithms across all test sets, using MFS decoding speed as a benchmark; Figure 18 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0015] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0016] The inventors discovered that the related technologies have at least the following drawbacks: First, the problem of low wake-up rate with low false wake-up count is particularly obvious in complex acoustic scenarios with high noise and low signal-to-noise ratio; second, the RNN-T (Recurrent Neural Transducer) / TDT-KWS decoding algorithm, which does not have a CTC (Connectionist Temporal Classification) decoding module, is prone to false wake-ups on negative example datasets; third, it only supports predefined keyword detection and cannot be extended to arbitrary keyword scenarios defined by users.
[0017] The above-mentioned defects are mainly due to the following reasons: on the one hand, the decoding algorithm is simple and is prone to false wake-up of negative audio; on the other hand, the model structure is not advanced enough and there is still room for optimization.
[0018] When addressing these shortcomings, industry professionals might increase training data to train better-performing and more robust models. They might also design models with larger parameters and higher computational demands, training larger models to mitigate overfitting and thus enhance performance in low signal-to-noise ratio scenarios. Furthermore, they might design multi-task, multi-level systems to further enhance performance.
[0019] The reason why professionals in this industry might not readily conceive of the solution proposed in this application is as follows: First, traditional decoding strategies are typically singular, either based on CTC or on a Transducer. The multi-head decoding strategy proposed in this application fuses the decoding results of CTC and Transducer. This innovative decoding approach requires a deep understanding of the advantages and disadvantages of both strategies and the design of an effective fusion method. Second, no research team has previously studied this problem in this complex scenario using this approach, and there is very little prior work to refer to. Third, frame-asynchronous decoding strategies in MFA-KWS (Multi-head Frame-asynchronous) such as TDT (Token-and-duration Transducer) and PSD (Fhone-Synchronous Decoding) can skip irrelevant frames, improving decoding efficiency. This frame-asynchronous decoding method is rare in traditional decoding strategies and requires in-depth research into the characteristics of speech signals and the decoding process to design an effective frame-asynchronous decoding algorithm.
[0020] In this embodiment, on the one hand, the training framework of MFA-KWS is based on multi-task learning, combining CTC (Connectionist Temporal Classification) and Transducer architectures. This joint training framework, by introducing the conditional independence assumption of CTC, helps reduce error accumulation and improves the robustness of the model in complex acoustic environments. Simultaneously, the Transducer architecture can perform more refined modeling of speech signals, capturing the temporal characteristics of speech. During training, the model uses a joint loss function, combining the CTC loss and the Transducer loss, and optimizing model performance by adjusting their weight coefficients. Furthermore, to further enhance the model's expressive power and adaptability, TDT is introduced, which can predict not only speech units but also their duration, enabling the model to more accurately capture the duration information of keywords. On the other hand, the decoding strategy employs a multi-head frame asynchronous decoding method, fusing the decoding results of CTC and Transducer. Specifically, the CTC branch skips frames with high blank probabilities by setting a blank probability threshold, thereby accelerating the decoding process while maintaining performance. The Transducer branch employs TDT for decoding, which can skip irrelevant frames based on the predicted duration, further improving decoding efficiency. Regarding fusion strategies, multiple methods were designed to more effectively combine the advantages of CTC and Transducer, enhancing decoding accuracy and robustness.
[0021] Please refer to Figure 1 The document illustrates the processing flow of an embodiment of this application.
[0022] In this embodiment, firstly, data preparation is performed, including the collection of datasets with fixed and arbitrary keywords, and data augmentation in noisy environments. Then, an MFA-KWS model based on the CTC-Transducer joint training framework is constructed. During training, a multi-task learning strategy is used to combine the CTC loss and the Transducer loss to optimize model performance. Next, a multi-head frame asynchronous decoding strategy is employed to fuse the decoding results of CTC and Transducer. Finally, the model is evaluated on multiple datasets, with evaluation metrics including recall and precision. The performance of different fusion strategies and decoding methods in noisy environments is analyzed, ultimately verifying the effectiveness and robustness of the MFA-KWS scheme in different scenarios.
[0023] The MFA-KWS scheme in this application demonstrates comprehensive superiority in terms of performance. Its direct effect is the accurate and efficient identification of predefined keywords in continuous speech streams. Compared to traditional methods, it achieves superior performance on both fixed-keyword and arbitrary-keyword datasets, such as higher recall and accuracy on datasets like Snips, MobvoiHotwords, and LibriKWS-20, while maintaining strong robustness in noisy environments. This scheme not only improves the accuracy and efficiency of keyword detection but also provides a more powerful front-end interaction entry point for intelligent interactive devices such as voice assistants, enabling users to wake up and operate devices more conveniently and accurately, greatly optimizing the user experience. At a deeper level, the MFA-KWS scheme promotes the development of speech recognition technology, providing new ideas and methods for the deep integration of multimodal interaction and intelligent devices, helping to build a more intelligent and natural human-computer interaction ecosystem, and thus promoting the progress of the entire intelligent technology field.
[0024] Please refer to Figure 2 The flowchart shown is an embodiment of the keyword detection system training and evaluation method of this application.
[0025] like Figure 2 As shown, in step 201, a fixed keyword dataset and an arbitrary keyword dataset are collected, as well as data augmentation under noisy conditions; In step 202, a multi-head frame asynchronous keyword detection model based on the CTC-Transducer joint training framework is constructed. During the training process, a multi-task learning strategy is used to combine the CTC loss and the Transducer loss to optimize the performance of the multi-head frame asynchronous keyword detection model. In step 203, a multi-header frame asynchronous decoding strategy is adopted to fuse the decoding results of CTC and Transducer; In step 204, the evaluation is performed on the fixed keyword dataset and the arbitrary keyword dataset, and the performance of different fusion strategies and decoding methods in noisy environments is analyzed to verify the effectiveness and robustness of the multi-head frame asynchronous keyword detection model in different scenarios.
[0026] In some optional embodiments, the construction of the multi-head frame asynchronous keyword detection model based on the CTC-Transducer joint training framework includes: extending the Transducer-based keyword detection model to the CTC-Transducer joint training framework and multi-head frame synchronous decoding to achieve robust performance in challenging scenarios including hard negative samples and noisy conditions. The multi-head frame synchronous decoding is transformed into a multi-head frame asynchronous decoding framework, which integrates the phoneme synchronous decoding of the CTC branch and the duration converter of the Transducer branch.
[0027] In some optional embodiments, the fusion of the decoding results of CTC and Transducer includes: the CTC branch skips frames with high blank probabilities by setting a blank probability threshold, thereby accelerating the decoding process while maintaining performance; the Transducer branch uses a duration converter for decoding, which can skip irrelevant frames according to the predicted duration, further improving decoding efficiency.
[0028] In some optional embodiments, the different fusion strategies include a naive score fusion strategy and a consistency-based fusion strategy. The naive score fusion strategy includes CTC-Dom, Transducer-Dom, and Equivalence-Dom, and the consistency-based fusion strategy includes CDC-Zero and CDC-Last. The consistency-based fusion strategy more effectively integrates the Transducer and CTC branches.
[0029] In some optional embodiments, the improvement to the multi-header frame asynchronous decoding strategy includes: adding a reward score S. bonus This is done to sharpen the activation time boundary and improve the accuracy of keyword endpoints; T is introduced. out To filter out excessively long decoding paths; the score of the best path is normalized according to its length and converted into a confidence level in the range of [0,1]. The decoding process is still fully streaming with no forced delay. In practice, the decoding grid state remains unchanged, and each step only inputs a single frame logarithm or posterior to the decoding module.
[0030] In some alternative embodiments, the Transducer includes an encoder, a predictor, and a joint network, wherein the encoder converts speech signals into latent acoustic representations, the predictor acts as a language model to process text input and provide text information in an autoregressive manner, and the latent acoustic representations and the text information are then combined by the joint network to generate a probability distribution for the next label.
[0031] In this embodiment, the MFA-KWS training framework is based on multi-task learning, combining CTC and Transducer architectures to reduce error accumulation and improve the model's robustness in complex acoustic environments. Simultaneously, the Transducer architecture enables more refined modeling of speech signals, capturing the temporal characteristics of speech. During training, the model employs a joint loss function, combining CTC and Transducer losses, and optimizing model performance by adjusting their weight coefficients. Furthermore, to further enhance the model's expressive power and adaptability, TDT is introduced, which can predict not only speech units but also their duration, enabling the model to more accurately capture the duration information of keywords. Further, the decoding strategy employs a multi-head-frame asynchronous decoding method, fusing the decoding results of CTC and Transducer. Regarding the fusion strategy, multiple methods are designed to more effectively combine the advantages of CTC and Transducer, improving decoding accuracy and robustness.
[0032] It should be noted that the above method steps are not intended to limit the execution order of each step. In fact, some steps may be executed simultaneously or in the reverse order of the steps, and this application does not impose any restrictions on this.
[0033] The following description uses a specific embodiment to illustrate the solution of this application, so that those skilled in the art can better understand the solution of this application.
[0034] Keyword detection (KWS) is a core technology for voice-driven applications, demanding high accuracy and efficiency. Traditional ASR-based KWS methods (such as greedy search and beam search) traverse the entire search space without explicitly prioritizing keyword detection, resulting in poor performance. In this application, we propose an effective keyword-specific KWS framework: a frame asynchronous system based on a streaming CTC Transducer, employing Multi-Head Frame Asynchronous Decoding (MFA-KWS). Specifically, MFA-KWS uses dedicated phoneme synchronous decoding (PSD) for keywords to optimize the CTC branch and replaces the traditional RNN-T (Recurrent Neural Network Transducer) with a Token-and-Duration Transducer (TDT), thereby improving performance and efficiency. Furthermore, we explore various score fusion strategies, including single-frame-based and consistency-based methods. Extensive experiments demonstrate that MFA-KWS achieves state-of-the-art (SOTA) performance on both fixed-keyword and arbitrary-keyword datasets, such as Snips, Mobvoi Hotwords, and LibriKWS-20, while exhibiting strong robustness in noisy environments. Among fusion strategies, the consistency-based CDC-Last method shows the best performance. Furthermore, MFA-KWS achieves inference speedups of 47%–63% compared to frame-synchronous baseline models on multiple datasets. These extensive experimental results further validate that MFA-KWS is an efficient KWS framework suitable for deployment on edge devices.
[0035] While related technologies have proposed multi-head decoding strategies, such as CTC / AED decoding, to further improve transcriptional accuracy during inference, the joint CTC-AED system is not universally applicable to our KWS objective because the attention mechanism is computationally expensive, and label-level inference introduces significant latency. Therefore, designing an efficient joint framework and multi-head decoding algorithm for KWS remains a challenge.
[0036] In this application, we propose Multi-head Frame Synchronous Decoding (MFS), extending the previous Transducer-based KWS system into a CTC-Transducer joint framework. To further improve KWS performance and computational efficiency, we introduce Multi-head Frame Asynchronous Decoding (MFA). Compared to a single Transducer-based KWS system, the condition-independent nature of CTC reduces error accumulation and ensures more stable performance in various challenging scenarios. Inspired by Phoneme Synchronous Decoding (PSD) in CTC and Term and Duration Transformers (TDT) in Transducer-based models, we introduce PSD for KWS for CTC streaming inference and term and duration prediction for Transducer streaming inference. These mechanisms enable frame skipping, improve robustness under noisy conditions, and accelerate decoding speed.
[0037] Among them, TDT-KWS introduces the KWS token and duration converter and the Transducer-based stream decoding algorithm. In the embodiments of this application, we extend the Transducer-based method to the CTC-Transducer joint framework, further improving performance and robustness. The main contributions are as follows: - We propose an efficient multi-task KWS system, MFS-KWS. Building on previous work, we extend Transducer-based KWS to a CTC-Transducer joint training framework and multi-head frame synchronous decoding, achieving robust performance in challenging scenarios including hard negative samples and noisy conditions.
[0038] To further improve performance and accelerate decoding speed, we propose MFA-KWS, a multi-header frame asynchronous decoding framework that integrates PSD from the CTC branch and TDT from the Transducer branch. To fully leverage the advantages of MFA-KWS, we designed and explored several novel fusion strategies.
[0039] We evaluated MFA-KWS on fixed-keyword English and Mandarin datasets (Hey Snips and MobvoiHotwords) and on arbitrary-keyword datasets (LibriKWS-20, derived from LibriSpeech). The results show that MFA-KWS achieves state-of-the-art (SOTA) performance on various datasets. Furthermore, evaluations at multiple signal-to-noise ratio (SNR) levels demonstrate that MFA-KWS significantly outperforms single-branch systems in terms of noise robustness.
[0040] - Compared to the frame-synchronous MFS framework, the proposed MFA-KWS improves inference speed by 47%-63% on different test sets. Compared to the single-branch KWS framework, MFA-KWS achieves a better balance between performance and efficiency, demonstrating its strong potential in on-device KWS applications.
[0041] - Open-source KWS decoding research remains limited, especially KWS decoding algorithms. To promote the development of KWS, we have released all decoding code on Github, including MFA, Transducer stream decoding, and other related algorithms (such as CTC dedicated stream decoding and various fusion strategies).
[0042] method This section provides a comprehensive overview of our work. We begin by introducing the design of Transducer-based KWS, detailing its architecture and decoding algorithm. Next, we propose MFS-KWS, leveraging multi-task learning within the CTC-Transducer joint training framework to enhance TDT-KWS. We also propose a streaming multi-head decoding algorithm for MFS-KWS, which merges scores in a frame-synchronized manner. Finally, we extend MFS-KWS to MFA-KWS, achieving more efficient and effective frame-synchronized decoding.
[0043] Transducer-based Keyword Spotting System (Sensor-based) A transducer consists of three parts: an encoder, a predictor, and a connector (or joint network). The encoder converts the speech signal into a latent acoustic representation, while the predictor acts as a language model, processing the text input (typically phonemes, syllables, or subwords) to provide linguistic context in an autoregressive manner. The latent sound and text information from these modules are then combined by the connector to generate the probability distribution for the next token. Given an audio input x = {x1, x2, ..., x...} T} ∈R T×D , where each x t It is a D-dimensional acoustic feature vector, where T represents the total number of frames, and a transcription y = {y1, y2, ..., y} U} ∈ R U×1 , where each x t Let x be a D-dimensional acoustic feature vector, and T represent the total number of frames. The optimization objective of Transducer is to maximize the log probability of ground-truth transcription y, given x. Figure 3 An overview of the MFA-KWS framework is shown. Among them, Figure 3 The left side shows the MFA-KWS framework, including the training and inference processes. Figure 3 The right side of the diagram introduces various fusion strategies for multi-head decoding, detailed in subsequent sections. PH (placeholder) indicates a special state where time steps are skipped during frame synchronization decoding. Blue blocks represent CTC scores, orange blocks represent Transducer scores, and red blocks represent activation steps.
[0044] Among them, B RNN-T Defined as a mapping from a valid Transducer-based augmented aligned πRNN-T (including special whitespace φRNN-T) to the label y, while B RNN-T -1 It refers to inverse mapping.
[0045] Previous Transducer-based KWS methods employed ASR inference strategies, such as greedy search or beam search, without considering the uniqueness of KWS. Therefore, these methods were not ideal for frame-level keyword detection. ASR aims to generate complete transcripts, while KWS focuses solely on detecting the presence of predefined keywords. To address this issue, the proposed streaming decoding method provides the predictor with only the keyword-tagged sequence, rather than providing partial hypotheses as in ASR decoding. This design is similar to the teacher-forcing strategy used in RNN-T loss calculation and autoregressive model training. By restricting the decoding space to the keyword search range, this approach reduces error accumulation and improves detection accuracy. We define the acoustic feature x = {x1, x2, ..., x...} T As before, the keyword transcription y={y0=φRNN-T,y1,y2,...,y U (y0 represents a blank symbol). Following the standard Transduce literature, we represent the token / blank emission probability of the decoding grid as follows: and For t∈[1,T] and u∈[0,U].
[0046] For the Transducer model, decoding uses a decoding grid, such as... Figure 6As shown. We define δ(t,u) as the path with the highest score among all paths to node (t,u). In streaming decoding, keywords can start from any point in the speech stream, so their start time needs special attention. To address this, we assign 1 point to δ(t,0) at each time step t, allowing keyword detection to begin at any time and facilitating seamless wake word detection in continuous speech. Figure 4 As shown, through dynamic programming, we can efficiently calculate the complete path score δ(t,U) of a keyword. Then, multiplying the path score δ(t,U) by the blank score φ(t,U) yields the keyword confidence at time t. Where φ(t,U) represents the final blank score at time t. Compared with related techniques, we made three improvements to the decoding algorithm. First, we added a reward score S. bonus This is done to sharpen the activation time boundary and improve the accuracy of keyword endpoints. Secondly, we introduce T... out This filters out excessively long decoding paths. Furthermore, we normalize the optimal path score by its length, converting it to a confidence level in the range [0,1]. This normalization is beneficial for the fusion process in multi-head decoding. For clarity, we describe... Figure 4 Assume the input consists of a complete posterior matrix p = [p1, p2, ... p2]. T ] ∈R T×U×V The vocabulary is composed of V, where V represents the vocabulary size including whitespace. However, the decoding process remains fully streaming, with no forced latency. In practice, the decoding grid state remains constant, with only a single frame logarithm or posterior being input to the decoding module at each step. See [link to documentation] for more details. Figure 4 .
[0047] Figure 4 This invention provides a streaming decoding algorithm based on a Transducer, as an embodiment of the present invention.
[0048] Figure 5 This is a multi-head streaming decoding algorithm provided in an embodiment of the present invention.
[0049] Figure 6The decoding path of the RNN-T KWS system is shown. Each node (t, u) represents the highest score δ(t, u) obtainable when the first u elements of the keyword are output at time t. The horizontal arrows emanating from node (t, u) represent the probability φ(t, u) of outputting a blank character, while the vertical arrows represent the probability y(t, u) of outputting the (u+1)th keyword element at time t. The optimal keyword path at time t (as shown by the red arrows) extends along the sequence with the highest score, corresponding to the keyword most likely to appear at that time.
[0050] Multi-head Frame-synchronous KWS System In this section, we will introduce how to extend the Transducer-based KWS system to the Multi-Head Frame Synchronization (MFS) CTC Transducer joint framework. CTC is a sequence-to-sequence training criterion widely used in non-autoregressive ASR models. Similar to RNN-T loss, CTC introduces a special blanking label φ. CTC It simulates non-speech output to facilitate alignment learning between speech and transcription. However, its registration transformation rules differ from those of RNN-T. Using the speech input x and corresponding label y defined in the aforementioned embodiments, the CTC loss is calculated as follows: Here, B CTC Align the effective CTC with π CTC (including φ) CTC ) is mapped to the label sequence y, and This indicates its reverse mapping. The KWS model based on CTC-Transducer is trained using a combination of RNN-T and CTC loss, as follows: Wherein, α is a coefficient that controls the influence of CTC. In the embodiments of this application, we set α to 0.3 by default.
[0051] CTC assumes conditional independence along the time axis, predicting the distribution of each frame without historical context. While this limits ASR (contextual information ensures semantic consistency), it can be beneficial for KWS by reducing error accumulation and mitigating false positives, thus preventing overfitting of keywords. Furthermore, CTC can serve as a regularization mechanism for acoustic modeling, improving the convergence of multi-task training frameworks.
[0052] In terms of inference, we propose a Multi-Head Frame Synchronization (MFS) decoding algorithm, extending Transducer-based decoding to a Transducer-CTC joint strategy. The CTC branch decoding concept, referred to as CTC-stream decoding in this embodiment, is a frame-level stream decoding strategy specifically designed for CTC-based KWS systems. Figure 4 Building upon the Transducer-based algorithm previously proposed, we draw inspiration from CTC-based methods and introduce MFS decoding. This performs inference on both branches simultaneously and merges frame-level confidence scores. Inference details are detailed in Algorithm 2. It's worth noting that while the pseudocode iterates over frame index t, this does not result in redundant computation in practical applications. The loop structure is included purely for clarity.
[0053] Multi-head frame-asynchronous system To improve performance in challenging environments and reduce computational overhead for faster inference, we extend the MFS KWS to a multi-head frame-asynchronous (MFA) version. Specifically, we modify the frame-level RNNT and CTC submodules to token-level systems. We introduce a variant of the token-and-duration transducer (TDT), which models token durations when predicting tokens, replacing the traditional RNN-T. For the CTC branch, we replace frame-synchronous decoding (FSD) with phoneme-synchronous decoding (PSD), which accelerates inference by skipping frames with a blank probability exceeding a preset threshold. Both methods skip irrelevant frames and improve decoding efficiency by reducing unnecessary frame computations.
[0054] TDT enhances the traditional Transducer by incorporating marker duration prediction into the connector output. Specifically, the traditional RNN-T simulates a posterior distribution P(v|t,u), where v represents a marker or blank symbol in the vocabulary φRNN-T, and t and u represent the acoustic frame and text marker indices, respectively. TDT extends this to a joint distribution P(v,d|t,u), where d represents the predicted marker duration at position (t,u). TDT treats duration prediction as a classification task, where d ranges from {0,1,2,---,D}. max Choose D from} max These are the predefined hyperparameters for duration modeling. The conditional independence assumption is adopted, resulting in factorization: Where PT(-) and PD(-) represent the label distribution and duration distribution, respectively. When training the KWS system based on CTC-TDT, the RNN-T loss L in equation (6) needs to be changed. RNN-T And RNN-T-based decoding is replaced with TDT loss L TDT And TDT-based decoding, while keeping all other parts unchanged. MFA loss is defined as Among them, L TDT For more detailed information, please refer to the relevant technical documentation.
[0055] For the CTC branch, we replace FSD with PSD, utilizing the peak characteristics of the CTC output to filter out blank frames irrelevant to the key speech frames. The probability p(π|x) of a complete registration can be divided into combinations of non-blank frames and blank frames, i.e. Here, non-blank frames represent frames containing speech information, while blank frames represent non-speech frames. Furthermore, we can utilize the peak phenomenon of CTC (Confirmation-Transcription-Based Transmission), where the prediction confidence and maximum probability are almost close to 1 within each time step, thus skipping some frames with high blank probabilities and accelerating decoding. Where, λ φ This is a preset threshold used to filter frames with a very high CTC blank probability. The total path can be approximated by frames containing speech content with a low blank probability. This is called Phoneme-Synchronous Decoding (PSD) for CTC decoding, a technique that maintains performance while significantly improving decoding speed.
[0056] The threshold for filtering blank frames determines the percentage of frames retained, and efficiency and performance can be balanced by adjusting the threshold. Compared to each original system (TDT vs RNN-T, PSD vs FSD), both TDT and PSD gain additional tag duration modeling capabilities. Combined with the fusion strategy proposed in the next section, the joint KWS system based on CTC-Transducer evolves into a multi-head frame asynchronous (MFA) KWS system.
[0057] Multi-head Joint Decoding In our MFA decoding framework, since the detection scores from CTC and Transducer are not aligned along the time axis, frame-synchronized decoding must address two key challenges: how and when to fuse the scores. The fusion method determines the system's effectiveness, while the timing of the fusion affects its efficiency. In this section, we will examine several different MFA decoding fusion strategies. The basic idea of score fusion can be expressed as follows: Where ⊕ represents the frame-by-frame fusion operation, S t Trans and S t CTC These represent the confidence scores of the TDT and CTC decoding strategies at time t, respectively. To handle frames skipped by TDT or PSD, we introduce a placeholder (PH), a special state representing the decoding score of the skipped frame. The placeholder value is determined by the selected fusion strategy. Since the score can be frame-synchronized or asynchronous, the fusion strategy must be carefully designed. We explore several fusion methods, such as... Figure 3 The details are shown on the right side below.
[0058] - CTC-Dom (CTC-Domination): The final fusion confidence score is primarily determined by the CTC branch. If the CTC branch produces a normal score S at time t... t CTC If so, it is directly used as the frame's built-in confidence. If the CTC branch skips the frame, and S t CTC If the score equals PH, then the score will be branched by the Transducer S. t Trans Fill in the blanks. If both branches are PH at time t, it means that both branches believe the current frame does not contain meaningful keyword information. In this case, we set the confidence score to 0 and mark it as an empty frame.
[0059] - Transducer-Dom: The fusion confidence score is primarily determined by the Transducer branch, the opposite of the logic in CTC-Dom. If the Transducer branch provides a normal score S... t Trans This will be directly used as the confidence level of the frame. If the Transducer outputs S... t Trans For pH value, then from CTC branch S t CTC Obtain the score. If both branches are PH, the confidence score is set to 0, and the frame is marked as empty.
[0060] - Equivalence-Dom: Both branches contribute equally to the final confidence score. If both branches provide the decoding score for frame t, their average is used. If only one branch provides a score, that score is used directly. If neither branch is a PH (Positive Confidence), 0 points are assigned to fill the frame.
[0061] These fusion strategies operate on a single frame basis, making them simple, straightforward, and easy to interpret. However, they do not consider historical activation states. High activation scores in isolated frames may indicate false alarms. Incorporating historical decoding states into the fusion strategy can improve reliability and discriminative power. Cross-Layer Discriminative Consistency (CDC) is a fractional fusion strategy that leverages the different behaviors of intermediate and output layers to improve performance. We extend CDC by capturing fractional trend discriminative power within a sliding window between CTC and the Transducer, using the degree of discriminative power as a weighting coefficient for fractional fusion. At time step t, the coefficient w... t CDC The calculation method is as follows: in, Indicates window size, f sim (-) is the cosine similarity function. Therefore, the fusion score is S. t Trans and S t CTC Weighted sum: For MFA decoding, PH must be replaced with an appropriate value before applying CDC-based fusion. We consider two consistency-based strategies: - CDC-Zero: In this method, if a PH is encountered during CDC calculation, it is assumed that no speech event has occurred in that frame, and 0 points are assigned.
[0062] - CDC-Last: In this method, skipped frames are treated as continuations of the most recent state. When performing a PH replacement, the nearest previous non-placeholder value is used.
[0063] We investigated multi-head decoding using various fusion strategies on multiple keyword datasets. Results are presented later.
[0064] Experimental setup Data Configuration We evaluate the MFA-KWS framework in different scenarios: 1) Fixed single-keyword utterances in English and Mandarin. 2) Arbitrary keyword detection in continuous speech, covering 20 different keywords. 3) Keyword detection in noisy environments. To ensure the comprehensiveness of the evaluation, we use multiple widely adopted datasets, the details of which are as follows.
[0065] - Hey-Snips. The Hey Snips dataset is an open-source KWS dataset that uses "Hey Snips" as the keyword, which is pronounced as a single phrase with no pause between the two words. The training set, development set and test set contain 5,799, 2,484 and 2,529 positive utterances respectively, and 44,860, 20,181 and 20,543 negative utterances respectively. Since the complete text of the negative utterances cannot be obtained, these segments are only used for false positive evaluation, not for training. The false positive dataset is constructed by aggregating all negative segments from the training set, development set and test set, yielding a total of approximately 97 hours of audio.
[0066] - MobvoiHotwords. MobvoiHotwords is a Mandarin KWS corpus that contains two keywords: "Xiaowen" (Hi Xiaowen) and "Wenwen" (Nihao Wenwen), as well as non-keyword speech. The corpus consists of approximately 262 hours of audio from 287k utterances. This dataset was recorded at different distances from smart speakers, with background noise at different signal-to-noise ratios (SNR), such as domestic sounds (e.g., music, television). In the training set, development set and test set, each keyword contains approximately 21.8k, 3.7k and 10.6k positive utterances respectively, and 131k, 31.2k and 52.2k negative utterances respectively. The negative test set contains approximately 63 hours of audio.
[0067] - LibriSpeech. LibriSpeech is a widely used speech corpus that contains 960 hours of read speech and corresponding transcripts. In addition to training a phoneme-based acoustic model on LibriSpeech, we selected 20 specific words from its two test sets as keywords to simulate arbitrary keyword detection in continuous speech. We refer to this evaluation dataset as LibriKWS20, which includes two subsets: test-clean and test-other. Both subsets contain the same 20 keywords, as listed in Table 1. To construct the false positive dataset, we merged the audio samples that do not contain the selected keywords from the test set, thereby obtaining datasets with a total duration of 3 hours respectively.
[0068] - AISELL-2. AISHELL-2 is an industry-scale ASR corpus containing 1000 hours of Mandarin speech. We used AISHELL-2 to train a good Mandarin acoustic seed model for subsequent KWS experiments with the words "Hi Xiaowen" and "Nihao Wenwen".
[0069] The WHAM! WHAM! dataset is a corpus of near-field and far-field ambient noise recorded in urban environments, including music, musical instruments, and background noise from restaurants and bars. It captures different acoustic conditions in various real-world scenarios. To evaluate the robustness of our models in noisy environments, we mixed the test portion of the WHAM! dataset with the Hey Snips and LibriKWS-20 datasets at different signal-to-noise ratio levels.
[0070] Data Preparation. For experiments with arbitrary keywords, we trained the model on LibriSpeech-960h and evaluated it on LibriKWS-20. For experiments with fixed keywords using Snips or MobvoiHotwords, we first pre-trained the model on the two general ASR datasets, LibriSpeech-960h or AISHELL-2. Then, we fine-tuned the model using positive samples from Hey Snips or Xiaowen / Wenwen, combined with an equal amount of non-keyword sentences as general data. For Mandarin MobvoiHotwords, due to its inherently challenging acoustic conditions, we needed to use its own non-keyword dataset as general data. Since the original MobvoiHotwords did not provide transcriptions of negative datasets, we used pseudotranscriptions generated by a powerful Mandarin ASR model and then fine-tuned them using randomly selected common ASR utterances.
[0071] In addition, we evaluated the performance under various noise conditions using the English Hey Snips and LibriKWS-20 datasets. (Note: MobvoiHotwords is not included because the original data contains noise, and the SNR adjustment cannot be controlled when adding additional noise). To simulate noisy conditions, each clean waveform from Snips and LibriSpeech was mixed with randomly selected noise samples from WHAM!, with the SNR taken from a uniform distribution of 0 to 20 dB. The generation process for the noisy negative test set was the same as the training process, while the positive test set was simulated at a fixed SNR level. Specifically, all positive test samples from Snips and LibriSpeech were mixed with noise at an SNR of {0, 5, 10, 15, 20} dB.
[0072] Figure 7 The selected keywords in librikws-20 are shown. The keywords are the same in both the TEST-CLEAN and TEST-OTHER datasets.
[0073] Training configuration Acoustic Features. The input acoustic features consist of 40-dimensional log-Mel filter bank coefficients (FBank), extracted with a window length of 25 milliseconds and a window jump of 10 milliseconds. During training, we employed two data augmentation strategies. We used online speech perturbation and randomly selected warp factors from the set {0.9, 1.0, 1.1}. Additionally, we employed SpecAugment with a maximum frequency mask range of F = 10 and a maximum time mask range of T = 50. Specifically, we used two masks of each type for a single data sample. We concatenated five frames from the left and right contexts to construct a 440-dimensional feature as the context acoustic frame, and set the frame jump parameter to 3 to perform cubic subsampling to reduce computational overhead.
[0074] Model Architecture. For the shared encoder, we adopted the encoder architecture and hyperparameter settings from Tiny Transducer. We used a Deep Feedforward Sequential Memory Network (DFSMN) as the speech encoder, a lightweight and efficient structure, and a mainstream structure for KWS tasks. The DFSMN-based shared encoder consists of 6 layers, with hidden layers of 512 and projection layers of 320. The auxiliary encoder for CTC consists of two additional DFSMN modules and a linear projection layer. We used a stateless predictor with a context size of 2 and an embedding dimension of 320. The connector transforms the 320-dimensional encoder and decoder outputs into a 256-dimensional representation, then activates and projects it onto the final output. For CTC and RNN-T, the final output includes 70 monophthongs extracted from the CMU pronunciation dictionary "cmudict0.7b", or 200 Mandarin syllables extracted from a widely used lexicon, and a special whitespace marker φ. CTC or φ RNN-T For TDT, the output units also include those from 0 to the maximum duration D. max Duration options.
[0075] Training details are set to a local batch size of 64, with a maximum of 12,288 frames per batch. We use the AdamW optimizer with a maximum learning rate of 1e-3 and a warm-up of 10k steps. The learning rate is halved if the evaluation loss does not improve. Training terminates if the loss does not improve for more than three epochs.
[0076] Assessment details Baselines. We conducted a comprehensive baseline evaluation to assess the performance of our proposed system. First, we compared our proposed multi-head decoding framework with various ASR-based decoding strategies, including greedy search and beam search for CTC and RNN-T. Furthermore, we compared CTC-Transducer multi-head decoding with single-branch stream KWS decoding strategies and compared our system with end-to-end (E2E) KWS models. Additionally, we evaluated frame asynchronous MFA-KWS and synchronous MFS-KWS.
[0077] Evaluation Metrics. We report recall and false negative rates at different false alarm rates (FAR) or a fixed number of false alarms. Recall is defined as the ratio of true positives to the total number of positives, while the false negative rate is given by (1 - recall). The false alarm rate is calculated as the ratio of false alarms to the total duration of negative samples. ASR-based KWS decoding strategies consist of two stages: hypothesis generation and keyword matching. Therefore, they cannot inherently control false alarms, making it impractical to set fixed thresholds for recall and false alarm rates. For the LibriKWS-20 dataset, we present macro-average metrics for 20 keywords, including macro-precision and macro-recall.
[0078] Results and Analysis Impact of PSD Filtering Threshold This section will investigate the optimal threshold λ for PSD on the Hey Snips dataset. φ .like Figure 8 As shown, recall gradually decreases with increasing frame skipping rate, and the rate of decrease accelerates. When approximately 35% of frames are skipped, the performance degradation remains acceptable. However, further reducing λ... φ This would lead to a significant performance drop. To balance accuracy and computational efficiency, we will use λ... φ The setup was to filter out 35% of PSD frames in all subsequent experiments, and to search for λ on the corresponding development set. φ .
[0079] Figure 8 The performance comparison of different frame skipping rates on the SNIPS dataset is shown when FAR = 0.02.
[0080] Single-branch training vs. joint training In this section, we will evaluate the effectiveness of TransducerCTC joint training on the KWS task. For example... Figure 9As shown, (CTC ⊕ TDT) represents the joint training method defined in Equation (8), while other baselines rely solely on CTC or Transducer loss. Each block shows a single-branch CTC-based baseline and a Transducer-based baseline. The main difference in this table compared to our previous Transducer-based work is the inclusion of a CTC branch during training. Under all FAR conditions, MFA-KWS consistently outperforms RNN-T / TDT KWS. Notably, under strict FAR constraints (FAR = 0.02 / h or 0.05 / h), joint training reduces the relative false negative rate by 70.14% and 72.17% respectively (recall: 98.56% vs. 99.57%, 98.85% vs. 99.68%) compared to TDT training alone. When further combined with multi-head MFA decoding, the false negative rate of MFA-KWS was reduced by 72.22% and 82.61% respectively (recall rate: 98.56% vs. 99.60%, 98.85% vs. 99.80%).
[0081] Figure 9 This section shows a comparison of recall rates for systems trained on the SNIPS dataset at different false acceptance rate (FAR) levels: single-branch training (CTC or Transducer) versus joint training (CTC and Transducer). Joint training (CTC and Transducer). TRANS. refers to the Transducer. ⊕ indicates multi-task training. The best results within the same decoder head are shown in bold; the best results in the table are shown in bold with an underline.
[0082] Figure 10 This shows a comparison of recall rates for different fusion strategies in multi-head decoding across three test sets, with FA fixed at 2. TDT-4 represents D... max =4 TDT. The best average (AVG.) in each table block is shown in bold, while the global best result in the table is indicated in bold and underlined.
[0083] The results show that the joint training framework effectively integrates the advantages of CTC and Transducer, achieving better convergence and continuously improved performance compared to single-branch training. Furthermore, in addition to the advantages of joint training, the three-row decoding results for (CTC ⊕ TDT) demonstrate that multi-head decoding further improves performance. Therefore, in the following sections, we will focus on evaluating systems using the joint training paradigms defined by Equation (6) or Equation (8). In this section, we adopt the CDC-Last fusion strategy because it provides the best performance. Different fusion strategies will be analyzed in detail in subsequent sections.
[0084] Comparison of different fusion strategies We are Figure 10 This paper evaluates various fusion strategies for multi-head stream decoding and presents results on multiple English datasets. This analysis focuses particularly on RNN-T and TDT backbones.
[0085] For RNN-T without PSD, it operates in frame synchronization mode and is actually an MFS decoding system. Figure 10 All other systems in the dataset are MFA decoding systems. For the MFS system, CDC-Zero and CDC-Last yield identical results because there are no frame skips. Compared to naive score fusion strategies (CTC-Dom, Transducer-Dom, Equivalence-Dom), consistency-based fusion methods (CDC-Zero and CDC-Last) more effectively integrate the Transducer and CTC branches. CDC-based methods can improve positive scores while suppressing false activations caused by branch misjudgments. In CDC-based strategies, on multiple datasets, using the closest previous non-skipped score for padding (CDC-Last) consistently outperforms CDC-Zero and other fusion strategies. This suggests that preserving the latest decoding state is preferable to resetting with skipped frames. One possible explanation is that TDT and PSD employ asynchronous frame skipping mechanisms, making the most recent score crucial for maintaining the detection state of each branch and ensuring accurate merging at each time step. Therefore, all subsequent multi-head decoding experiments will use the CDC-Last fusion strategy.
[0086] Figure 11 This paper presents a comparison of recall performance of a joint RNN-T and TDT model at FA=2, as well as fsd / psd decoding on snips, test-clean, and test-other test sets. TRANS. refers to the converter. RNN-T represents the standard RNN-T model, while TDT-n represents D... max = n's TDT.
[0087] Figure 12 The recall of the proposed multi-head system is compared with that of end-to-end KWS systems at different levels on the HEY-SNPS dataset. TRANS. denotes the Transducer. POS.Snips denotes the positive portion of the Snips. EQU.LS denotes the equivalent number of librispeech segments.
[0088] Figure 13 The comparison of recall rates for the Mandarin dataset mobvoihotwords is shown. At FAR = 0.5 / h, the recall rates are compared.
[0089] Performance Comparison: MFS vs. MFA This section explores the optimal model configurations that maximize performance on three keyword datasets, and the results are summarized in... Figure 11 The Transducer branch offers two options: traditional RNN-T or TDT, while the CTC branch supports frame-by-frame FSD or word-by-word PSD decoding. For example... Figure 11 As shown in the first row, systems using the original RNN-T and FSD decoding in CTC belong to MFS-KWS, while configurations using TDT or PSD decoding belong to MFA-KWS. The results show that TDT consistently outperforms RNN-T, especially on more challenging datasets such as the LibriKWS-20 test set. Furthermore, optimizing the maximum duration D in TDT... max It can improve keyword detection, D max = 4 yielded the best results on all datasets. Compared to MFS-KWS, MFA-KWS (TDT-4 with PSD) reduced the relative false negative rate by an average of 23.0% (2.55% vs. 3.31%). In all subsequent experiments, we used TDT-4 (TDT, D max = 4) As the backbone of MFA-KWS.
[0090] Fixed keyword benchmark performance In this section, we will further demonstrate the performance comparison of the proposed MFA-KWS framework with various KWS baselines. Figure 12 and Figure 13 The results for Snips and MobvoiHotwords are listed separately.
[0091] For the fixed English keyword dataset Snips, in addition to Figure 9In addition to comparing ASR-based systems, we also compared them with end-to-end KWS systems. To bridge the gap between E2E KWS and ASR-based KWS systems, we first reimplemented forward Snips and the equivalent LibriSpeech (denoted as Pos.) for comparison. Figure 12 The results in rows 3 and 4 show that MDTC still achieves comparable performance under the new data strategy (99.88% vs. 98.85%; 99.92% vs. 99.29%). This indicates that the construction of the new dataset has little impact on performance, and may even make the dataset more challenging. We then compared MFA-KWS with E2E KWS. While E2E models such as RIL-KWS, WaveNet, and MDTC also achieve good performance under lenient FAR testing conditions, our proposed MFA-KWS can even correctly detect almost all keywords (99.96% vs. 99.88% or 99.92% under FAR=0.5 / h or FAR=1.0 / h). This demonstrates the powerful keyword detection capability of MFA-KWS. Furthermore, since all systems perform well under these FAR conditions, we finally evaluated the performance under extremely stringent conditions (FAR=0.02 / h or FAR=0.05 / h), which is more applicable to real-world scenarios. We found that the recall rate of MDTC in line 4, with FAR=0.05, was 89.52%, significantly lower than that of MFA-KWS under the same test conditions. This indicates that the E2E KWS system cannot handle difficult test scenarios well. Furthermore, compared with... Figure 9 CTC-based systems, WFST-based systems, and Transducer-based systems, as well as Figure 12 Compared to various systems such as E2E systems, MFA-KWS consistently maintains superior performance under almost all FAR conditions, setting a new state-of-the-art (SOTA) benchmark on the Snips dataset.
[0092] Figure 14 The display shows heatmaps of wakefulness scores for the Transducer and CTC branches at each (t,u), along with the joint decoding score of the MFA. The text is selected from test-clean, with the keyword "everything". The vertical yellow dashed lines represent word boundaries derived from force alignment.
[0093] Table 7 further evaluates the performance of MFA-KWS on the Mandarin benchmark MobvoiHotwords. While all systems achieved relatively high recall rates (FAR=0.5 / h) on short texts and classical Chinese texts, the MFA strategy consistently outperformed them, demonstrating its effectiveness. Compared to the best-performing models on these two datasets, MFA-KWS reduced the false negative rate by 27.66% and 68%, respectively (recall: 99.53% vs. 99.53% : 99.53% vs. 99.66%). (99.75% vs. 99.92%). These state-of-the-art results confirm that MFA-KWS can be well generalized to Mandarin, highlighting its robustness across different modeling units and languages.
[0094] Figure 15 The performance comparison of MFA-KWS with the ASR-based CTC and Transducer KWS baselines is shown. Macro accuracy for 20 keywords is provided on the test-clean and test-other subsets of librikws-20.
[0095] Comparison of arbitrary keyword detection In this section, we evaluate the performance of the arbitrary KWS dataset LibriKWS-20. To achieve arbitrary keyword detection, all models are trained on the general ASR dataset LibriSpeech. End-to-end KWS systems are unsuitable for this task because they require keyword-specific data for training. Therefore, we compare the proposed MFA-KWS with various ASR-based KWS systems employing different architectures (CTC and Transducer) and decoding algorithms, including greedy search, beam search, and streaming KWS decoding. Since selecting the optimal operating point for ASR decoding is challenging, we present the average accuracy at FAR=0 for a fair comparison. Figure 15 As shown, MFA-KWS achieved the highest accuracy on both test datasets and demonstrated the best average accuracy among all decoding strategies.
[0096] Figure 16 The performance of different decoding strategies under varying noise conditions is compared on the Snips, test-clean, and test-other datasets with a fixed FA=2. MFS is trained using a combination of CTC and RNN-T, while the others use CTC and TDT (D... max =4) Conduct training.
[0097] Overall, the proposed streaming MFA-KWS not only outperforms ASR-based decoding strategies but also allows for flexible adjustment of the KWS task's operation point. Furthermore, with FAR=0, MFA-KWS surpasses single-branch streaming decoding, demonstrating its ability to effectively integrate the advantages of both branches to improve performance. For FAR̸=0, we will further present the results for arbitrary keywords in the subsequent noise robustness analysis section. The best results further highlight the powerful capabilities of MFA-KWS in arbitrary keyword detection.
[0098] We are Figure 14 The figure shows an inference example of LibriKWS-20 to demonstrate MFA decoding in a continuous speech stream. The figure includes TDT-based streaming decoding, CTC-based streaming decoding, and the fusion detection score obtained by CDC-Last at each frame t, clearly showing the decoding state of each node in the decoding grid, as well as the start and end times of activation events.
[0099] Performance analysis under noisy conditions This section will further evaluate the noise robustness of the MFA-KWS. Figure 16 The performance of Snips and LibriKWS-20 under different signal-to-noise ratios (SNR) using different decoding strategies is reported. CTC-FSD and CTCPSD in the table indicate whether PSD is enabled in CTC-based decoding. Results confirm that the proposed MFA framework consistently outperforms all other methods across the three test sets, demonstrating strong robustness in high-noise environments. Several key observations include: 1) Transducer decoding exhibits higher stability and reliability on continuous speech datasets compared to CTC-based decoding, especially on more challenging test sets. 2) Under noisy conditions, frame skipping strategies consistently improve performance (MFA vs. MFS, CTC-PSD vs. CTC-FSD), likely by reducing noise interference and increasing focus on critical speech frames. Frame synchronization methods may help the model prioritize important speech content frames while filtering out irrelevant or noisy frames. 3) In noisy environments, multi-head decoding significantly outperforms single-branch methods, and MFA-KWS maintains stable performance across various SNR levels. These results highlight the strong noise robustness of the MFA-KWS and demonstrate its ability to effectively handle complex acoustic scenarios.
[0100] Reasoning efficiency This section evaluates the efficiency of the proposed MFAKWS by measuring the relative speeds of all datasets. Figure 17As shown, MFA-KWS achieves a speed improvement of 1.47× to 1.63× compared to the basic MFS-KWS. It is faster than single-branch RNN-T, comparable to TDT-4, and only faster than CTC-PSD. Given that MFA-KWS significantly outperforms all single-branch systems, these results confirm that the proposed framework not only maintains strong performance but also achieves a substantial improvement in efficiency.
[0101] Figure 17 The relative speed comparison of all decoding algorithms across all test sets is shown, with MFS decoding speed as the benchmark.
[0102] In this application, we propose MFA-KWS, a joint CTCTransducer KWS system with multi-head frame asynchronous decoding capability for dedicated keywords. MFA-KWS mitigates error accumulation in Transducer-based KWS systems and significantly improves the resolution between keywords and challenging negative samples. Furthermore, we explore several fusion strategies for multi-head frame asynchronous decoding, with CDCLast consistently maintaining robust performance under various test conditions. Compared to multiple baselines, MFA-KWS demonstrates superior performance under diverse datasets and noise conditions. Moreover, the MFA framework achieves a 47%–63% speedup over MFS-KWS. In conclusion, our results demonstrate that MFA-KWS is an efficient KWS framework well-suited for on-device deployment.
[0103] In other embodiments, the present invention also provides a non-volatile computer storage medium storing computer-executable instructions that can execute the keyword detection system training and evaluation method in any of the above method embodiments; Non-volatile computer-readable storage media may include a stored program area and a stored data area. The stored program area may store an operating system and an application program required for at least one function. The stored data area may store data created based on the keyword detection system training and evaluation method and the use of the system. Furthermore, the non-volatile computer-readable storage medium may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some embodiments, the non-volatile computer-readable storage medium may optionally include memory remotely configured relative to the processor, and these remote memories can be connected to the keyword detection system training and evaluation method via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0104] This invention also provides a computer program product, which includes a computer program stored on a non-volatile computer-readable storage medium. The computer program includes program instructions, which, when executed by a computer, cause the computer to perform any of the above-described keyword detection system training and evaluation methods.
[0105] Figure 18 This is a schematic diagram of the structure of the electronic device provided in the embodiment of the present invention, such as... Figure 18 As shown, the device includes one or more processors 1810 and a memory 1820. Figure 18 Taking a processor 1810 as an example, the device for the keyword detection system training and evaluation method and system may further include an input device 1830 and an output device 1840. The processor 1810, memory 1820, input device 1830, and output device 1840 can be connected via a bus or other means. Figure 18 Taking a bus connection as an example, the memory 1820 is the aforementioned non-volatile computer-readable storage medium. The processor 1810 executes various server functions and data processing by running non-volatile software programs, instructions, and modules stored in the memory 1820, thereby implementing the keyword detection system training and evaluation method described in the above embodiment. The input device 1830 can receive input numeric or character information and generate key signal inputs related to user settings and function control of the keyword detection system training and evaluation device. The output device 1840 may include a display screen or other display device.
[0106] The above-described product can execute the method provided in the embodiments of the present invention, and has the corresponding functional modules and beneficial effects for executing the method. Technical details not described in detail in this embodiment can be found in the method provided in the embodiments of the present invention.
[0107] In one embodiment, the above-described electronic device is used in a keyword detection system training and evaluation apparatus, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to: Collect fixed keyword datasets and arbitrary keyword datasets, as well as data augmentation in noisy environments; A multi-head frame asynchronous keyword detection model based on the CTC-Transducer joint training framework is constructed. During the training process, a multi-task learning strategy is used to combine the CTC loss and the Transducer loss to optimize the performance of the multi-head frame asynchronous keyword detection model. A multi-head frame asynchronous decoding strategy is adopted to fuse the decoding results of CTC and Transducer; The model is evaluated on both the fixed keyword dataset and the arbitrary keyword dataset, and the performance of different fusion strategies and decoding methods in noisy environments is analyzed to verify the effectiveness and robustness of the multi-head frame asynchronous keyword detection model in different scenarios.
[0108] The electronic devices described in this application exist in various forms, including but not limited to: (1) Mobile communication devices: These devices are characterized by their mobile communication capabilities and primarily aim to provide voice and data communication. These terminals include: smartphones, multimedia phones, feature phones, and low-end phones, etc.
[0109] (2) Ultra-mobile personal computer devices: These devices fall under the category of personal computers, possessing computing and processing capabilities, and generally also have mobile internet access features. These terminals include PDAs, MIDs, and UMPCs, etc.
[0110] (3) Portable entertainment devices: These devices can display and play multimedia content. This category includes: audio and video players, handheld game consoles, e-book readers, as well as smart toys and portable car navigation devices.
[0111] (4) Server: A device that provides computing services. The components of a server include a processor, hard disk, memory, system bus, etc. Servers are similar to general computer architectures, but because they need to provide highly reliable services, they have higher requirements in terms of processing power, stability, reliability, security, scalability, and manageability.
[0112] (5) Other electronic devices with data interaction functions.
[0113] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0114] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods of various embodiments or some parts of embodiments.
[0115] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for training and evaluating a keyword detection system, comprising: Collect fixed keyword datasets and arbitrary keyword datasets, as well as data augmentation in noisy environments; A multi-head frame asynchronous keyword detection model based on the CTC-Transducer joint training framework is constructed. During training, a multi-task learning strategy is used to combine the CTC loss and the Transducer loss to optimize the performance of the multi-head frame asynchronous keyword detection model. Specifically, the Transducer-based keyword detection model is extended to a CTC-Transducer joint training framework and multi-head frame synchronous decoding to achieve robust performance in challenging scenarios, including hard negative samples and noisy conditions. The multi-head frame synchronous decoding is transformed into a multi-head frame asynchronous decoding framework, which integrates phoneme synchronous decoding of the CTC branch and duration converter of the Transducer branch. A multi-head frame asynchronous decoding strategy is adopted to fuse the decoding results of CTC and Transducer. The CTC branch skips frames with high blank probabilities by setting a blank probability threshold, thereby accelerating the decoding process while maintaining performance. The Transducer branch uses a duration converter for decoding, which can skip irrelevant frames based on the predicted duration, further improving decoding efficiency. The model is evaluated on both the fixed keyword dataset and the arbitrary keyword dataset, and the performance of different fusion strategies and decoding methods in noisy environments is analyzed to verify the effectiveness and robustness of the multi-head frame asynchronous keyword detection model in different scenarios.
2. The method according to claim 1, wherein, The different fusion strategies include a naive score fusion strategy and a consistency-based fusion strategy. The naive score fusion strategy includes CTC-Dom, Transducer-Dom, and Equivalence-Dom. The consistency-based fusion strategy includes CDC-Zero and CDC-Last. The consistency-based fusion strategy more effectively integrates the Transducer and CTC branches.
3. The method according to claim 1, wherein, The improvements made by adopting a multi-header frame asynchronous decoding strategy include: Add bonus score S bonus This is done to sharpen the activation time boundary and improve the accuracy of keyword endpoints; Introducing T out To filter out excessively long decoding paths; The optimal path score is normalized by its length and converted to a confidence level in the range of [0,1]. The decoding process remains fully streaming with no forced delay. In practice, the decoding grid state remains unchanged, and each step only inputs a single frame logarithm or posterior to the decoding module.
4. The method according to any one of claims 1-3, wherein, The Transducer includes an encoder, a predictor, and a joint network, wherein the encoder converts speech signals into latent acoustic representations, the predictor acts as a language model to process text input and provides text information in an autoregressive manner, and the latent acoustic representations and the text information are then merged by the joint network to generate a probability distribution for the next label.
5. An electronic device comprising: At least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor to enable the at least one processor to perform the steps of the method according to any one of claims 1-4.
6. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1-4.
Citation Information
Patent Citations
A multi-example keyword detection method based on a multi-task neural network
CN108538285A
Keyword detection method, electronic equipment and storage medium
CN116072114A