Speech keyword detection method, storage medium and electronic device

By identifying and modeling different noise types, a finite state converter for noise perception is constructed, which solves the problems of overfitting and false alarm rate in speech keyword detection systems under high noise environments, and achieves high-accuracy detection in complex acoustic scenarios.

CN119339715BActive Publication Date: 2025-11-21AISPEECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411448198.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-16
Publication Date
2025-11-21
Estimated Expiration
2044-10-16

AI Technical Summary

Technical Problem

Existing speech keyword detection systems perform poorly in high-noise, low-signal-to-noise ratio environments, easily leading to overfitting and high false alarm rates, making it difficult to maintain high accuracy in complex acoustic scenarios.

Method used

By identifying different types of background noise, a noise-aware linear grammar finite-state converter is constructed, and wildcard arcs are introduced to configure a composite weighted finite-state converter, dynamically adjusting the system's sensitivity to noise and optimizing the decoding process.

Benefits of technology

Maintaining high keyword detection accuracy in high-noise, low-signal-to-noise ratio environments reduces overfitting and improves system performance in complex acoustic environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119339715B_ABST
    Figure CN119339715B_ABST
Patent Text Reader

Abstract

The application provides a speech keyword detection method, a storage medium and an electronic device, and relates to the technical field of speech signal processing. The method comprises the following steps: determining a phoneme score path and a background sound path corresponding to an acoustic feature sequence of a target speech to be detected; identifying noise types corresponding to each noise background sound; constructing an initial first linear grammar finite state transducer according to the phoneme score path, and configuring corresponding wildcard arcs for the first linear grammar finite state transducer by using each extracted noise type, so as to obtain a second linear grammar finite state transducer based on noise perception; creating a composite weighted finite state transducer through a composition operation; and decoding the composite weighted finite state transducer to predict whether a target keyword exists in the target speech. Thus, the keyword detection system can perceive different noise environments, dynamically adjust the sensitivity of the system to noise, and guarantee the keyword detection accuracy in a high-noise environment.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of speech signal processing, and in particular to a speech keyword detection method, a storage medium and an electronic device. BACKGROUND

[0002] Keyword spotting (KWS), also known as wake word detection (WWD), is an important interface for human-computer interaction. KWS systems are usually deployed on devices and run continuously. With the increasing diversification of application scenarios, improving noise robustness under resource-constrained conditions has become a key research focus in the KWS field.

[0003] To enhance the noise robustness of KWS systems, various methods have been developed. One method is to introduce a single / multi-channel speech enhancement (SE) module before the KWS module. Another method is to design a new training strategy or architecture to achieve a noise-robust KWS system. Regardless of the method, simulated noise data plays a crucial role in improving KWS performance. The preliminary experiments of the present application highlight two key observations: first, at low signal-to-noise ratio (SNR) levels, strong noise energy naturally leads to a decrease in recall; second, simulated data with too low SNR causes the model to overfit to noise, ultimately reducing overall performance. These phenomena occur because the noise in the challenging training data masks part of the keyword speech, causing the model to confuse noise with the target keyword.

[0004] Therefore, speech errors caused by noise can lead to mismatches between speech content and paired text transcriptions during training. In complex acoustic environments, especially in complex acoustic scenarios with high noise and low signal-to-noise ratio, the model overfits to noise during training, resulting in a low wake-up rate.

[0005] To address the above problems, the industry has not yet proposed a better solution. SUMMARY

[0006] The present application provides a speech keyword detection method, a storage medium and an electronic device to at least solve the problem of poor performance of speech keyword detection in a noisy environment in the related art.

[0007] In a first aspect, an embodiment of the present application provides a speech keyword detection method, comprising: determining a phoneme score path and a background path corresponding to an acoustic feature sequence of a target speech to be detected; the phoneme score path comprises a plurality of keyword phonemes in sequence and corresponding keyword activation scores, the keyword activation scores being used to indicate activation scores of keyword phonemes for corresponding keyword fragments in a target keyword; identifying at least one noise background sound in each background sound in the background path, and identifying a noise type corresponding to each noise background sound; constructing an initial first linear grammar finite state transducer according to the phoneme score path, and configuring corresponding wildcard arcs for the first linear grammar finite state transducer by using each extracted noise type, to obtain a second linear grammar finite state transducer based on noise perception; performing a composition operation on the second linear grammar finite state transducer, a vocabulary table and the phoneme score path, to create a composite weighted finite state transducer; the vocabulary table comprises a mapping relationship between normal phonemes of each keyword fragment and corresponding labels; and decoding the composite weighted finite state transducer to predict whether the target keyword exists in the target speech.

[0008] In a second aspect, an embodiment of the present application provides an electronic device, comprising: at least one processor, and a memory connected to the at least one processor in communication, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform steps of the speech keyword detection method of any embodiment of the present application.

[0009] In a third aspect, an embodiment of the present application provides a storage medium having a computer program stored thereon, wherein the program is executed by a processor to implement steps of the speech keyword detection method of any embodiment of the present application.

[0010] In a fourth aspect, an embodiment of the present application provides a computer program product, comprising computer programs / instructions, which are executed by a processor to implement steps of the speech keyword detection method of any embodiment of the present application.

[0011] The embodiment of the present application has the following beneficial effects:

[0012] By distinguishing different types of noise background, introducing noise type as a wildcard arc configuration in the grammar finite state transducer, constructing the search space through the composition operation, and decoding in the search space of the composite weighted finite state transducer, the noise information can be effectively integrated while the phoneme information is preserved, so that the keyword detection system can perceive different noise environments, thereby dynamically adjusting the sensitivity of the system to noise, and maintaining high keyword detection accuracy in high noise and low signal-to-noise ratio environments, and reducing the incidence of overfitting. BRIEF DESCRIPTION OF DRAWINGS

[0013] In order to more clearly illustrate the technical solutions of the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or prior art description will be briefly introduced. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0014] Figure 1 A flow chart of an example of a voice keyword detection method according to an embodiment of the present application is shown.

[0015] Figure 2 A comparative effect diagram of an example of noise data simulation and decoding of NTC based on CTC and WFST according to an embodiment of the present application is shown.

[0016] Figure 3 A structural schematic diagram of an embodiment of the electronic device of the present application. DETAILED DESCRIPTION

[0017] In order to make the purpose, technical scheme and advantages of the embodiments of the present application more clear, the technical scheme of the embodiments of the present application will be described clearly and completely in the following with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some embodiments of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0018] It should be noted that the presence of the preset keyword is continuously detected in the audio stream. The wake-up word detection (Keyword Spotting, KWS) system of various voice providers and the wake-up system on the mobile device side, such as Huawei "Xiao Yi", Baidu "Xiaoduo Xiaoduo", Xiaomi "Xiaoyi classmates" and Apple "Hey Siri". At present, the KWS system runs a low-power wake-up word detection model on the chip on the end side for 24 hours without interruption to respond to the user's instructions.

[0019] In recent years, there has been growing interest in designing small and efficient KWS systems based on Connectionist Temporal Classification (CTC). These systems are often deployed on low-resource computing platforms, and the limitations in model size and computational power create bottlenecks in complex acoustic scenarios. These limitations often lead to overfitting and confusion between keywords and background noise, resulting in high false positive rates.

[0020] Figure 1 An operation flowchart of an example of a voice keyword detection method according to an embodiment of the present application is shown.

[0021] As shown in step S110, the phoneme score path and the background path corresponding to the acoustic feature sequence of the target voice to be detected are determined. Figure 1

[0022] Here, the phoneme score path contains a plurality of keyword phonemes in sequence and corresponding keyword activation scores, which are used to indicate the activation scores of the keyword phonemes for the corresponding keyword segment in the target keyword.

[0023] In some embodiments, the target voice can be processed by employing various non-limiting acoustic models (such as deep neural networks, etc.) to extract the corresponding factor score path and background path. Illustratively, the acoustic feature sequence of the target voice to be detected is first extracted. Feature extraction is performed on the voice signal using methods such as Mel-frequency cepstral coefficients (MFCCs) or linear predictive coding (LPC). Then, according to these acoustic features, the phoneme score path is constructed. Each keyword phoneme is labeled as a node, and is connected to other keyword factors according to the timing, and each factor can also be represented by a keyword activation score indicating its relevance to the target keyword, which can be measured by a confidence score.

[0024] In step S120, at least one noise background sound in each background sound in the background path is identified, and the noise type corresponding to each noise background sound is identified.

[0025] ​Here, the noise of each frame of background sound can be identified by various noise detection algorithms, such as by short-time Fourier transform, spectral subtraction, or blind source separation algorithm, etc. Further, for each identified noise background sound, its noise type can also be identified by various machine learning methods. It should be understood that the noise type can be diversified, which should not be limited here. On the one hand, the noise type can be a description based on environmental information, such as traffic noise, human voice interference, etc. On the other hand, the noise type can also be described according to the interference or error caused by its recognition of the keyword, such as keyword insertion error caused by noise or phoneme masking error caused by too large noise, etc., and all of them are within the implementation scope of the embodiments of the present application.

[0026] In step S130, an initial first linear grammar FST (Finite State Transducer) is constructed according to the phoneme scoring path, and a corresponding wildcard arc is configured for the first linear grammar FST using each extracted noise type, to obtain a second linear grammar FST based on noise perception.

[0027] Here, according to the phoneme scoring path, an initial first linear grammar finite state transducer (FST) is constructed. At this time, according to the identified noise type, a wildcard arc is configured to form a second linear grammar finite state transducer, which realizes a noise modeling path based on noise type perception, so that the keyword detection system can adapt to various noise environments and improve the accuracy and sensitivity of keyword detection.

[0028] In step S140, a composite operation is performed on the second linear grammar FST, the vocabulary, and the phoneme scoring path to create a composite WFST (Weighted Finite State Transducer).

[0029] Here, the vocabulary contains the mapping relationship between the normal phonemes of each keyword fragment and the corresponding labels. The second linear grammar FST based on noise perception is combined with the vocabulary, and the mapping relationship between the phonemes and the labels in the vocabulary is used to set the weight for each phoneme, so as to give the corresponding probability weight in the decoding process. Further, using a weighting strategy, the importance of each phoneme and the influence degree (or transition cost) of the background noise are dynamically adjusted. In this way, the existence of different keyword fragments in the target speech can be more accurately evaluated, thereby improving the overall detection performance.

[0030] In step S150, the composite WFST search space is decoded to predict whether the target keyword exists in the target speech.

[0031] Here, a dynamic programming algorithm, a Viterbi algorithm or other efficient decoding algorithm can be applied to search the composite WFST, to select the path with the highest score by calculating the scores of different decoding paths, and to determine whether the target keyword is contained, so as to quickly and accurately determine the target keyword and optimize the real-time performance of the speech keyword detection.

[0032] In some examples of the embodiments of the present application, more abundant decoding path information can be supplemented to the composite WFST before decoding, to further improve the speech decoding accuracy. Specifically, based on an acoustic posterior model, an emission dense map ε(x) corresponding to the acoustic feature sequence x of the target speech is determined. Here, the acoustic posterior model is a pre-trained acoustic posterior model. Then, the emission dense map is intersected with the composite weighted finite state transducer to obtain an updated graph structure. Further, the Viterbi decoding algorithm is used to decode the updated graph structure to predict whether the target keyword exists in the target speech.

[0033] In some embodiments, the total score of each decoding path in the updated graph structure is calculated by the Viterbi decoding algorithm. Illustratively, the Viterbi algorithm is used to traverse the updated graph structure, starting from the initial state, to evaluate each possible path, and to update the score of the path at each step according to the current state, input feature and transition probability. Here, the path score is obtained by accumulating the product of the phoneme activation score and the transition probability, and the best predecessor state of each state needs to be recorded during the calculation process for subsequent backtracking. Then, the optimal decoding path is searched from the decoding paths according to the total scores of the decoding paths. Illustratively, the total scores of the decoding paths are compared, and the path with the highest score is selected as the optimal decoding path, which ensures that the best phoneme combination can be captured by accurately calculating the score of each path. Then, the recognition result for the target speech is determined according to the optimal decoding path to predict whether the target keyword exists in the target speech. Illustratively, according to the selected optimal decoding path, the initial state is backtracked to extract the corresponding phoneme sequence, which is matched with the keyword to output the final recognition result, so as to determine whether the target keyword exists in the target speech. More details about the decoding process of the algorithm combined with Viterbi will be described below in combination with other examples.

[0034] It should be noted that during the decoding of the updated graph structure, the correct phoneme alignment should be preferentially selected and the selection of the path or state of the wildcard should be reduced, so as to reduce the interference of noise on the keyword detection performance.

[0035] In some examples of the embodiments of the present application, the decoding paths of pure noise in the update graph structure are screened out, which are composed of wildcard symbols. Then, the total scores of the decoding paths are calculated by the Viterbi algorithm for each screened decoding path. By the embodiments of the present application, the noise decoding paths composed of only wildcard symbols without keyword phoneme symbols are screened out to avoid invalid analysis of pure noise background, so as to ensure that the decoding process is carried out for potential effective decoding paths, and the effectiveness and efficiency of the decoding process are ensured.

[0036] In some examples of the embodiments of the present application, the transition costs corresponding to different types of wildcard arcs are respectively configured to quantify the influence degree of different noise types on the keyword decoding process of the voice, so as to improve the adaptability and accuracy of the decoding system. The setting details of the transition costs (or influence weights) of different wildcard arcs can be referred to the example description in combination with other parts below. Specifically, the transition costs corresponding to different types of wildcard arcs are obtained. Then, for the noise modeling path in each decoding path, the total score corresponding to the noise modeling path is calculated according to the transition cost, and the noise modeling path is a decoding path containing wildcard arcs and keyword phoneme symbols, such as a state node in the path configured with a wildcard arc. In this way, in the decoding process, for each noise modeling path, the path score can be combined by the activation score of the keyword phoneme and the transition cost of the wildcard arc in combination with the wildcard arc and the keyword phoneme, so as to comprehensively consider the effectiveness of the keyword phoneme and the influence of noise on decoding, thereby improving the detection accuracy of the target keyword.

[0037] In some embodiments, the types of wildcard arcs include self-loop arcs and bypass arcs. The self-loop arc is used to indicate the noise type of noise insertion error, and the bypass arc is used to indicate the noise type of masking or interference caused by excessive noise. In this way, the potential errors between the keyword fragment sequence and the adjacent subwords are modeled by the self-loop arc (@) and the bypass arc (*), the self-loop arc @ is used to solve the noise insertion error, and the bypass arc * is used to process the masking and interference caused by excessive noise.

[0038] Regarding the implementation details of step S110, in some embodiments, the phoneme score path and the background sound path corresponding to the acoustic feature sequence of the target voice are determined based on a CTC (Connectionist Temporal Classification) acoustic model, and the encoder of the CTC acoustic model adopts a deep feedforward sequence memory network.

[0039] It should be noted that the CTC acoustic model is a deep learning model for sequence data processing, which allows the input and output sequence length to be mismatched. In speech keyword detection, the number of input audio frames is usually more than the number of characters of the target keyword, and this variable-length feature makes CTC very suitable for processing such tasks. In addition, CTC does not require manually labeled alignment information (i.e., the correspondence between input and output), simplifying the data preparation process, especially when the speech dataset is large, reducing the labeling cost. Moreover, CTC can automatically learn how to align the input and the target output during training. In the training phase, the model can be optimized by using the CTC loss function, so that it can accurately predict the timing of the keyword.

[0040] In addition, the encoder of the CTC acoustic model adopts DFSMN (Deep Feedforward Sequential Memory Networks) layers, which have good memory and modeling capabilities and can effectively capture long-time dependencies in speech signals. By introducing a memory module, historical information in the sequence data can be more effectively utilized, thereby improving the modeling accuracy and performance of the signal. In addition, in the deep encoding structure based on the DFSMN layer, each DFSMN layer can perform specific processing and feature extraction on the input data, and the combination of multiple layers can gradually improve the encoding performance.

[0041] As to the implementation details of step S130, in some embodiments, an initial first linear FST is constructed according to the phoneme score path. Exemplarily, the first linear FST contains multiple states and transitions between the states, the states can be defined by the respective phoneme score path nodes in the phoneme score path, and the transitions (or edges) can be defined by the order of the respective phoneme score path nodes. Then, for each extracted noise type, the matching target state is determined according to the phoneme score node corresponding to the noise type, and the target wildcard arc matching the extracted noise type is automatically determined according to the preset noise wildcard relationship table, and the target wildcard is configured to the target state. Exemplarily, the noise wildcard relationship table records multiple noise types and corresponding wildcard arcs, for example, noise type 1 corresponds to wildcard arc a, noise type 2 corresponds to wildcard arc b, noise type 3 corresponds to wildcard arc c, and so on. Specifically, the corresponding keyword phonemes and phoneme score nodes are located according to the noise frame, and then the wildcard arc corresponding to the noise type of the noise frame is configured to the corresponding state, realizing the association matching between the wildcard arc and the state information based on noise perception, so that different types of noise can be individually modeled during the decoding process, the influence of different noise environments can be perceived, and thus the sensitivity of the system to noise can be dynamically adjusted, improving the detection performance of speech keyword detection in a noisy environment.

[0042] Through the embodiments of the present application, a noise-aware CTC-based KWS (NTC-KWS) framework is provided, aiming to enhance the robustness of the model in noisy environments, especially in the case of extremely low signal-to-noise ratio. In the embodiments of the present application, two additional noise modeling wildcard arcs are introduced in the training and decoding process based on the weighted finite state transducer (WFST) graph: a self-loop arc for solving noise insertion errors, and a bypass arc for processing masking and interference caused by excessive noise. Experiments on clean and noisy Hey Snips datasets show that NTC-KWS outperforms the state-of-the-art (SOTA) end-to-end system and CTC-KWS baseline under various acoustic conditions, especially in low signal-to-noise ratio scenarios.

[0043] Specifically, first, the possible error categories in the training data of low signal-to-noise ratio are classified, and then different noises are explicitly modeled in training and testing. Specifically, two additional noise modeling wildcard arcs are introduced in the training based on the weighted finite state transducer (WFST) graph. A self-loop arc is used to solve noise insertion errors, and a bypass arc is used to process masking and interference caused by excessive noise. The original CTC loss function is modified so that the model can model noise.

[0044] Further, in the testing or inference stage, the two noise arcs are also introduced. On the one hand, it can reduce the decoding difficulty of the model under low signal-to-noise ratio. If a certain part of the wake-up word is not easily recognized, it can be skipped by the noise modeling edge. On the other hand, it ensures the consistency of the training and testing of NTC-KWS, so that the overall performance of the system will be better.

[0045] Through the NTC-KWS system provided by the embodiments of the present application, the possibility of jointly modeling noise in the training process of the wake-up system is illustrated, which promotes further mining of more fine-grained noise modeling. For example, how to model human voice interference, how to model pure environmental noise, and whether the wake-up system can perceive different noise environments to dynamically adjust the sensitivity of the system to noise.

[0046] The specific framework and details of the NTC-KWS system provided by the embodiments of the present application will be described below.

[0047] Figure 2 A comparison effect diagram of an example of noise data simulation and NTC decoding graph based on CTC and WFST according to the embodiments of the present application is shown.

[0048] In Figure 2The left part of Figure 1 illustrates three types of errors, masking and interference (in red) and insertion (in blue) due to excessive noise, through a noise simulation example. In the right part of Figure 1, the keyword is set to "A B C" in conjunction with the example. The decoding transfer rules at the grammar and label levels are denoted by Figure 2 and S, respectively. In the NTC, the two additional wildcard arcs are highlighted in red and blue, corresponding to the error colors in the left simulation example. In the composite graph S, λ1 and λ2 denote the wildcard transition costs. ε on the left denotes the null transition, while on the right denotes the null output.

[0049] A. WFST-based CTC implementation

[0050] Connectionist temporal classification (CTC) is a widely used loss function for training sequence learning models where the input and label lengths do not match, especially for automatic speech recognition (ASR) or keyword detection (KWS). For ASR, the input sequence of acoustic features can be represented as where each x t is a D-dimensional acoustic feature vector, and T denotes the total number of frames. The output label sequence can be represented as where y U is a label token, and U is the total number of tokens. It is worth noting that T must be much larger than U, because CTC introduces a special blank token φ to model the null output for each frame. The actual alignment sequence for the input is where each π t denotes a normal phoneme or the blank token φ in the CTC vocabulary V ctc . The goal of CTC is to minimize the sum of the negative log-likelihoods of the alignment sequence given the input sequence. Therefore, the CTC loss can be formulated as:

[0051]

[0052] where B is defined as the mapping of a legal CTC alignment π to the label y, and B -1 denotes the inverse mapping.

[0053] Moreover, the decoding is to find the most likely target for the speech, which can be formulated as:

[0054]

[0055] where the approximate derivation introduces the Viterbi decoding to find the most likely alignment.

[0056] ​Weighted Finite State Transducers (WFSTs) aim at mapping input sequences to output sequences and provide an efficient implementation for the training and decoding of CTCs. Specifically, the search space S is defined by the composition of and represent the CTC merge (removing φ and merging duplicate tokens), the vocabulary (mapping from words to tokens) and the grammar FST, respectively. In addition, an emission dense graph ε is constructed to model the acoustic posterior. Thus, the training and decoding of CTCs can be reformulated as:

[0057]

[0058]

[0059] where ∩ and ° represent the intersection and composition operation on WFST graphs, respectively. is represented as a linear FST constructed from the transcription y. Forward and Viterbi represent the forward computation and Viterbi decoding on the composed WFST graph, respectively. Specifically, Token Passing is an efficient method to implement Viterbi decoding within the WFST framework. When a token reaches the final node, the decoding path is considered complete.

[0060] B. CTC-based keyword spotting

[0061] Compared to the CTC-based automatic speech recognition (ASR) described in Section A above, the training process of KWS is almost identical and the decoding stage has only two small differences.

[0062] In terms of the decoding graph, the grammar FST of KWS is much more compact and lightweight, containing only the keyword fragment paths and an optional background path, instead of the entire acoustic space.

[0063] In terms of the confidence score metric, the search results are converted into keyword activation scores by computing the confidence scores of the speech segments and different strategies can be evaluated to compute the confidence scores for CTC-based KWS.

[0064] C. Training and decoding of NTC-KWS

[0065] Although CTC-based KWS performs well under typical acoustic conditions, its performance significantly degrades when speech segments are drowned in noise. As Figure 2errors, including content masking, noise artifacts, and pronunciation distortions. This leads to a mismatch between the speech and the transcription, increasing the likelihood of noise-induced false triggers. To enhance noise robustness, the present application proposes NTC-KWS, which aims to introduce two types of wildcard arcs in both the training and decoding processes: self-loop arcs and bypass arcs. These arcs represent different label-level graph errors, as shown in Figure 1 When a portion of the keyword is masked by noise, the self-loop arc (@) models insertion errors in the keyword label (indicated in blue error and arc). In scenarios where a phoneme is distorted or masked by noise, the bypass arc (*) considers substitution or deletion errors (indicated in red error and arc). These wildcards effectively capture the notion of “excess speech noise and interference”. The acoustic posterior of the * and @ arcs, representing the average posterior of all phonemes, is computed as follows:

[0066]

[0067] where π i (j) denotes the label j in the vocabulary V i predicted in frame x ctc . To ensure that the model prefers to learn the correct alignment over the wildcard paths, the wildcard arcs are assigned a large initial penalty. These penalties are exponentially decayed over training epochs:

[0068]

[0069] where l denotes the current epoch number, β @ , β * is a hyperparameter that can be set empirically.

[0070] While the model learns to handle the noise-dominant portion of the keyword through NTC-based training, the standard CTC-KWS decoding graph fails to fully exploit the enhanced training, especially in low signal-to-noise ratio (SNR) scenarios. To better align the training and decoding processes, the decoding search space can also be extended by modifying the grammar graph to introduce wildcard paths, as shown on the right side of Figure 1 . For example, assume the keyword is “A B C”, and the vocabulary V ctc includes the label “AB C D E” as well as a blank label φ, introducing self-loop arcs (@) and bypass arcs (*) on the graph to model potential errors between the “A B C” keyword sequence and adjacent subwords. Then, the NTC grammar graph G ntc transfers the keyword component paths and the vocabulary Combining to create the NTC-KWS search space S ntc With CTC-KWS search space S ctc In comparison, S ntc An alternative noise modeling path is provided for token propagation, thereby improving the decoding topology of noisy KWS in several ways. First, consistency between the training and decoding methods is guaranteed. Second, identifying each token in a keyword becomes challenging when noise dominates. Therefore, relaxing decoding constraints in complex acoustic environments is reasonable to improve recall, as this allows the model to select wildcard arcs against noise. To balance recall and false positives, a transfer cost λ is introduced for each type of arc in the NTC decoding graph. @ and λ * ( Figure 1 (λ1 and λ2 in the code). However, if the decoding path consists only of wildcard markers and does not contain any keyword phonemes, the path will be discarded.

[0071] This application proposes an NTC-KWS framework to address speech errors caused by noise. It combines noise-aware training and decoding strategies to enhance the model's robustness in noisy environments. More specifically, this application incorporates two noise modeling wildcard arcs into the CTC training framework to adaptively model noise and further improves the WFST-based decoding search space, thereby enhancing KWS performance. Thus, the noise-aware CTC training criterion can model noise-dominated speech segments in complex acoustic scenarios, preventing noise overfitting. Furthermore, by employing a training-free decoding strategy, the noise modeling path is integrated into the WFST-based KWS decoding, improving performance in noisy environments without requiring NTC training and providing further improvements when applied within the NTC-KWS framework. Compared to state-of-the-art end-to-end systems and the CTC-KWS baseline, the proposed NTC-KWS achieves superior performance under both clean and noisy conditions, particularly excelling in highly challenging scenarios.

[0072] To demonstrate the advanced nature of the technical solutions in the embodiments of this application, the inventors of this application also conducted experiments on the superior performance of the NTC-KWS system during the practice of this application. The relevant experimental details will be described in detail below.

[0073] A. Dataset

[0074] To evaluate the system's performance and noise robustness, this application uses the following open-source corpus to construct the dataset.

[0075] LibriSpeech (LS): LibriSpeech is an open-source, widely used ASR corpus containing 960 hours of English speech and its corresponding transcribed text. This dataset is used for pre-training of the base model and for general ASR data for the KWS model.

[0076] Hey Snips (Snips): Hey-Snips is an open-source KWS dataset centered around the keyword "Hey Snips." Since negative utterances do not have transcribed text, these negative data are only used for evaluation, not for training. This application combines all negative data in the training, development, and test sets to create a large test set. The length of the new negative test set is approximately 97 hours.

[0077] WHAM!: The WHAM! dataset contains various types of environmental noise, such as music, restaurant sounds, etc. This application uses this dataset to synthesize noisy speech data for training and evaluation.

[0078] During the training process, the base speech encoder is first pre-trained using the clean LibriSpeech dataset. Then, the KWS model is fine-tuned using the Snips and an equal amount of LibriSpeech dataset. To simulate the noisy dataset, each clean speech waveform of Snips and LS is mixed with a randomly selected noise sample from the WHAM! dataset, with a signal-to-noise ratio (SNR) selected from a uniform distribution from 0 to 20 dB. The clean and noisy speech data are collected to build the training dataset.

[0079] During the evaluation process, the noisy negative test set is simulated in the same way as the training, and the positive speech of the test is simulated at different fixed SNR levels. All positive test samples in the Snips dataset are mixed with noise at different SNR levels ({0, 5, 10, 15, 20} dB). To further evaluate the robustness of the model, an additional dataset is also simulated at -5 dB, which represents an extremely low and out-of-domain SNR.

[0080] B. Configuration

[0081] In terms of models, the toolkit K2 is used to build the differential WFSTs for training, and OpenFST is utilized to build the decoding system. Deep Feedforward Sequence Memory Network (DFSMN) is a lightweight and efficient speech encoder, which is widely used in KWS tasks. This architecture is used as the backbone of CTC and NTC KWS systems. All models are initialized from the same pre-trained checkpoint and optimized by SGD. The DFSMN architecture contains 6 layers, with the hidden layer and projection layer size set to 512 and 320, respectively. The left and right context orders for each DFSMN layer are 8 and 2, respectively. The final output units of the CTC-based model include 70 phonemes derived from the CMU pronunciation dictionary “cmudict-0.7b”, and a special blank symbol φ. In addition, the NTC model includes two additional special symbols: a self-loop symbol @ and a bypass symbol *.

[0082] In terms of acoustic features, 40-dimensional log-Mel filter bank coefficients (FBank) are extracted from the raw audio using a 25ms window and a 10ms hop. In addition, two data augmentation strategies are adopted in the present application: online speech perturbation, with the warping factors uniformly sampled from {0.9, 1.0, 1.1}, and SpecAugment, with 2 frequency masks (max F = 10) and 2 time masks (max T = 50) used for each speech segment. The current frame is concatenated with the previous 5 frames and the next 5 frames to generate the model input. Here, downsampling with a frame skipping factor of 3 is also applied to reduce the computational cost.

[0083] In terms of hyperparameters, the decay coefficient β @ and β * are set to 0.999 and 0.975, respectively. The initial wildcard costs λ(0) @ and λ(0) * are both set to -4 for training. The decoding costs λ @ and λ * are adjusted during testing. During the label passing inference process of the CTC and NTC systems, there are at most 20 active labels.

[0084] C. Evaluation details

[0085] In terms of baselines, to evaluate the NTC-KWS system provided by the embodiments of the present application, a comprehensive baseline is constructed. First, the CTC-based model is compared with the state-of-the-art (SOTA) end-to-end (E2E) system on the Hey Snips dataset, which is open-sourced in WeKws, to demonstrate that the KWS model built by CTC or its variants can achieve excellent performance. Then, further detailed comparisons of NTC-KWS and CTC-related baselines are made under different in-domain and out-of-domain SNR levels to evaluate the improved training and decoding methods.

[0086] In terms of evaluation metrics, recall values at different SNR levels are reported at a specific false alarm rate (FAR). The published SOTA models usually present recall at a FAR level of 0.5 or 1 per hour. To ensure a fair comparison, the present application adopts the same setting when benchmarking against these systems. However, this evaluation criterion is relatively lenient as performance tends to be stronger at this FAR level. To minimize the impact of bias and randomness, the present application evaluates all systems under more stringent conditions. Therefore, after a preliminary comparison, recall at a FAR = 0.05 per hour is demonstrated.

[0087] Table 1. Comparison of E2E baseline and the present application’s CTC and NTC KWS systems on clean Hey Snips dataset. “POS. SNIPS” refers to the positive portion in the Snips dataset, while “EQU. LS” indicates that the present application added an equal amount of LibriSpeech speech as general-purpose ASR data.

[0088]

[0089] The experimental results and related analysis conclusions will be specifically expanded below.

[0090] A. Comparison with end-to-end (E2E) SOTA baseline

[0091] The upper half of the table shows the results using the official Snips data, while the lower half reports the results based on the reconstructed data due to the lack of transcription text for the negative samples in Snips. As described in the above, the CTC model cannot be trained on speech data without transcription text, so this reconstruction is necessary. The CTC and NTC models are compared with the SOTA E2E baseline on the clean Hey Snips dataset, with a false alarm rate (FAR) of 0.5 or 1.0 per hour as the test condition. The models of the present application can achieve comparable performance to the SOTA model (D & E vs. B or C1). When the models are trained on the same data, the results in the lower half of Table I show that the CTC and NTC models outperform the SOTA baseline at a FAR of 0.5 or 1.0 per hour (D & E vs. C2). Furthermore, under the more stringent test condition of FAR = 0.05, the recall of the CTC and NTC models improves by more than 9%. The results in Table I further demonstrate that the models of the present application outperform the baseline, especially under stringent test conditions, verifying the effectiveness of the CTC-KWS baseline and the proposed NTC-KWS.

[0092] B. Impact of NTC training and decoding

[0093] Here, the respective effects of NTC training and decoding are investigated. Table II shows the performance of CTC and NTC models under different combinations of training and inference strategies. For CTC inference (F vs. D), updating the training criterion to NTC did not improve performance because NTC training does not force the model to treat noise as part of the speech signal. As a result, the performance of model F is similar to that of a CTC model trained using clean high-quality speech. Training the model using the CTC loss and testing it using the NTC decoding method (D vs. G) yields an absolute improvement in recall of 2.7%. This gain is particularly significant at low SNR levels (SNR = -5, 0), highlighting the challenges of label decoding under such conditions. The proposed decoding strategy relaxes the constraints in complex environments, resulting in a significant performance improvement. This demonstrates that if the original CTC KWS model cannot be retrained, but performance improvement is needed in noisy environments, using the NTC decoding strategy is a practical solution. Furthermore, applying NTC decoding to the NTC model results in an absolute increase in average recall of 2.2% (E vs. G), mainly due to the improvement brought by the consistency between the training and inference processes. This improvement highlights the benefits of NTC training and decoding consistency. In summary, the proposed NTC-KWS framework achieves an absolute performance improvement of 4.9% over the CTC-KWS baseline (E vs. D). Notably, at low SNR levels of -5 and 0, significant performance improvements of 15% and 8.3% are achieved, respectively, highlighting the robustness of NTC-KWS in extreme noisy conditions.

[0094] Table II. Evaluation of different combinations of CTC-KWS and NTC-KWS. "TRAIN" refers to the specific training criterion used during the model training phase, while "TEST" indicates the type of decoding graph used during the inference phase.

[0095]

[0096] C. Finding the optimal decoding configuration

[0097] Here, the importance of two hyperparameters in the decoding configuration will be further explored. A grid search was performed on the values of λ @ and λ * . When λ @ and λ *When both are set to positive values, the performance is best, which indicates that the NTC model does not overfit the shortcuts during the NTC training process and still needs to go through the wildcard path to enhance the performance during inference. In addition, the applicant also tried to set the decoding transition coefficients to negative values during the implementation of the present application. Although the improvement is not as obvious as when the values are positive, the performance is still better than the CTC-KWS baseline. When both coefficients are set to -∞, the NTC-KWS model behaves like model F in Table II, which indicates that the NTC decoding strategy degenerates to the CTC method. Another conclusion drawn from Table III is that the sensitivity of the decoding stage to λ * is higher than λ @ , and it is not feasible to set λ * too high.

[0098] Table III. Different configurations of the self-loop weight λ @ and the bypass weight λ * in the inference stage. Negative values represent penalties, and positive values represent encouragements.

[0099]

[0100] The embodiments of the present application provide NTC-KWS, which is a CTC-based KWS framework variant optimized for noisy environments, especially in very low signal-to-noise ratio (SNR) conditions. By introducing multiple wildcards in the training and decoding stages, the risk of overfitting the model to noise is reduced. Modeling noise-dominant segments as wildcard tokens enables the model to better distinguish keyword speech from background noise. Compared with previous end-to-end (E2E) systems, the system provided by the embodiments of the present application achieves SOTA*level performance on the clean Hey Snips dataset. Specifically, the NTC-KWS framework has an average absolute recall rate improvement of 4.9% compared with the standard CTC-KWS system at all SNR levels. In low SNR conditions, such as -5dB and 0dB, the absolute gain in recall rate reaches 15% and 8.3%, respectively, which highlights the robustness of the system provided by the embodiments of the present application in challenging, noisy environments.

[0101] It should be noted that, for the foregoing method embodiments, in order to simply describe, they are all described as a series of actions, but those skilled in the art should know that the present application is not limited to the order of the actions described, because according to the present application, certain steps can be performed in other orders or at the same time. Secondly, those skilled in the art should know that the embodiments described in the specification all belong to preferred embodiments, and the actions and modules involved are not necessarily essential to the present application. In the above embodiments, the description of each embodiment is focused on, and the parts not described in detail in a certain embodiment can be referred to the related description of other embodiments.

[0102] In some embodiments, the embodiments of the present application provide a non-volatile computer readable storage medium, wherein one or more programs including execution instructions are stored in the storage medium, and the execution instructions can be read and executed by an electronic device (including but not limited to a computer, a server, or a network device, etc.) to execute any of the voice keyword detection methods described above.

[0103] In some embodiments, the embodiments of the present application also provide a computer program product, which includes a computer program stored on a non-volatile computer readable storage medium, and the computer program includes program instructions, which, when executed by a computer, cause the computer to execute any of the voice keyword detection methods described above.

[0104] In some embodiments, the embodiments of the present application also provide an electronic device, which includes at least one processor and a memory connected to the at least one processor in communication, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute a voice keyword detection method.

[0105] Figure 3 is a hardware structure schematic diagram of an electronic device for executing a voice keyword detection method provided by another embodiment of the present application, as shown in Figure 3 The device includes:

[0106] one or more processors 310 and a memory 320, Figure 3 In an example, the processor 310 is taken as an example.

[0107] The device for executing the voice keyword detection method can also include an input device 330 and an output device 340.

[0108] The processor 310, the memory 320, the input device 330 and the output device 340 can be connected through a bus or other means, Figure 3 In an example, the connection through the bus is taken as an example.

[0109] The memory 320 is a kind of non-volatile computer readable storage medium, which can be used to store non-volatile software programs, non-volatile computer executable programs and modules, such as program instructions / modules corresponding to the voice keyword detection method in the embodiments of the present application. The processor 310 executes various functional applications and data processing of the server by running the non-volatile software programs, instructions and modules stored in the memory 320, that is, the voice keyword detection method of the above method embodiment is realized.

[0110] The memory 320 can include a program storage area and a data storage area. The program storage area can store an operating system, application programs required by at least one function, and the like. The data storage area can store data created according to the use of the electronic device, and the like. In addition, the memory 320 can include a high-speed random access memory, and can also include a non-volatile memory such as at least one of a magnetic disk storage device, a flash memory device, or other non-volatile solid state storage device. In some embodiments, the memory 320 can optionally include a memory disposed remotely from the processor 310, which can be connected to the electronic device through a network. Examples of the network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.

[0111] The input device 330 can receive input digital or character information, and generate a signal associated with user settings and function control of the electronic device. The output device 340 can include a display device such as a display screen.

[0112] The one or more modules are stored in the memory 320, and when executed by the one or more processors 310, perform the voice keyword detection method in any of the method embodiments described above.

[0113] The above products can perform the methods provided in the embodiments of the present application, and have the corresponding function modules and beneficial effects of performing the methods. Technical details not described in detail in the embodiments can be referred to the methods provided in the embodiments of the present application.

[0114] The electronic device of the embodiments of the present application exists in various forms, including but not limited to:

[0115] (1) Mobile communication device: This type of device is characterized by having mobile communication function, and providing voice and data communication as the main target. This type of terminal includes: smart phone, multimedia phone, functional phone, and low-end phone, etc.

[0116] (2) Ultra-mobile personal computer device: This type of device belongs to the category of personal computers, has computing and processing functions, and generally also has mobile Internet features. This type of terminal includes: PDA, MID and UMPC device, etc.

[0117] (3) Portable entertainment device: This type of device can display and play multimedia content. This type of device includes: audio and video players, handheld game consoles, electronic books, smart toys, and portable car navigation devices.

[0118] (4) Other onboard electronic devices with data interaction function, such as vehicle-mounted devices installed on vehicles.

[0119] The device embodiments described above are merely illustrative, wherein the units described as separate components can or can not be physically separate, and the components displayed as units can or can not be physical units, i.e., can be located in one place, or can be distributed to multiple network units. Part or all of the modules can be selected according to actual needs to achieve the purpose of the embodiment scheme.

[0120] Through the description of the above embodiments, those skilled in the art can clearly understand that the embodiments can be implemented by means of software plus a general hardware platform, and of course can also be implemented by hardware. Based on such understanding, the above technical solutions can be embodied in the form of a software product, which can be stored in a computer readable storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, etc., and includes a plurality of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute the methods described in each embodiment or some parts of the embodiments.

[0121] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A voice keyword detection method, comprising: determining a phoneme score path and a background path corresponding to an acoustic feature sequence of a target voice to be detected; the phoneme score path comprising a plurality of sequential keyword phonemes and corresponding keyword activation scores, the keyword activation scores being used to indicate activation scores of keyword phonemes for corresponding keyword segments in a target keyword; identifying at least one noise background sound in each background sound in the background path, and identifying a noise type corresponding to each noise background sound; constructing an initial first linear grammar finite state transducer according to the phoneme score path, and configuring a corresponding wildcard arc for the first linear grammar finite state transducer using each extracted noise type, to obtain a second linear grammar finite state transducer based on noise perception; performing a composition operation on the second linear grammar finite state transducer, a vocabulary, and the phoneme score path, to create a composite weighted finite state transducer; the vocabulary comprising a mapping relationship between normal phonemes of each keyword segment and corresponding labels; decoding the composite weighted finite state transducer to predict whether the target keyword exists in the target voice.

2. The method of claim 1, wherein, The decoding of the composite weighted finite state transducer to predict whether the target keyword exists in the target voice comprises: determining an emission dense graph corresponding to the acoustic feature sequence of the target voice based on an acoustic posterior model; performing an intersection operation on the emission dense graph and the composite weighted finite state transducer to obtain an updated graph structure; decoding the updated graph structure by a Viterbi decoding algorithm to predict whether the target keyword exists in the target voice.

3. The method of claim 1, wherein, The construction of the initial first linear grammar finite state transducer according to the phoneme score path, and the configuration of the corresponding wildcard arc for the first linear grammar finite state transducer using each extracted noise type, to obtain the second linear grammar finite state transducer based on noise perception, comprises: constructing an initial first linear grammar finite state transducer according to the phoneme score path; the states of the first linear grammar finite state transducer being defined by each phoneme score path node in the phoneme score path; for each extracted noise type, determining a target state matched with the noise type according to a phoneme score node corresponding to the noise type, and automatically determining a target wildcard arc matched with the extracted noise type according to a preset noise wildcard relationship table, and configuring the target wildcard to the target state; the noise wildcard relationship table recording a plurality of noise types and corresponding wildcard arcs.

4. The method of claim 2, wherein, The decoding of the updated graph structure by the Viterbi decoding algorithm to predict whether the target keyword exists in the target voice comprises: calculating total scores of each decoding path in the updated graph structure by the Viterbi decoding algorithm; searching an optimal decoding path from each decoding path according to the total scores of the decoding paths. Determine a recognition result of the target speech according to the optimal decoding path, to predict whether the target keyword exists in the target speech.

5. The method of claim 4, wherein, The total score of each decoding path in the updated graph structure is calculated by the Viterbi decoding algorithm, including: Purify pure noise decoding paths in the updated graph structure, which are composed of wildcard markers; For each purified decoding path, the total score of the decoding path is calculated by the Viterbi decoding algorithm.

6. The method of claim 5, wherein, The total score of the decoding path is calculated by the Viterbi decoding algorithm, including: Obtain transition costs respectively corresponding to various types of wildcard arcs; For each noise modeling path in the decoding path, calculate the total score corresponding to the noise modeling path according to the transition costs; the noise modeling path is a decoding path containing wildcard arcs and keyword phoneme markers.

7. The method of claim 6, wherein, The types of the wildcard arcs include self-loop arcs and bypass arcs; the self-loop arcs are used to indicate noise types of noise insertion errors, and the bypass arcs are used to indicate noise types of masking or interference caused by excessive noise.

8. The method of any one of claims 1-7, wherein, The phoneme score path and the background sound path corresponding to the acoustic feature sequence of the target speech to be detected are determined, including: Determine the phoneme score path and the background sound path corresponding to the acoustic feature sequence of the target speech based on a CTC acoustic model; an encoder of the CTC acoustic model adopts a deep feedforward sequence memory network. 9.A storage medium having stored thereon a computer program, which, when executed by a processor, implements the steps of the method of any one of claims 1-8.

10. An electronic device comprising: At least one processor, and a memory connected with the at least one processor in communication, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the steps of the method of any one of claims 1-8.