End-to-end voice keyword detection method, device and medium
By dynamically adding auxiliary categories, the problem of poor training results of high false alarm rates and low-quality samples in the speech keyword detection model is solved, achieving higher detection accuracy and performance improvement.
Patent Information
- Application Number
- CN202510210659.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-02-25
AI Technical Summary
In the prior art, when the speech keyword detection model deals with non-keyword categories, it is difficult to determine the decision boundary, resulting in a high false alarm rate; at the same time, the training effect of low-quality keyword samples is poor, affecting the model performance.
The method of dynamically adding auxiliary categories is used, and it is divided into the first and second categories of auxiliary categories, which are used to deal with non-keyword samples that are prone to false alarms and low-quality keyword samples that are prone to missed detection. Improve the classification capabilities of the model by dynamically adding auxiliary categories during the training process, adjusting labels and training strategies.
It effectively reduces the false alarm rate, improves the accuracy of keyword detection, utilizes low-quality samples, and improves the overall performance of the model.
Smart Images

Figure CN119993129A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of speech signal processing, and in particular to an end-to-end speech keyword detection method, device and medium. Background Art
[0002] Voice keyword detection technology refers to detecting pre-defined keywords from a continuous audio stream. It is one of the current research hotspots in the voice field and is widely used in terminal device control, smart home, voice information retrieval, audio monitoring and other fields.
[0003] However, for speech keyword detection with a fixed vocabulary, both traditional methods and deep learning methods generally use keyword-filler models to model keywords and non-keywords; specifically, the keywords in the keyword table are split according to their modeling units and used as keyword categories, and the sentences other than keywords are split according to modeling units and all used as non-keyword categories. This modeling method results in a smaller number of overall classification categories, less data required, and faster decoding speed, but it also has some defects. Since the non-keyword category contains too much content, some non-keyword features are very similar to keyword features, making it difficult for the model to determine a better decision boundary during training, thereby damaging the performance of the keyword detection model.
[0004] Taking tonal syllables as modeling units as an example, on the one hand, the keyword-filler model will cause some non-keyword syllables that are very similar to the acoustic pronunciation of the syllables belonging to the keyword category to be included in the non-keyword category, which will increase the classification difficulty of the model during the training and reasoning processes, resulting in a higher false alarm rate for the keyword detection model.
[0005] On the other hand, in the training corpus, there are some training sentences containing keywords, and the sample quality itself is not high, which makes the pronunciation of the keyword fragment unclear, resulting in a high loss value of the sample during training, that is, the keyword sample is always missed, which is not conducive to the training of the model and will have a negative impact on the classification ability of the model. Summary of the invention
[0006] In order to solve at least one of the technical problems existing in the prior art to a certain extent, the present invention aims to provide an end-to-end speech keyword detection method, device and medium based on dynamically adding auxiliary categories.
[0007] The first technical solution adopted by the present invention is:
[0008] An end-to-end speech keyword detection method comprises the following steps:
[0009] S1. Process the training sample labels and pre-add placeholders for auxiliary categories to add auxiliary categories during subsequent training. Among them, the auxiliary categories are divided into two major categories, which are respectively used to process non-keyword samples prone to false alarms and low-quality keyword samples prone to missed detections.
[0010] S2. Use the tonal syllables as modeling units to construct a speech keyword detection model, and normally train the model for E1 rounds to enable the model to have the basic ability to distinguish between keyword and non-keyword categories.
[0011] S3. Starting from the E1 round, add the first type of auxiliary category. Add 1 auxiliary category every E gap rounds until N1 auxiliary categories are added. The method of adding auxiliary categories is as follows: for a tonal syllable belonging to the non-keyword category, its label is changed from the non-keyword category label to the corresponding label of the auxiliary category. N1 is a hyperparameter, which is similar to the number of syllable categories N keyword corresponding to the keyword table.
[0012] S4. Starting from the E2 round, add the second type of auxiliary category, where E1 < E2. The method of adding auxiliary categories is as follows: for a low-quality training sample containing a keyword, change the label of the syllable corresponding to its keyword from the keyword category label to the auxiliary category label. Low-quality samples with multiple keyword syllables can belong to the same second type of auxiliary category label.
[0013] S5. Input the test sample into the trained speech keyword detection model to obtain the posterior probability output matrix, use the prefix beam search algorithm for decoding, and determine whether the decoding result contains the predefined keyword.
[0014] Further, in step S1, the original label set C contains N start = N keyword + N nonkeyword + N blank output categories; where N nonkeyword is 1, representing the non-keyword category, and N blank is 1, representing the CTC blank label; the label set C1 for adding the first type of auxiliary category has N1 auxiliary categories, and the label set C2 for adding the second type of auxiliary category has N2 auxiliary categories; after pre-adding the auxiliary categories, the label set becomes C ′ = {C, C1, C2}, with a total of N class = N start + N1 + N2 output categories.
[0015] Further, the original label set C = {C kw1 , C kw2 , …, C kwN , C filler , Cblank}; Among them, C kw1 ,C kw2 ,…,C kwN are N pre-defined keyword category labels, C filler is the non-keyword category label, C blank It is a blank label for CTC;
[0016] After adding the auxiliary category placeholder, the label set C ′ ={C,C1,C2}, where C2={C ax2}, There are N1 pre-set first auxiliary categories, C ax2 There is a pre-set second auxiliary category.
[0017] Furthermore, the speech keyword detection model comprises a convolutional downsampling layer, a backbone layer and a classification layer in sequence, and the features of the input model adopt the logarithmic Mel spectrum of the audio signal;
[0018] The backbone layer of the model uses the convolution-enhanced Transformer model, referred to as Conformer, and other neural network models that support end-to-end training can also be used; the classification layer uses a fully connected neural network;
[0019] The training process calculates the CTC loss of the posterior probability output matrix and the target label sequence output by the model:
[0020]
[0021] Among them, x represents the input feature sequence corresponding to the speech sample, l represents the target label sequence corresponding to the annotated text, and B represents the many-to-one mapping function between the target label sequence l and all feasible paths π; the loss is specifically calculated by a forward-backward algorithm based on dynamic programming.
[0022] Furthermore, the step S3 comprises:
[0023] Starting from round E1, based on the recognition results of the model on the validation set, each E gap In each round, select one syllable with the highest number of false alarms from the syllables in the non-keyword category and which is not set as an auxiliary category, and set the label of the selected syllable in all training corpora to the corresponding label in the auxiliary category set C1 before the next round of training begins; E gap Set it according to the total number of training rounds to ensure that the model adds N1 auxiliary categories before the end of training, and continue training for a certain number of rounds to make the model fully converge.
[0024] Furthermore, for weakly supervised speech keyword annotated samples, that is, the sample annotation only contains the text transcription of the corresponding speech sample but does not contain the frame-level annotation of the speech sample or the keyword position information of the speech sample, when the sample generates a false alarm, the similarity between the segment in the text A that generates the false alarm sample and the keyword text W that is the false alarm is calculated by the text pronunciation similarity calculation tool, and a segment of the same length as W in the text A is selected using a sliding window, and the similarity is calculated with W. The higher the similarity score, the higher the similarity between the two text segments, and the text segment A with the highest similarity score is selected. max As a false alarm segment in the false alarm sample, the text segment A max The included syllables are used to count the number of false alarms in the validation set of this round.
[0025] Furthermore, the first auxiliary category is selected in the following manner:
[0026] After adding the first auxiliary category, when the model is evaluated using the validation set, the posterior probability matrix output by the model is decoded, and it is determined whether the decoded result contains the predefined keywords. If the target label sequence does not contain the keywords but the keywords appear in the decoded result, it is determined to be a false alarm.
[0027] In the weakly annotated target sequence, according to the keyword text sequence of the false alarm, the fuzzy sound matching tool library is used to perform similarity matching in the target sequence, and the segment with the highest similarity to the keyword of the false alarm in the target sequence is found as the segment of the false alarm, and each syllable contained in the segment is recorded as a false alarm;
[0028] Count all the syllables that generate false alarms and their false alarm times in the current round of verification set, and add the syllable with the highest false alarm number that does not belong to the keyword category and has been added to the auxiliary category to the auxiliary category;
[0029] After the auxiliary categories are added, according to the updated label set C ′ Regenerate the labels of the training samples and save the label set C ′ .
[0030] Furthermore, the step S4 comprises:
[0031] Starting from round E2, the keyword confidence score S output by the keyword samples in the training set during the training process is counted, and the score S is used to determine whether the keyword is well learned by the model. cont If the keyword confidence score S of a keyword sample in a round is less than the threshold θ, the keyword sample is regarded as a low-quality sample, and the keyword syllable with a low confidence score in the keyword segment is changed to the corresponding label in the auxiliary category set C2 before the next round of training begins.
[0032] Furthermore, the second auxiliary category is selected as follows:
[0033] After the second auxiliary category is added, the CTC loss is calculated after the forward propagation of the training process, and the keyword confidence score is obtained at the same time. The keyword sample number and number of times the keyword confidence is lower than the preset value are counted;
[0034] If the keyword confidence score of a keyword sample is lower than θ for multiple consecutive rounds ′ , then the keyword sample is regarded as a low-quality sample and added to the low-quality sample list;
[0035] After this round of training is completed, the labels of all training samples are regenerated, and the specified keyword fragments in the target sequence of low-quality samples are added to the second auxiliary category set C2. When generating labels, the labels of the low-quality keyword fragments of the target sequence are changed from the original keyword labels to the second auxiliary category labels.
[0036] Furthermore, the step S5 comprises:
[0037] In the posterior probability matrix output by the model, the label types include keyword category labels, non-keyword labels, labels corresponding to auxiliary category sets C1 and C2, and CTC blank labels. During decoding, the CTC blank labels are first removed, and then consecutive identical labels are merged and searched for keyword category labels. At this time, the labels corresponding to auxiliary categories C1 and C2 are treated as non-keyword labels.
[0038] The second technical solution adopted by the present invention is:
[0039] An electronic device comprises a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, at least one program, the code set or the instruction set is loaded and executed by the processor to implement an end-to-end voice keyword detection method as described above.
[0040] The third technical solution adopted by the present invention is:
[0041] A computer-readable storage medium stores at least one instruction, at least one program, a code set or an instruction set, wherein the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by a processor to implement an end-to-end voice keyword detection method as described above.
[0042] The fourth technical solution adopted by the present invention is:
[0043] A computer program product or a computer program includes computer instructions stored in a computer-readable storage medium. A processor of a computer device can read the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the above method.
[0044] The beneficial effects of the present invention are as follows: the present invention improves the keyword-filling model commonly used in speech keyword detection, and expands the non-keyword category by two additional subcategories, namely the first auxiliary category and the second auxiliary category; the first auxiliary category targets non-keyword samples that are prone to false alarms, thereby reducing the overall false alarm rate; the second auxiliary category targets keyword samples that are prone to missed detection, thereby achieving waste utilization of low-quality samples. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the embodiments of the present invention or the drawings of related technical solutions in the prior art are introduced below. It should be understood that the drawings introduced below are only for the convenience of clearly describing some embodiments of the technical solutions of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.
[0046] Figure 1 It is a simplified flow chart of an end-to-end voice keyword detection method based on dynamically adding auxiliary categories in an embodiment of the present invention;
[0047] Figure 2 Detailed flowchart of an end-to-end speech keyword detection method based on dynamically adding auxiliary categories in an embodiment of the present invention;
[0048] Figure 3 The present invention is a flowchart of an end-to-end voice keyword detection method according to an embodiment of the present invention. DETAILED DESCRIPTION
[0049] The embodiments of the present invention are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and are not to be construed as limitations of the present invention. For the step numbers in the following embodiments, they are only provided for the convenience of explanation, and the order between the steps is not limited in any way, and the execution order of each step in the embodiment can be adaptively adjusted according to the understanding of those skilled in the art.
[0050] In the description of the present invention, it should be understood that descriptions involving orientations, such as up, down, front, back, left, right, etc., and orientations or positional relationships indicated are based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be understood as a limitation on the present invention.
[0051] In the description of the present invention, "several" means one or more, "more" means more than two, "greater than", "less than", "exceed" etc. are understood as not including the number itself, and "above", "below", "within" etc. are understood as including the number itself. If there is a description of "first" or "second", it is only used for the purpose of distinguishing the technical features, and cannot be understood as indicating or implying the relative importance or implicitly indicating the number of the indicated technical features or implicitly indicating the order of the indicated technical features.
[0052] In the description of the present invention, unless otherwise clearly defined, terms such as setting, installing, connecting, etc. should be understood in a broad sense, and technicians in the relevant technical field can reasonably determine the specific meanings of the above terms in the present invention based on the specific content of the technical solution.
[0053] In response to the existing technical problems, the present invention dynamically adds auxiliary categories during the training process of the model, thereby reducing the negative impact of acoustic similarity between keyword and non-keyword category samples and low-quality keyword samples on training without increasing the number of output categories too much. The method is also used for weakly supervised training corpus to automatically identify difficult samples and set them as auxiliary categories during the training process, and no manual participation is required during the training process.
[0054] The present invention adds two major categories of auxiliary categories, namely the first category of auxiliary categories and the second category of auxiliary categories; the first category of auxiliary categories is used to process non-keyword samples that are prone to false alarms, thereby reducing the false alarm rate; the second category of auxiliary categories is used to process keyword samples that are prone to missed detection in the training set, thereby realizing waste utilization of low-quality samples.
[0055] Example 1
[0056] like Figure 3 As shown, this embodiment provides an end-to-end voice keyword detection method, using basic language units (such as toned syllables) as recognition units, including the following steps:
[0057] S1. Process the training sample labels and pre-add placeholders for auxiliary categories for adding auxiliary categories in the subsequent training process; the auxiliary categories are divided into two categories, which are used to process non-keyword samples that are prone to false alarms and low-quality keyword samples that are prone to missed detection.
[0058] For example, in step S1, the initial tag set C includes Nstart =N keyword +N nonkeyword +N blank output categories, where N nonkeyword 1, indicating a non-keyword category, N blank 1, indicating a CTC blank label; adding the label set C1 of the first auxiliary category, there are N1 auxiliary categories, adding the label set C2 of the second auxiliary category, there are N2 auxiliary categories; after adding auxiliary categories in advance, the model output label set becomes C ′ ={C,C1,C2}, total N class =N start +N1+N2 output categories.
[0059] In some embodiments, step S1 uses connectionist temporal classifier loss for training, where the connectionist temporal classifier loss is referred to as CTC loss. The original label set C = {C kw1 ,C kw2 ,…,C kwN ,C filler ,C blank}; Among them, C kw1 ,C kw2 ,…,C kwN are N pre-defined keyword category labels, C filler is the non-keyword category label, C blank Blank label for CTC.
[0060] After adding the auxiliary category placeholder, the label set C ′ ={C,C1,C2}, where C2={C ax2}, There are N1 pre-set first auxiliary categories, C ax2 There is a pre-set second auxiliary category.
[0061] S2. Use tonal syllables as modeling units to build a speech keyword detection model. Train the model normally for E1 rounds to enable the model to have basic ability to distinguish between keyword and non-keyword categories.
[0062] In some embodiments, the speech keyword detection model includes a convolutional downsampling layer, a backbone layer, and a classification layer in sequence, and the input model features use the logarithmic Mel spectrum of the audio signal. Specifically, the backbone layer uses a convolution-enhanced Transformer model, referred to as Conformer, and other neural network models that support end-to-end training can also be used; the classification layer uses a fully connected neural network.
[0063] The training process calculates the CTC loss of the posterior probability output matrix and the target label sequence output by the model:
[0064]
[0065] Among them, x represents the input feature sequence corresponding to the speech sample, l represents the target label sequence corresponding to the annotated text, and B represents the many-to-one mapping function between the target label sequence l and all feasible paths π; the loss is specifically calculated by a forward-backward algorithm based on dynamic programming.
[0066] S3, starting from round E1, add the first auxiliary category, every interval E gap Add one auxiliary category per round until N1 auxiliary categories are added; the method of adding auxiliary categories is: for a tonal syllable belonging to the non-keyword category, its label is changed from the non-keyword category label to the auxiliary category corresponding label; N1 is used as a hyperparameter, and the number of syllable categories N corresponding to the keyword table keyword similar.
[0067] In step S3, starting from round E1, according to the recognition results of the model on the validation set, each E gap In each round, select one syllable with the highest number of false alarms from the syllables in the non-keyword category and which is not set as an auxiliary category, and set the label of the selected syllable in all training corpora to the corresponding label in the auxiliary category set C1 before the next round of training begins; E gap Set it according to the total number of training rounds to ensure that the model adds N1 auxiliary categories before the end of training, and continue training for a certain number of rounds to ensure that the model fully converges.
[0068] For weakly supervised speech keyword annotated samples, that is, the sample annotation only contains the text transcription of the corresponding speech sample but does not contain the frame-level annotation of the speech sample or the keyword position information of the speech sample, when the sample generates a false alarm, the similarity between the segment in the text A that generates the false alarm sample and the keyword text W that is falsely alarmed is calculated by a text pronunciation similarity calculation tool (such as Dimsim), and a segment of the same length as W in the text A is selected using a sliding window and similarity is calculated with W. The higher the similarity score, the higher the similarity between the two text segments, and the text segment A with the highest similarity score is selected. max Assuming A is a false alarm segment in the false alarm sample, max The included syllables are used to count the number of false alarms in the validation set of this round.
[0069] In some embodiments, the start round E1 and the number of auxiliary categories N1 in step S3 are set as hyperparameters. The setting of E1 cannot be too small or too large. If the time point for starting to add the first type of auxiliary category is too early, the model does not yet have basic discrimination ability, and the selected difficult non-keyword syllable categories may not have obvious acoustic confusion at this time. If the time point for starting to add the first type of auxiliary category is too late, it will slow down the convergence speed of the model, resulting in a longer training time.
[0070] As an implementation manner, the specific process of selecting the first type of auxiliary category in step S3 includes:
[0071] S3.1. After starting to add the first type of auxiliary category, when using the validation set to evaluate the model, decode the posterior probability matrix output by the model and determine whether the decoded result contains a predefined keyword. If the target label sequence does not contain the keyword but the keyword appears in the decoded result, it is determined that a false alarm has occurred.
[0072] S3.2. In the weakly labeled target sequence, there is no position information of the keyword, but it is necessary to determine the position where the false alarm occurs in the input speech. The present invention adopts a fuzzy matching method to search for the false alarm segment of the target sequence. The specific method is to perform similarity matching in the target sequence using the fuzzy sound matching tool library Dimsim according to the keyword text sequence of the false alarm, find the segment with the highest similarity to the false alarm keyword in the target sequence as the false alarm segment, and record each syllable included in this segment once for the false alarm.
[0073] S3.3. According to the above method, count all the syllables that generate false alarms and their false alarm times on the validation set of the current round, and add the syllable with the highest false alarm times and not belonging to the keyword category and already added to the auxiliary category to the auxiliary category.
[0074] S3.4. After the addition of the auxiliary category is completed, according to the updated label set C ′ Regenerate the labels of the training samples and save the label set C ′ .
[0075] S4. Start adding the second type of auxiliary category from round E2, where E1 < E2; the method of adding the auxiliary category is: for a low-quality training sample containing a keyword, modify the label of the syllable corresponding to the keyword to the auxiliary category label. Multiple low-quality samples of keyword syllables can belong to the same second type of auxiliary category label.
[0076] In step S4, starting from round E2, count the keyword confidence score S output by the keyword samples in the training set during the training process, and use this score to determine whether the keyword is well learned by the model. If for consecutive E contIf the keyword confidence score S of a keyword sample in a round is less than the threshold θ, the keyword sample is regarded as a low-quality sample and the keyword syllable with a low confidence score in the keyword segment is changed to the corresponding label in the auxiliary category set C2 before the next round of training begins.
[0077] In some embodiments, the start round E2 in step S4 is set as a hyperparameter, because the setting of E2 should also be moderate. If it is started too early, the model does not have basic distinguishing ability, and the confidence score of the keyword has little reference value for whether the sample is of low quality; if it is started too late, the model convergence speed will slow down, resulting in longer training time.
[0078] As an optional implementation, the specific process of selecting the second auxiliary category in step S4 includes:
[0079] S4.1. After adding the second auxiliary category, after the forward propagation of the training process, the CTC loss is calculated and the keyword confidence score is obtained while decoding, and the keyword sample number and number with keyword confidence lower than 0.1 are counted; the keyword confidence score calculation formula is:
[0080]
[0081] Wherein, k′ (i=1, 2, ... n) represents the characters contained in the keyword kw, and there are n characters in total; Indicates the output category k i The corresponding probability in the tth frame.
[0082] Since the length of defined keywords may be inconsistent, the confidence score needs to be normalized according to the number of characters contained in the keyword:
[0083]
[0084] Among them, Score kw This is the final value of S mentioned above, which is used as the keyword confidence score in the present invention.
[0085] Since the Softmax function is used to calculate the keyword score, the score of the CTC peak will be affected when the number of categories is different. When the number of categories is large, the absolute value of the peak will become smaller accordingly, and when the number of categories is small, the absolute value of the peak will become larger. Taking this factor into account, the actual threshold used is modified according to the threshold θ combined with the number of categories:
[0086] θ′=θlnN class
[0087] θ′ is the threshold finally adopted by the present invention, where N classRepresents the total number of output categories.
[0088] S4.2. If the keyword confidence score of a keyword sample is lower than θ′ for five consecutive rounds, the keyword sample is regarded as a low-quality sample and added to the low-quality sample list.
[0089] S4.3. After this round of training is completed, the labels of all training samples are regenerated, and the specified keyword fragments in the target sequence of low-quality samples are added to the second category auxiliary category set C2. When generating labels, the labels of the low-quality keyword fragments of the target sequence are changed from the original keyword labels to the second category auxiliary category labels.
[0090] S5. Input the test sample into the trained speech keyword detection model to obtain the posterior probability output matrix, use the prefix beam search algorithm to decode it, and determine whether the decoding result contains the predefined keywords.
[0091] In step S5, in the posterior probability matrix output by the model, the label types include keyword category labels, non-keyword labels, labels corresponding to auxiliary category sets C1 and C2, and CTC blank labels; during decoding, the CTC blank labels are first removed, and then consecutive identical labels are merged, and keyword category labels are searched among them. At this time, the labels corresponding to the auxiliary categories C1 and C2 are all treated as non-keyword labels.
[0092] During inference, the number of categories in the model output decoding graph after adding auxiliary categories has increased compared to when no auxiliary categories are added. These auxiliary categories are treated as non-keyword categories during decoding, which does not affect the keyword search.
[0093] The following is combined with Figure 1-2 The above method is supplemented by the following specific implementation examples.
[0094] See also Figure 1 and Figure 2 This embodiment provides an end-to-end voice keyword detection method based on dynamically adding auxiliary categories, comprising the following steps:
[0095] Step 1: Preparation before training: build a network model, extract the acoustic features of the data set, and generate training sample labels;
[0096] In this embodiment, the convolution downsampling layer uses a 2-layer convolutional neural network with a convolution kernel size of [3,3], and performs 4 times of temporal downsampling to reduce the computational complexity; the backbone layer uses a 6-layer Conformer neural network with an input dimension of 128, 4 attention heads, a feedforward layer dimension of 512, and a convolution kernel size of 31; the classification layer uses a 2-layer fully connected neural network, and finally has N class =N start+N1+N2 output categories.
[0097] In this embodiment, the training set, validation set and test set are all from the AISHELL-2 Chinese Mandarin speech dataset, and 8 keywords are selected, namely: "company", "Internet", "reporter", "country", "Beijing", "Shanghai", "city", and "capital"; the training set is 72.43 hours, the validation set is 7.78 hours, and the test set is 12.98 hours. The ratio of positive samples to negative samples is 1:3, where positive samples refer to samples containing keywords, and negative samples refer to samples that do not contain keywords; the speech dataset extracts 80-dimensional logarithmic Mel spectrum features.
[0098] In this embodiment, the modeling unit of the keyword adopts a tonal syllable.
[0099] Step 2: Perform regular training on the model for a certain number of rounds to enable the model to have basic ability to distinguish between keywords and non-keyword categories.
[0100] In this embodiment, the optimizer uses AdamW during training, the maximum number of training rounds is 80 rounds, the number of samples in each small batch is 64, the learning rate increases linearly from 0 to 0.001 in the first 10 rounds, and then decays by cosine, decaying to 0 in 80 rounds; before adding auxiliary categories, the number of output categories is 1.
[0101] Step 3: After the model has a certain ability to distinguish, start adding the first type of auxiliary categories. According to the classification results of the model on the verification set, count the syllable categories with more false alarms in the current round of verification set, obtain a list of false alarms in the verification set of this round, and add the syllables with the most false alarms that are not keywords and have been added as auxiliary categories to the auxiliary categories.
[0102] In this embodiment, the model starts to add the first auxiliary category from the 10th round, and a total of 10 auxiliary categories are set, and one auxiliary category is added every other round.
[0103] In this embodiment, when the model detects a false alarm on the validation set, it is necessary to find out which segment of the non-keyword sample segments is falsely alarmed as a keyword. However, the samples do not have frame-level annotation information, so it is impossible to determine which word in the target sequence is falsely alarmed as a keyword based on the frame-level label output by the model. This embodiment adopts a fuzzy matching algorithm to find out the corresponding words of the target sequence that are falsely alarmed as keywords: a fuzzy sound Chinese matching tool Dimsim is used, which can calculate the acoustic similarity of two given Chinese words. According to the length of the keyword that generates the false alarm, Chinese words are extracted one by one in the target sequence according to a sliding window of the same length, and the acoustic similarity with the false alarm keyword is calculated. Finally, the Chinese word segment with the highest similarity is identified as the syllable that generates the false alarm.
[0104] For example, the non-keyword target sequence is "This background is really beautiful", but the keyword "Beijing" is detected, resulting in a false alarm. According to the acoustic similarity calculation, it is found that the "background" in the target sequence has the highest similarity with the keyword "Beijing", so the false alarm of "background" in the target sequence is determined to be the keyword "Beijing", and the syllables "bei4" and "jing3" corresponding to "background" are recorded once each.
[0105] In this embodiment, after the syllables corresponding to the auxiliary categories of the first round are added in the current round, sample labels are regenerated before the next round of training begins, and the labels of the corresponding syllables are changed from the non-keyword category to the corresponding auxiliary category.
[0106] Step 4: After the model has a certain degree of distinguishing ability, it starts to add the second category of auxiliary categories. According to the keyword confidence score obtained by the forward propagation of the model on the training set, if a keyword sample has a low keyword confidence score for multiple consecutive rounds, the keyword fragment label in the sample is modified to the auxiliary category label.
[0107] In this embodiment, the model adds samples to the second auxiliary category starting from the 50th round. If the confidence score of the keyword of a positive sample for five consecutive rounds is less than 0.1, it is considered to be a low-quality sample and the sample is added to the low-quality sample list. The sample labels are regenerated before the next round of training, and the labels corresponding to the keyword fragments in the low-quality samples are changed from the keyword category to the second auxiliary category.
[0108] Step 5: After training is completed, the test sample is input into the speech keyword detection model to obtain the posterior probability output matrix, and the prefix beam search algorithm is used for decoding. The auxiliary categories encountered in the decoding are regarded as non-keyword categories.
[0109] In this embodiment, when the model has no performance improvement on the validation set for 10 consecutive rounds, the training is considered to have converged and the training is stopped.
[0110] The present invention proposes two auxiliary categories, focusing on training non-keyword samples that are prone to false alarms and keyword samples that are prone to missed detections, thereby improving the utilization efficiency of training samples and enhancing the performance of voice keyword detection.
[0111] Example 2
[0112] An embodiment of the present invention further provides an electronic device, the electronic device comprising a processor and a memory, the memory storing at least one instruction, at least one program, a code set or an instruction set, the at least one instruction, the at least one program, the code set or the instruction set being loaded and executed by the processor to implement the following Figure 3 An end-to-end speech keyword detection method is shown.
[0113] It is understood that the memory may include a random access memory (RAM) or a read-only memory (ROM). Optionally, the memory includes a non-transitory computer-readable storage medium. The memory may be used to store instructions, programs, codes, code sets, or instruction sets. The memory may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function, instructions for implementing the above-mentioned various method embodiments, etc.; the data storage area may store data created according to the use of the server, etc.
[0114] The processor may include one or more processing cores. The processor uses various interfaces and lines to connect the various parts of the entire server, and executes various functions of the server and processes data by running or executing instructions, programs, code sets or instruction sets stored in the memory, and calling data stored in the memory. Optionally, the processor can be implemented in at least one hardware form of digital signal processing (DSP), field programmable gate array (FPGA), and programmable logic array (PLA). The processor can integrate one or a combination of a central processing unit (CPU) and a modem. Among them, the CPU mainly processes the operating system and application programs; the modem is used to process wireless communications. It can be understood that the above-mentioned modem may not be integrated into the processor, but implemented separately through a chip.
[0115] Since the electronic device is an electronic device corresponding to an end-to-end voice keyword detection method of an embodiment of the present invention, and the principle of solving the problem by the electronic device is similar to that of the method, the implementation of the electronic device can refer to the implementation process of the above-mentioned method embodiment, and the repeated parts will not be repeated.
[0116] Example 3
[0117] The embodiment of the present invention further provides a computer-readable storage medium, wherein the storage medium stores at least one instruction, at least one program, code set or instruction set, and the at least one instruction, the at least one program, the code set or instruction set is loaded and executed by a processor to implement the following Figure 3 An end-to-end speech keyword detection method is shown.
[0118] Those skilled in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium, and the storage medium includes a read-only memory (ROM), a random access memory (RAM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), a one-time programmable read-only memory (OTPROM), an electronically erasable rewritable read-only memory (EEPROM), a compact disc (CD-ROM) or other optical disc storage, magnetic disk storage, magnetic tape storage, or any other computer-readable medium that can be used to carry or store data.
[0119] Since the storage medium is a storage medium corresponding to an end-to-end voice keyword detection method of an embodiment of the present invention, and the principle of solving the problem by the storage medium is similar to that of the method, the implementation of the storage medium can refer to the implementation process of the above-mentioned method embodiment, and the repeated parts will not be repeated.
[0120] Example 4
[0121] In some possible implementations, various aspects of the method of the embodiment of the present invention can also be implemented in the form of a program product, which includes a program code. When the program product is run on a computer device, the program code is used to enable the computer device to perform the steps of an end-to-end voice keyword detection method according to various exemplary embodiments of the present application described above in this specification. Among them, the executable computer program code or "code" for executing each embodiment can be written in a high-level programming language such as C, C++, C#, Smalltalk, Java, JavaScript, Visual Basic, structured query language (e.g., Transact-SQL), Perl, or in various other programming languages.
[0122] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, a plurality of steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0123] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.
[0124] The above embodiments are only for illustrating the technical concept and features of the present invention, and their purpose is to enable ordinary technicians in the field to understand the content of the present invention and implement it accordingly, and they cannot be used to limit the protection scope of the present invention. Any equivalent changes or modifications made based on the essence of the content of the present invention should be included in the protection scope of the present invention.
Claims
1. An end-to-end speech keyword detection method, characterized in that: It includes the following steps: S1. Process the training sample labels and pre-add placeholders for auxiliary categories to add auxiliary categories during subsequent training. Among them, the auxiliary categories are divided into two major categories, which are respectively used to process non-keyword samples prone to false alarms and low-quality keyword samples prone to missed detections; S2. Build a voice keyword detection model and normally train the model for E1 rounds to enable the model to have the basic ability to distinguish between keyword and non-keyword categories; S3. Start adding the first type of auxiliary category from round E1. The method of adding the auxiliary category is: for a tonal syllable belonging to the non-keyword category, its label is changed from the non-keyword category label to the corresponding label of the auxiliary category; S4. Start adding the second type of auxiliary category from round E2, where E1 < E2. The method of adding the auxiliary category is: for a low-quality training sample containing a keyword, change the label of the syllable corresponding to the keyword in it from the keyword category label to the auxiliary category label; S5. Input the test sample into the trained voice keyword detection model to obtain a posterior probability output matrix, perform decoding, and determine whether the decoding result contains a predefined keyword.
2. The end-to-end voice keyword detection method according to claim 1, characterized in that: The original label set C = {C kw1 ,C kw2 ,…,C kwN ,C filler ,C blank }; Among them, C kw1 ,C kw2 ,…,C kwN are N predefined keyword category labels, C filler is the non-keyword category label, C blank It is a blank label for CTC; After adding the auxiliary category placeholder, the label set C ′ ={C,C1,C2}, where There are N1 pre-set first auxiliary categories, C ax2 There is a pre-set second auxiliary category.
3. The end-to-end voice keyword detection method according to claim 1, characterized in that: The step S3 includes: Starting from round E1, based on the recognition results of the model on the validation set, each E gap In each round, a syllable with the highest number of false alarms and not set as an auxiliary category is selected from the syllables included in the non-keyword category, and the label of the selected syllable in the entire training corpus is set to the corresponding label in the auxiliary category set C1 before the next round of training begins; Ensure that N1 auxiliary categories are added before the end of the model training, and continue training for a certain number of rounds to make the model fully converge.
4. The end-to-end voice keyword detection method according to claim 3, characterized in that: For weakly supervised speech keyword annotated samples, that is, the sample annotation only contains the text transcription of the corresponding speech sample but does not contain the frame-level annotation of the speech sample or the keyword position information of the speech sample, when the sample generates a false alarm, the similarity between the segment in the text A that generates the false alarm sample and the keyword text W that is the false alarm is calculated by the text pronunciation similarity calculation tool, and a segment of the same length as W in the text A is selected using a sliding window, and the similarity is calculated with W. The higher the similarity score, the higher the similarity between the two text segments, and the text segment A with the highest similarity score is selected. max As a false alarm segment in the false alarm sample, the text segment A max The included syllables are used to count the number of false alarms in the validation set of this round.
5. The end-to-end voice keyword detection method according to claim 1, characterized in that: The selection method of the first type of auxiliary category is: After starting to add the first type of auxiliary category, when evaluating the model using the validation set, decode the posterior probability matrix output by the model and determine whether the decoding result contains a predefined keyword. If the target label sequence does not contain a keyword but a keyword appears in the decoding result, it is judged that a false alarm has occurred; In the weakly labeled target sequence, according to the keyword text sequence of the false alarm, use the fuzzy sound matching tool library to perform similarity matching in the target sequence, find the segment with the highest similarity to the false alarm keyword in the target sequence as the false alarm segment, and record each syllable included in this segment once for false alarm; Count all the syllables that generate false alarms and their false alarm times on the validation set in the current round, and add the syllable with the highest false alarm times that does not belong to the keyword category and has not been added to the auxiliary category to the auxiliary category; After the auxiliary categories are added, according to the updated label set C ′ Regenerate the labels of the training samples and save the label set C ′ .
6. The end-to-end voice keyword detection method according to claim 1, characterized in that: The step S4 includes: Starting from round E2, the keyword confidence score S output by the keyword samples in the training set during the training process is counted, and the score S is used to determine whether the keyword is well learned by the model. cont If the keyword confidence score S of a keyword sample in a round is less than the threshold θ, the keyword sample is regarded as a low-quality sample, and the keyword syllable with a low confidence score in the keyword segment is changed to the corresponding label in the auxiliary category set C2 before the next round of training begins.
7. The end-to-end voice keyword detection method according to claim 1, characterized in that: The selection method of the second type of auxiliary category is: After starting to add the second type of auxiliary category, calculate the CTC loss while performing decoding after the forward propagation during the training process to obtain the confidence score of the keyword, and count the keyword sample numbers and times with the keyword confidence score lower than the preset value; If the keyword confidence score of a keyword sample is lower than θ′ for multiple consecutive rounds, then regard this keyword sample as a low-quality sample and add this sample to the low-quality sample list; After the training of this round ends, regenerate the labels of all training samples. The specified keyword segment in the target sequence of the low-quality sample is added to the second type of auxiliary category set C2, and when generating the label, the label of the low-quality keyword segment part of the target sequence is changed from the original keyword label to the second type of auxiliary category label.
8. The end-to-end voice keyword detection method according to claim 1, characterized in that: The step S5 includes: In the posterior probability matrix output by the model, the label types include keyword category labels, non-keyword labels, labels corresponding to auxiliary category sets C1 and C2, and CTC blank labels. During decoding, the CTC blank labels are first removed, and then consecutive identical labels are merged and searched for keyword category labels. At this time, the labels corresponding to auxiliary categories C1 and C2 are treated as non-keyword labels.
9. An electronic device, characterized in that: The electronic device includes a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the method described in any one of claims 1 to 8.
10. A computer-readable storage medium, characterized in that: The storage medium stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Low-resource voice keyword detection method based on unsupervised learning and transfer learning
CN116434742A
Keyword spotting system having filler model by keyword model and method for making filler model by keyword model
KR1020110112890A