Reinforcement Method for Wake Word Recognition Model of Intelligent Voice Assistant Based on Genetic Algorithm
The voice assistant wake word recognition model is optimized through genetic algorithms, and the missed awakening words are searched using the dissimilarity measurement between words and the text-to-speech system, which solves the problem of false wakeup by voice assistants and improves the accuracy and security of wake word recognition.
Patent Information
- Application Number
- CN202111299802.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-04
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2041-11-04
AI Technical Summary
The existing voice assistant wake-up model is easily accidentally awakened, resulting in user privacy and security issues, and it is difficult for the existing technology to effectively optimize the phenomenon of false wake-up.
Genetic algorithms are used to define the dissimilarity measure between words, solve multi-objective optimization problems, and use the text-to-voice system to perform automated searches to mine false wake-up words and reinforce the wake-up word recognition model.
It improves the accuracy of the recognition of the voice assistant wake-up words, reduces the frequency of false wake-ups, and enhances the security and privacy protection of the voice assistant.
Smart Images

Figure CN114187899B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of intelligent voice assistant security, and particularly relates to a method and device for strengthening an intelligent voice assistant wake-up word recognition model based on a genetic algorithm. Background Art
[0002] With the rapid development of artificial intelligence, existing devices such as speakers are becoming increasingly intelligent and can interact with users through voice assistants. Users can use voice commands to achieve various functions, such as playing music, searching the web, making phone calls, etc. Therefore, the security of voice assistants is crucial for user privacy and security.
[0003] Existing voice assistants are usually activated by wake-up words. The voice assistant will detect the surrounding voices. Only after the user says the preset wake-up word, the voice assistant will be activated to receive further instructions. Therefore, the correct wake-up of the voice assistant is the key to protecting user privacy and security. Once the voice assistant is accidentally woken up by a program broadcast or during a conversation, it will start recording the surrounding voices, infringing on user privacy and even possibly receiving incorrect instructions, affecting user safety.
[0004] Existing voice assistant wake-up models usually consist of a lightweight model deployed locally and a detection model in the cloud, which are jointly used to identify possible wake-up words in the environment. However, due to reasons such as insufficient training data, the wake-up models of voice assistants are often accidentally woken up, misidentifying non-wake-up words as wake-up words and being wrongly activated, thus bringing many security problems. Summary of the Invention
[0005] The present application aims to at least solve one of the technical problems in the related art to some extent.
[0006] To this end, the first object of the present application is to propose a method for strengthening an intelligent voice assistant wake-up word recognition model based on a genetic algorithm, which solves the problems of difficult mining of accidentally woken-up words in existing voice assistants and frequent accidental wake-up phenomena that are difficult to optimize, provides an efficient and low-cost method for strengthening the voice assistant model, realizes the mining of accidentally woken-up words in the voice assistant, and further strengthens the accidentally woken-up word detection model of the voice assistant.
[0007] The present application uses the phoneme features of wake-up words, defines the dissimilarity measure between different words, adopts a genetic algorithm to solve the multi-objective optimization problem covering the accidental wake-up rate and dissimilarity, and uses a text-to-speech system (TTS) for fast automated search to find as many accidentally woken-up words located on the Pareto front as possible. The accidentally woken-up words are used as retraining samples to strengthen the original wake-up word recognition model, and repeated iteration is performed to improve the accuracy of the model in recognizing wake-up words.
[0008] The second object of this application is to propose a method and device for strengthening the wake-up word recognition model of an intelligent voice assistant based on a genetic algorithm.
[0009] The third object of this application is to propose a non-transitory computer-readable storage medium.
[0010] To achieve the above object, the first aspect embodiment of this application proposes a method for strengthening the wake-up word recognition model of an intelligent voice assistant based on a genetic algorithm, including: Step S10: According to the speaker type, select phoneme features and the range of feature values, and then according to the phoneme features and the range of feature values, select an appropriate number of features and define the dissimilarity between different words; Step S20: Based on the genetic algorithm, design a solution algorithm for optimizing the two objectives of false wake-up rate and dissimilarity simultaneously; Step S30: Connect the Raspberry Pi with the voice assistant, deploy a false wake-up word mining platform, and perform efficient false wake-up word mining by running the solution algorithm to generate a false wake-up sample set; Step S40: After correctly labeling the false wake-up sample set, train the wake-up word detection model; Step S50: Repeat Step S20, Step S30, and Step S40 to repeatedly strengthen the wake-up word detection model until the number of mined false wake-up words is within an acceptable range.
[0011] Optionally, in an embodiment of this application, according to the speaker type, selecting phoneme features and the range of feature values includes:
[0012] For Chinese speakers, select initials, finals, and tones as phoneme features, and the range of feature values is the initials, finals, and four tones used in Chinese;
[0013] For English speakers, select letters as phoneme features, and the range of feature values is Arabic letters and the placeholder " / ".
[0014] Optionally, in an embodiment of this application, for Chinese speakers, the number of selected features is 3 times the number of Chinese characters in the wake-up word, corresponding to the initials, finals, and tones of each Chinese character respectively;
[0015] For Chinese speakers, the dissimilarity between two Chinese words is expressed as:
[0016]
[0017] Wherein, represent two Chinese words, c i represents the feature at the corresponding position, represents the predefined distance between two features.
[0018] Optionally, in an embodiment of this application, for English speakers, the number of selected features is 1.5 times the number of letters in the word, corresponding to the letter or placeholder at that position respectively;
[0019] For English speakers, the dissimilarity between two English words is expressed as:
[0020]
[0021] in, Represents two English words, c i Indicates the phonemes corresponding to the words, represents the predefined dissimilarity between two phonemes, and D, I, and E are the sets of deletion, insertion, and substitution operations required to transform word W1 into word W2.
[0022] Optionally, in one embodiment of the present application, a solution algorithm for simultaneously optimizing both false awakening rate and dissimilarity is designed based on a genetic algorithm, comprising the following steps:
[0023] Step S21: taking the wake-up word, words close to the wake-up word, and the initialized words as initial samples, wherein the dissimilarity of the wake-up word is calculated, and features with a dissimilarity less than a preset value are selected and recombined to obtain words close to the wake-up word;
[0024] Step S22: Evaluate the false awakening rate and dissimilarity of the samples respectively, wherein the false awakening rate is defined as the playing sample;
[0025] Step S23: Select samples according to the Pareto dominance and crowding ranking method to obtain retained samples;
[0026] Step S24: performing a mutation operation on the retained sample set to obtain a sample set of the next generation, wherein the mutation operation includes: randomly selecting two samples in the set and randomly exchanging a feature, or randomly updating a feature of a sample to another feature within a value range;
[0027] Step S25: Repeat steps S22, S23, and S24 until the maximum number of iterations of the algorithm is reached to generate a final sample set.
[0028] Optionally, in one embodiment of the present application, samples are selected according to the Pareto dominance and crowding ranking method, specifically:
[0029] The Pareto front in the set is selected as the retained samples. After the selection, the selected samples are deleted from the set, and the Pareto front of the remaining samples is selected. If the number of selected samples on the Pareto front exceeds the preset number of retained samples after a certain selection, the selected samples are sorted in descending order according to the congestion degree, and the samples are retained one by one until the number of selected samples reaches the preset number of retained samples.
[0030] Optionally, in an embodiment of the present application, a Raspberry Pi is used to connect to a voice assistant, and a miswake word mining platform is deployed. An efficient miswake word mining is performed by running a solution algorithm to generate a miswake sample set, including the following steps:
[0031] Step S31: Run a solution algorithm on a computer to generate samples to be tested, and play the generated samples to the smart speaker through a speaker;
[0032] Step S32: Connect to the smart speaker through a light sensor, determine whether the voice assistant of the smart speaker is activated, and the Raspberry Pi returns the activation result to the computer;
[0033] Step S33: After the preset number of algorithm iterations is reached, record and save the dissimilarity and miswake rate of the samples during the test, and retain the samples with a certain miswake rate as the miswake sample set.
[0034] Optionally, in an embodiment of the present application, correct sample marking is performed on the miswake sample set, specifically:
[0035] After marking the miswake samples as negative classes, positive samples are randomly selected from the initial training dataset so that the ratio of positive and negative samples in the new dataset is the same as the ratio of positive and negative samples in the original dataset.
[0036] Use the cross-entropy loss function as the objective to train the wake word detection model, where the cross-entropy loss function is expressed as:
[0037]
[0038] where y i represents whether the sample is a positive class, and p i represents the probability that the sample is predicted to be a positive class.
[0039] To achieve the above object, an embodiment of the second aspect of the present application proposes a reinforcement device for an intelligent voice assistant wake word recognition model based on a genetic algorithm, including: a feature preprocessing module, an algorithm design module, a miswake word mining module, a training module, and a repetition module. Among them,
[0040] The feature preprocessing module is used to select phoneme features and feature value ranges according to the speaker type, and then select an appropriate number of features and define the dissimilarity between different words according to the phoneme features and feature value ranges;
[0041] The algorithm design module is used to design a solution algorithm for simultaneously optimizing two objectives of miswake rate and dissimilarity based on the genetic algorithm;
[0042] The false wake-up word mining module is used to connect to the voice assistant using a Raspberry Pi, deploy a false wake-up word mining platform, and perform efficient false wake-up word mining by running a solution algorithm to generate a false wake-up sample set;
[0043] The training module is used to train the wake-up word detection model after correctly labeling the samples in the false wake-up sample set;
[0044] The repetition module is used to repeatedly call the algorithm design module, the false wake-up word mining module, and the training module to repeatedly strengthen the wake-up word detection model until the number of mined false wake-up words is within an acceptable range.
[0045] To achieve the above object, a third aspect embodiment of the present application proposes a non-transitory computer-readable storage medium, which can execute an intelligent voice assistant wake-up word recognition model reinforcement method based on a genetic algorithm when the instructions in the storage medium are executed by a processor.
[0046] The intelligent voice assistant wake-up word recognition model reinforcement method, the intelligent voice assistant wake-up word recognition model reinforcement device, and the non-transitory computer-readable storage medium according to the embodiments of the present application solve the problems of difficult false wake-up word mining and frequent false wake-up phenomena that are difficult to optimize in existing voice assistants, provide an efficient and low-cost voice assistant model reinforcement method, realize the false wake-up word mining of the voice assistant, and further strengthen the false wake-up word detection model of the voice assistant.
[0047] The present application uses the phoneme features of wake-up words to define the dissimilarity measure between different words, adopts a genetic algorithm to solve the multi-objective optimization problem covering the false wake-up rate and dissimilarity, and uses a text-to-speech system (TTS) for fast automated search to find as many false wake-up words as possible located on the Pareto front, and uses the false wake-up words as samples for retraining to strengthen the original wake-up word recognition model, and iterates repeatedly to improve the accuracy of the model in recognizing wake-up words.
[0048] The additional aspects and advantages of the present application will be partially given in the following description, partially become obvious from the following description, or be understood through the practice of the present application. Description of the Drawings
[0049] The above and / or additional aspects and advantages of the present application will become obvious and easy to understand from the following description of the embodiments in conjunction with the drawings, where:
[0050] Figure 1 is a flowchart of an intelligent voice assistant wake-up word recognition model reinforcement method provided by Embodiment 1 of the present application;
[0051] Figure 2Another flowchart of the method for strengthening the wake-up word recognition model of the intelligent voice assistant based on the genetic algorithm according to the embodiment of the present application;
[0052] Figure 3 The diagram of the false wake-up word mining platform for the method of strengthening the wake-up word recognition model of the intelligent voice assistant based on the genetic algorithm according to the embodiment of the present application;
[0053] Figure 4 The diagram of the false wake-up word mining efficiency before and after training the model for strengthening with the open-source dataset according to the embodiment of the present application;
[0054] Figure 5 The structural schematic diagram of a device for strengthening the wake-up word recognition model of the intelligent voice assistant based on the genetic algorithm provided in the second embodiment of the present application. Detailed implementation manners
[0055] The embodiments of the present application will be described in detail below. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals indicate the same or similar elements or elements with the same or similar functions from beginning to end. The embodiments described by referring to the accompanying drawings are exemplary and are intended to explain the present application, but should not be construed as a limitation to the present application.
[0056] The method and device for strengthening the wake-up word recognition model of the intelligent voice assistant based on the genetic algorithm according to the embodiments of the present application will be described below with reference to the accompanying drawings.
[0057] Figure 1 The flowchart of a method for strengthening the wake-up word recognition model of the intelligent voice assistant based on the genetic algorithm provided in the first embodiment of the present application.
[0058] As Figure 1 shown, the method for strengthening the wake-up word recognition model of the intelligent voice assistant based on the genetic algorithm includes the following steps:
[0059] Step S10: Select phoneme features and the range of feature values according to the speaker type, and then select an appropriate number of features and define the dissimilarity between different words according to the phoneme features and the range of feature values;
[0060] Step S20: Based on the genetic algorithm, design a solution algorithm for optimizing the two objectives of the false wake-up rate and the dissimilarity simultaneously;
[0061] Step S30: Connect the Raspberry Pi with the voice assistant, deploy the false wake-up word mining platform, and perform efficient false wake-up word mining by running the solution algorithm to generate a false wake-up sample set;
[0062] Step S40: After correctly labeling the false wake-up sample set, train the wake-up word detection model;
[0063] Step S50: Repeat Step S20, Step S30, and Step S40 to repeatedly reinforce the wake word detection model until the number of mis-awakened words mined is within an acceptable range.
[0064] The method for reinforcing the wake word recognition model of the intelligent voice assistant based on the genetic algorithm in the embodiments of the present application includes Step S10: According to the speaker type, select phoneme features and the range of feature values, and then according to the phoneme features and the range of feature values, select an appropriate number of features and define the dissimilarity between different words; Step S20: Based on the genetic algorithm, design a solution algorithm for simultaneously optimizing the two objectives of the false wake-up rate and dissimilarity; Step S30: Connect the Raspberry Pi with the voice assistant, deploy a false wake-up word mining platform, and perform efficient false wake-up word mining by running the solution algorithm to generate a false wake-up sample set; Step S40: After correctly labeling the false wake-up sample set, train the wake word detection model; Step S50: Repeat Step S20, Step S30, and Step S40 to repeatedly reinforce the wake word detection model until the number of mis-awakened words mined is within an acceptable range. Thus, it can solve the problems of difficult false wake-up word mining and frequent false wake-up phenomena that are difficult to optimize in existing voice assistants, provide an efficient and low-cost method for reinforcing the voice assistant model, realize the false wake-up word mining of the voice assistant, and further reinforce the false wake-up word detection model of the voice assistant.
[0065] The present application uses the phoneme features of the wake word to define the dissimilarity measure between different words, adopts the genetic algorithm to solve the multi-objective optimization problem covering the false wake-up rate and dissimilarity, and uses the text-to-speech system (TTS) for fast automated search to find as many false wake-up words as possible on the Pareto front, and uses the false wake-up words as samples for retraining to reinforce the original wake word recognition model, and iterates repeatedly to improve the accuracy of the model in recognizing wake words.
[0066] Further, in the embodiments of the present application, according to the speaker type, selecting phoneme features and the range of feature values includes:
[0067] For Chinese speakers, select initials, finals, and tones as phoneme features, and the range of feature values is the initials, finals, and four tones used in Chinese;
[0068] For English speakers, select letters as phoneme features, and the range of feature values is Arabic letters and the placeholder " / ".
[0069] Further, in the embodiments of the present application, for Chinese speakers, the number of selected features is 3 times the number of Chinese characters in the wake word, corresponding to the initials, finals, and tones of each Chinese character respectively;
[0070] For Chinese speakers, the dissimilarity between normalized Chinese words is expressed as:
[0071]
[0072] Among them, represents two Chinese words, c i represents the feature at the corresponding position, represents the distance between two predefined features. Select α = 100 to obtain a more uniform dissimilarity distribution.
[0073] Furthermore, in the embodiments of the present application, for an English speaker, the number of selected features is 1.5 times the number of word letters, corresponding to the letters or placeholders at that position respectively;
[0074] For an English speaker, the dissimilarity between normalized English words is expressed as:
[0075]
[0076] Among them, represents two English words, c i represents the phoneme corresponding to the word, represents the dissimilarity between two predefined phonemes. D, I, E are the sets of deletion, insertion, and substitution operations required to transform word W1 into word W2 respectively.
[0077] For an English speaker, the number of features of the English wake-up word is selected to be slightly larger than the number of word letters. Specifically, it is selected to be 1.5 times the number of word letters (rounded down).
[0078] Furthermore, in the embodiments of the present application, based on the genetic algorithm, a solution algorithm for simultaneously optimizing the two objectives of false wake-up rate and dissimilarity is designed, including the following steps:
[0079] Step S21: The wake-up word and the word recombined from features close to the wake-up word features (dissimilarity less than ∈), where f can represent initial consonants, final vowels, tones (c) or phonemes (p), and a part of the initialized words are used as the initial population. Among them, by calculating the dissimilarity of the wake-up word, features with a dissimilarity less than the preset value are selected and recombined to obtain a word close to the wake-up word;
[0080] Step S22: Evaluate the false wake-up rate and dissimilarity of the samples respectively. Among them, the false wake-up rate S(W i ) is defined as the proportion of times the voice assistant is falsely woken up during 10 playbacks of the sample, and the dissimilarity D(W i ): = dis(W i , W0) is calculated as follows:
[0081]
[0082]
[0083] Step S23: Select and retain a part of the samples according to the method of Pareto domination and crowding degree sorting. A sample W i is dominated by sample W j if and only if S(W j )≥S(W i ) and D(W j )≥D(W i ). The set of all non-dominated samples in the set is called the Pareto front. The crowding degree C(W i ) of a sample can be calculated from the metrics of its surrounding samples:
[0084] C(W i ) = 2(S a -S b +D a -D b )
[0085] where
[0086]
[0087]
[0088] Similarly, by changing the above metrics to dissimilarity, D a , D b can be obtained;
[0089] Step S24: Perform a mutation operation on the retained sample set to obtain the next generation of sample sets. Among them, the mutation operation includes: randomly selecting two samples in the set and randomly swapping a segment of features, or randomly updating a certain feature of a sample to other features within the value range;
[0090] Step S25: Repeat Step S22, Step S23, and Step S24 until the maximum number of algorithm iterations is reached to generate the final sample set.
[0091] Furthermore, in the embodiment of the present application, the samples are selected according to the method of Pareto domination and crowding degree sorting, specifically:
[0092] Select the Pareto front in the set as the samples to be retained. After selection, delete the selected and retained samples from the set, and continue to select the Pareto front of the remaining samples. If the number of samples selected after a certain selection of the Pareto front samples exceeds the preset number of retained samples, then sort the samples selected this time in descending order according to the crowding degree, and retain the samples one by one until the number of selected samples reaches the preset number of retained samples.
[0093] Furthermore, in an embodiment of the present application, a Raspberry Pi is connected to a voice assistant, a false wake-up word mining platform is deployed, and efficient false wake-up word mining is performed by running a solution algorithm to generate a false wake-up sample set, including the following steps:
[0094] Step S31: running the solution algorithm on the computer to generate samples to be tested, and playing the generated samples to the smart speaker through the speaker;
[0095] Step S32: Connecting to the smart speaker through the light sensor to determine whether the voice assistant of the smart speaker is activated, and the Raspberry Pi returns the activation result to the computer;
[0096] Step S33: After the preset number of algorithm iterations is reached, the dissimilarity and false awakening rate of the samples during the test are recorded and saved, and samples with a certain false awakening rate are retained as a false awakening sample set:
[0097]
[0098] Furthermore, in the embodiment of the present application, after the false awakening sample set is correctly marked as a sample, the existing network model is retrained, specifically:
[0099] After marking the false awakening samples as negative, randomly select positive samples in the initial training data set so that the ratio of positive and negative samples in the new data set is consistent with that of the original data set.
[0100] The cross entropy loss function is used as the target to train the wake-up word detection model, where the cross entropy loss function is expressed as:
[0101]
[0102] Among them, y i Indicates whether the sample is positive, p i Represents the probability that a sample is predicted to be a positive class.
[0103] This application provides a reinforcement method for a trained model. The dataset here is the dataset used for the initial training, which is usually collected by the trainer rather than generated in this application.
[0104] The process of wake-up model reinforcement in this application includes feature preprocessing, multi-objective optimization algorithm design, false wake-up word mining, false wake-up word retraining and other steps. Among them, false wake-up word mining is carried out through a multi-objective genetic algorithm, different features are selected for Chinese and English speakers, and different dissimilarity measures are defined. The model reinforcement method of this application is based on an automated false wake-up word mining platform, which improves security by repeatedly iteratively searching for false wake-up words and then retraining the wake-up word detection model.
[0105] Figure 2 Another flowchart of the method for strengthening the wake-up word recognition model of the intelligent voice assistant based on the genetic algorithm according to the embodiment of the present application.
[0106] As Figure 2 shown, in the method for strengthening the wake-up word recognition model of the intelligent voice assistant based on the genetic algorithm, first, according to the speaker type, select the phoneme features and the feature value range; for Chinese speakers, select initial consonants, finals, and tones as phoneme features, and the feature value range is the initial consonants, finals, and four tones used in Chinese; for English speakers, select letters as phoneme features, and the feature value range is Arabic letters and the placeholder " / "; select an appropriate number of features and define the dissimilarity between different words; use the wake-up word, words close to the wake-up word, and randomly generated words as the initial population; connect to the voice assistant through a Raspberry Pi, deploy a false wake-up word mining platform, evaluate the false wake-up rate and dissimilarity of the samples respectively, then select the samples according to the method of Pareto domination and crowding degree sorting to obtain the retained samples, perform mutation operations on the retained sample set to obtain the next generation of sample sets, judge whether the maximum number of algorithm iterations is reached, if not, re-evaluate, select, and mutate the samples; if so, generate a false wake-up sample set; after correctly marking the false wake-up samples and correct samples, retrain the existing network model; judge whether the false wake-up rate of the model is low enough, if not, run the genetic algorithm again, repeatedly strengthen the wake-up word detection model until the number of mined false wake-up words is within an acceptable range; if so, end.
[0107] Figure 3 A diagram of the false wake-up word mining platform for the method of strengthening the wake-up word recognition model of the intelligent voice assistant based on the genetic algorithm according to the embodiment of the present application.
[0108] As Figure 3 shown, in the method for strengthening the wake-up word recognition model of the intelligent voice assistant based on the genetic algorithm, connect to the voice assistant through a Raspberry Pi, deploy a false wake-up word mining platform, and perform efficient false wake-up word mining, specifically: run the false wake-up word mining algorithm on the computer side to generate samples to be tested; play the samples to be tested to the intelligent speaker through the speaker; connect to the intelligent speaker through a light sensor to judge whether the voice assistant of the intelligent speaker is activated; the Raspberry Pi returns the activation result to the computer.
[0109] Figure 4 A diagram of the false wake-up word mining efficiency before and after training the model with an open-source dataset for strengthening according to the embodiment of the present application.
[0110] As Figure 4 shown, the horizontal axis represents the wake-up words in different open-source datasets, and the vertical axis represents the proportion of false wake-up words obtained after running the false wake-up word mining algorithm, that is, the proportion of false wake-up words in the total test samples, which has decreased significantly after strengthening.
[0111] Figure 5 This is a schematic structural diagram of a reinforcement device for an intelligent voice assistant wake-up word recognition model provided in the second embodiment of the present application.
[0112] As Figure 5 shown, the reinforcement device for the intelligent voice assistant wake-up word recognition model based on the genetic algorithm includes: a feature preprocessing module, an algorithm design module, a false wake-up word mining module, a training module, and a repetition module. Among them,
[0113] The feature preprocessing module 10 is used to select phoneme features and feature value ranges according to the speaker type, and then select appropriate numbers of features and define the dissimilarity between different words according to the phoneme features and feature value ranges;
[0114] The algorithm design module 20 is used to design a solution algorithm for simultaneously optimizing two objectives of false wake-up rate and dissimilarity based on the genetic algorithm;
[0115] The false wake-up word mining module 30 is used to connect to the voice assistant using a Raspberry Pi, deploy a false wake-up word mining platform, and perform efficient false wake-up word mining by running the solution algorithm to generate a false wake-up sample set;
[0116] The training module 40 is used to train the wake-up word detection model after correctly labeling the false wake-up sample set;
[0117] The repetition module 50 is used to repeatedly call the algorithm design module, the false wake-up word mining module, and the training module to repeatedly reinforce the wake-up word detection model until the number of mined false wake-up words is within an acceptable range.
[0118] The reinforcement device for the wake-up word recognition model of the intelligent voice assistant based on the genetic algorithm according to the embodiment of the present application includes: a feature preprocessing module, an algorithm design module, a false wake-up word mining module, a training module, and a repetition module. Among them, the feature preprocessing module is used to select phoneme features and feature value ranges according to the speaker type, and then select appropriate feature quantities and define the dissimilarity between different words according to the phoneme features and feature value ranges; the algorithm design module is used to design a solution algorithm for simultaneously optimizing the two objectives of false wake-up rate and dissimilarity based on the genetic algorithm; the false wake-up word mining module is used to connect to the voice assistant using a Raspberry Pi, deploy a false wake-up word mining platform, and perform efficient false wake-up word mining by running the solution algorithm to generate a false wake-up sample set; the training module is used to train the wake-up word detection model after correctly labeling the false wake-up sample set; the repetition module is used to repeatedly call the algorithm design module, the false wake-up word mining module, and the training module to repeatedly reinforce the wake-up word detection model until the number of mined false wake-up words is within an acceptable range. Thus, it can solve the problems that it is difficult to mine false wake-up words of the existing voice assistant and the false wake-up phenomenon is frequent and difficult to optimize, provides an efficient and low-cost method for reinforcing the voice assistant model, realizes the mining of false wake-up words of the voice assistant, and further reinforces the false wake-up word detection model of the voice assistant.
[0119] The present application uses the phoneme features of the wake-up word to define the dissimilarity measure between different words, adopts the genetic algorithm to solve the multi-objective optimization problem covering the false wake-up rate and dissimilarity, and uses the text-to-speech system (TTS) for rapid automated search to find as many false wake-up words located on the Pareto front as possible. The false wake-up words are used as samples for retraining to reinforce the original wake-up word recognition model, and are iterated repeatedly to improve the accuracy of the model in recognizing wake-up words.
[0120] To implement the above embodiment, the present application also proposes a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the method for reinforcing the wake-up word recognition model of the intelligent voice assistant based on the genetic algorithm in the above embodiment.
[0121] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc., mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of this application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0122] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of this application, "a plurality of" means at least two, such as two, three, etc., unless otherwise specifically defined.
[0123] Any process or method description depicted in a flowchart or otherwise described herein can be understood to represent a module, segment, or portion of code including one or more executable instructions for implementing a customized logical function or process. And the scope of the preferred embodiments of this application includes additional implementations, where the functions may be executed in a manner that is not shown or discussed, including in a substantially simultaneous manner or in a reverse order according to the functions involved, which should be understood by those skilled in the art to which the embodiments of this application pertain.
[0124] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in conjunction with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of the computer-readable medium include the following: an electrical connection part with one or more wirings (electronic device), a portable computer disk cartridge (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or otherwise processing as appropriate, and then storing it in a computer memory.
[0125] It should be understood that various parts of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), and the like.
[0126] Those of ordinary skill in the art of this technology can understand that all or part of the steps carried by the methods of the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.
[0127] In addition, each functional unit in various embodiments of the present application may be integrated into a processing module, or each unit may exist physically alone, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.
[0128] The above-mentioned storage medium may be a read-only memory, a magnetic disk, an optical disc, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.
Claims
1. An enhancement method for the wake word recognition model of an intelligent voice assistant based on a genetic algorithm, characterized in that, Including the following steps: Step S10: According to the speaker type, select phoneme features and the range of feature values. Then, based on the phoneme features and the range of feature values, select an appropriate number of features and define the dissimilarity between different words; Step S20: Based on the genetic algorithm, design a solution algorithm for optimizing both the false wake-up rate and dissimilarity; Step S30: Connect a Raspberry Pi to a voice assistant, deploy a false wake-up word mining platform, and perform efficient false wake-up word mining by running the solution algorithm to generate a false wake-up sample set; Step S40: After correctly labeling the false wake-up sample set, train the wake-up word detection model; Step S50: Repeat Step S20, Step S30, and Step S40 to repeatedly strengthen the wake-up word detection model until the number of mined false wake-up words is within an acceptable range.
2. The method according to claim 1, characterized in that, The selection of phoneme features and the range of feature values according to the speaker type includes: For Chinese speakers, select initials, finals, and tones as phoneme features, and the range of feature values is the initials, finals, and four tones used in Chinese; For English speakers, select letters as phoneme features, and the range of feature values is Arabic letters and the placeholder " / ".
3. The method according to claim 2, wherein For Chinese speakers, the number of selected features is 3 times the number of Chinese characters in the wake-up word, corresponding to the initials, finals, and tones of each Chinese character respectively; For Chinese speakers, the dissimilarity between two Chinese words is expressed as: Among them, represents two Chinese words, c i represents the feature at the corresponding position, represents the distance between two predefined features.
4. The method according to claim 2, wherein For English speakers, the number of selected features is 1.5 times the number of letters in the word, corresponding to the letters or placeholders at the corresponding positions of the word letters; For English speakers, the dissimilarity between two English words is expressed as: Among them, represents two English words, represents the dissimilarity between two predefined phones. D, I, and E are the sets of deletion, insertion, and substitution operations required to transform word W1 into word W2, respectively.
5. The method according to claim 1, characterized in that, The design of a solution algorithm for optimizing both the false wake-up rate and dissimilarity based on the genetic algorithm includes the following steps: Step S21: Use the wake-up word, words close to the wake-up word, and initialized words as initial samples. Among them, by calculating the dissimilarity of the wake-up word, select features with dissimilarity less than a preset value and recombine them to obtain the words close to the wake-up word; Step S22: Evaluate the false wake-up rate and dissimilarity of the samples respectively. Among them, the false wake-up rate is defined as playing the sample; Step S23: Select samples in the order of Pareto dominance and crowding degree sorting to obtain the retained samples; Step S24: Perform a mutation operation on the retained sample set to obtain the next generation of sample sets. Among them, the mutation operation includes: randomly select two samples in the set and randomly exchange a section of features, or randomly update a certain feature of a sample to other features within the range of values; Step S25: Repeat Step S22, Step S23, and Step S24 until the maximum number of algorithm iterations is reached to generate the final sample set.
6. The method according to claim 5, wherein The selection of samples in the order of Pareto dominance and crowding degree sorting is specifically: The Pareto front in the set is selected as the retained samples. After the selection, the selected samples are deleted from the set, and the Pareto front of the remaining samples is selected. If the number of selected samples on the Pareto front exceeds the preset number of retained samples after a certain selection, the selected samples are sorted in descending order according to the congestion degree, and the samples are retained one by one until the number of selected samples reaches the preset number of retained samples.
7. The method according to claim 1, wherein The method of using a Raspberry Pi to connect with a voice assistant, deploying a false wake-up word mining platform, and running the solution algorithm to perform efficient false wake-up word mining to generate a false wake-up sample set includes the following steps: Step S31: running the solution algorithm on the computer to generate samples to be tested, and playing the generated samples to the smart speaker through the speaker; Step S32: Connecting to the smart speaker through the light sensor to determine whether the voice assistant of the smart speaker is activated, and the Raspberry Pi returns the activation result to the computer; Step S33: After the preset number of algorithm iterations is reached, the dissimilarity and false awakening rate of the samples during the test are recorded and saved, and samples with a certain false awakening rate are retained as a false awakening sample set.
8. The method according to claim 1, characterized in that, The correct sample marking of the false wakeup sample set is specifically as follows: After marking the false awakening sample set as a negative class, randomly select positive samples from the initial training data set so that the ratio of positive and negative samples in the new data set is consistent with that of the original data set. The wake-up word detection model is trained using a cross entropy loss function as a target, wherein the cross entropy loss function is expressed as: Among them, y i indicates whether the sample is a positive class, and p i indicates the probability that the sample is predicted to be a positive class.
9. An intelligent voice assistant wake-up word recognition model reinforcement device based on a genetic algorithm, characterized in that, It includes feature preprocessing module, algorithm design module, false wake-up word mining module, training module and repetition module. in, The feature preprocessing module is used to select phoneme features and feature value ranges according to the speaker type, and then select an appropriate number of features and define the dissimilarity between different words according to the phoneme features and feature value ranges; The algorithm design module is used to design a solution algorithm for simultaneously optimizing the false awakening rate and the dissimilarity based on a genetic algorithm; The false wake-up word mining module is used to connect the Raspberry Pi to the voice assistant, deploy a false wake-up word mining platform, and perform efficient false wake-up word mining by running the solution algorithm to generate a false wake-up sample set; The training module is used to train the wake-up word detection model after correctly marking the false wake-up sample set; The repetition module is used to repeatedly call the algorithm design module, the false wake-up word mining module, and the training module to repeatedly reinforce the wake-up word detection model until the number of false wake-up words mined is within an acceptable range.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Voice awakening method and device
CN106782536A
Anti-error awakening method based on multiple acoustic models and voice recognition module
CN112102812A