Recognition Method, System, Device and Storage Medium for Multi-Model Voice Command Words

Through the multi-model voice command word recognition method, multi-threaded parallel computing and neural network recognition are used to solve the problems of degraded recognition performance and low efficiency in traditional technologies, and efficient and flexible command word recognition is achieved.

CN116189677BActive Publication Date: 2025-07-08SUZHOU QIMENGZHE NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310174256.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-28
Publication Date
2025-07-08
Estimated Expiration
2043-02-28

AI Technical Summary

Technical Problem

Traditional command word recognition technology has degraded recognition performance when there is vocal interference or other voice interference, and it is impossible to implement multi-core parallel computing, resulting in high error recognition rate and low efficiency.

Method used

Multi-model voice command word recognition method is used to build multiple models by dividing voice command words. Each model is assigned a separate thread to compute in parallel, and features are extracted and calculated using neural networks, and the recognition result is finally determined by the command words with the highest score.

Benefits of technology

It realizes increasing the number of command words without losing recognition performance, improving recognition efficiency and flexibility, reducing misrecognition, and no need to retrain all models when adding or deleting command words.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116189677B_ABST
    Figure CN116189677B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, system, device and storage medium for recognizing multi-model voice command words. The recognition method includes the following steps: constructing a corresponding number of models and the command words supported by each model based on the voice command words to be supported; obtaining the maximum number of models that need to run in parallel during system operation according to the division result, creating a thread pool according to the maximum number of models, and loading the models that need to run. Each model is allocated a separate thread from the thread pool; the main thread performs feature extraction and calculation of the common part on the audio input, and the remaining multiple threads are respectively calculated by the neural networks of the corresponding models; when only one model recognizes the command word, filtering the misrecognition to obtain the final recognition result; when multiple models detect the command word at the same time, the one with the highest score of the command word is used as the final recognition result; the present invention quickly and accurately recognizes voice command words through the models built with command words.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of speech recognition, and in particular, to a method, system, device, and storage medium for recognizing multi-model speech command words. Background Art

[0002] Command word recognition technology is an artificial intelligence technology that enables machines to recognize and understand speech commands. Command word recognition technology has been widely used in our lives, such as smart homes, wearable devices, intelligent vehicle systems, and so on.

[0003] Traditional command word recognition technology requires voice activity detection (VAD) and then obtains the recognized content through acoustic modeling and WFST decoding. It can only accurately recognize command words. If there is human voice interference before and after the command word or other speech spoken by the speaker, the performance of speech recognition will drop significantly. Although fillers can be added before and after when building WFST to absorb additional speech, a large number of misrecognitions will occur in actual use. At the same time, this decoding recognition requires loading a model containing all command words and can only run on one CPU, unable to perform multi-core parallel processing. Summary of the Invention

[0004] The purpose of the present invention is to overcome the deficiencies of the prior art and provide a method, system, device, and storage medium for recognizing multi-model speech command words that models with command words, then recognizes speech command words through the model, and finally obtains the speech recognition result quickly and accurately. At the same time, the number of command words can be increased without sacrificing recognition performance.

[0005] To achieve the above object, the technical solution adopted by the present invention is: A method for recognizing multi-model speech command words, comprising the following steps:

[0006] Based on the speech command words to be supported, divide and construct the corresponding number of models and the command words supported by each model;

[0007] According to the division result, obtain the maximum number of models that need to run in parallel during system operation, create a thread pool according to the maximum number of models, load the models that need to run, and allocate a separate thread for each model from the thread pool;

[0008] The main thread performs feature extraction and calculation of common parts on the audio input, and the remaining multiple threads are respectively calculated by the neural networks of the corresponding models;

[0009] When only one model recognizes the command word, filter out misrecognitions to obtain the final recognition result; when multiple models detect the command word at the same time, use the one with the highest score of the command word as the final recognition result.

[0010] Further, the voice command words are divided according to function, category, and the length of the command words.

[0011] Further, the number of command words for each model does not exceed 10.

[0012] Further, the calculation method of the score of the command word is as follows:

[0013] Where N is the number of frames of the selected command word segment, C h represents the set of selected command word segments, P(c|w) is the probability of a certain command word, that is, the output of the corresponding node of the neural network, and P(b|w) is the probability of a non-command word.

[0014] Further, different length limits are set for different models. Since command words of similar lengths are grouped together, when the duration of the selected command word segment does not reach the set length, it is considered a misrecognition.

[0015] Further, the voice command words include main command words, and secondary command words are hierarchically divided under each main command word;

[0016] Among them, the models constructed by each secondary command word are command words for the specific functions of the main command word or the model constructed by the previous secondary command word in the usage scenario.

[0017] A recognition system for multi-model voice command words, including:

[0018] Group modeling module: Based on the divided voice command words, construct the corresponding number of models and the command words supported by each model;

[0019] Thread pool construction module: Create a thread pool based on the maximum number of models that need to run in parallel during system operation, load the models that need to run, and allocate a separate thread for each model from the thread pool;

[0020] Parallel decoding module: The main thread performs feature extraction and calculation of common parts on the audio input, and the remaining multiple threads are respectively calculated by the neural networks of the corresponding models;

[0021] Recognition result acquisition module: When only one model recognizes a command word, filter out misrecognition to obtain the final recognition result; when multiple models detect command words at the same time, use the command word with the highest score as the final recognition result.

[0022] A processing device includes: one or more processors; a memory for storing one or more programs; wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the foregoing method.

[0023] A readable storage medium stores a computer program, wherein the computer program is configured to execute the foregoing method when running.

[0024] Due to the application of the above technical solution, the present invention has the following advantages compared with the prior art:

[0025] The method, system, device and storage medium for recognizing multi-model voice command words according to the solution of the present invention do not require voice endpoint detection. The number of output layer parameters is smaller. After obtaining the model result, it does not require wfst decoding and only requires simple post-processing to obtain the recognition result. It can increase the number of command words without sacrificing recognition performance. Multi-threaded parallel computing can also improve efficiency, and it is more flexible in use and later maintenance and update. In addition, when adding or deleting command words subsequently, only local adjustments need to be made, and it is not necessary to retrain all models. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] The technical solution of the present invention will be further described below with reference to the drawings:

[0027] Figure 1 It is a flowchart of an embodiment of the present invention;

[0028] Figure 2 It is a flowchart of voice command word recognition after constructing command words in an embodiment of the present invention;

[0029] Figure 3 It is a flowchart of parallel operation of two models A and B constructed with the main command word in step S3 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0030] The present invention will be further described in detail below with reference to the drawings and specific embodiments.

[0031] Refer to Figure 1 , a method for recognizing multi-model voice command words according to an embodiment of the present invention includes the following steps:

[0032] S1 Divide and construct a corresponding number of models and the command words supported by each model based on the voice command words to be supported;

[0033] S2 Obtain the maximum number of models that need to be parallelized during system operation according to the division result, create a thread pool according to the maximum number of models, load the models to be run, and allocate a separate thread from the thread pool to each model;

[0034] S3 The main thread performs feature extraction and calculation of the common part on the audio input, and the remaining multiple threads are respectively calculated by the neural networks of the corresponding models;

[0035] In S4, when only one model recognizes the command word, filtering is performed on misrecognition to obtain the final recognition result; when multiple models detect the command word at the same time, the one with the highest score for the command word is used as the final recognition result.

[0036] As a further preferred embodiment of the present invention, in step S1, the voice command words are divided according to function and the length of the command word or according to category and the length of the command word, and the command words with similar or the same length are grouped together; at the same time, since the command words with similar lengths are grouped together, different length limits can also be set for different models. When the duration of the selected command word segment does not reach the set length, it is considered a misrecognition, which improves the recognition accuracy.

[0037] In addition, to ensure the accuracy of model recognition, the number of command words for each model does not exceed 10.

[0038] As a further preferred embodiment of the present invention, in step S2, since the number of thread pools is equal to the maximum number of model parallelisms during system operation, when a certain model needs to be used, an idle thread is taken out from the thread pool for calculation. When unloading, the thread corresponding to the model is recycled into the thread pool, thereby saving system overhead as much as possible.

[0039] As a further preferred embodiment of the present invention, in step S4, the calculation method of the score of the command word is as follows:

[0040] where N is the number of frames of the selected command word segment, C h represents the set of selected command word segments, P(c|w) is the probability of a certain command word, that is, the output of the corresponding node of the neural network, and P(b|w) is the probability of a non-command word. At the same time, a threshold for the score of the command word can be set when scoring the recognition to suppress misrecognition.

[0041] The final recognition effect is determined by the score of the command word, and the recognition accuracy is high in this way.

[0042] As a further preferred embodiment of the present invention, the voice command words include main command words, and secondary command words are hierarchically divided under each of the main command words; among them, the models constructed by each secondary command word are command words for the specific functions of the main command word or the model constructed by the previous secondary command word in the usage scenario.

[0043] Specifically, the voice command words can include multiple main command words. There are multiple first-level secondary command words under the main command words, and multiple second-level secondary command words under the first-level secondary command words. By analogy, there can be multiple levels of secondary command words. Among them, the model constructed by the first-level secondary command words is a command word for the specific function of the model constructed by the main command words in terms of usage scenarios, and the model constructed by the second-level secondary command words is a command word for the specific function of the model constructed by the first-level secondary command words in terms of usage scenarios.

[0044] See Figure 2 FIG.

[0045] is a specific working flowchart of an embodiment of the present invention. In this embodiment, two models A and B are constructed based on the length of the main command words for use. At the same time, there are also three models C, D, and E composed of the next-level command words of the main command words. Among them, C and D are command words for implementing the specific functions of model A in terms of usage scenarios, and E is a command word for implementing the specific function of B. C and D are also divided according to length. Figure 2 When the system starts, the two initial models A and B run in parallel. When the recognition result is the command word of A, it switches to the parallel operation of C and D; when the recognition result is the command word of B, it switches to E. If A and B or C and D throw recognition results at the same time, the two results will be scored to select the best result. Of course,

[0046] only a specific example is given. In actual use, they can be combined arbitrarily: for example, the number of initial models is three or more, and the models constructed by the next-level command words of the main command words are two, three, or four.

[0047] See Figure 3 , which is a description of a specific embodiment of step S3. Taking the parallel operation of two models A and B constructed by the main command words as an example, it can be extended to the parallel operation of multiple models.

[0048] The recognition method of the present invention has the following advantages: First, it does not require voice endpoint detection, the number of output layer parameters is smaller, and after obtaining the model result, it does not require wfst decoding, and only simple post-processing is needed to obtain the recognition result; it can increase the number of command words without sacrificing recognition performance, and multi-thread parallel computing can also improve efficiency. At the same time, it is more flexible in use and later maintenance and update; in addition, when adding or deleting command words later, only local adjustments need to be made, and it is not necessary to retrain all models.

[0049] The present invention also discloses a recognition system for multi-model voice command words, including:

[0050] Sub-group modeling module: Based on the voice command words to be supported, a corresponding number of models and the command words supported by each model are constructed after division;

[0051] ThreadPool construction module: Create a thread pool based on the maximum number of models that need to run in parallel during the operation of the system, load the models that need to run, and each model is allocated a separate thread from the thread pool;

[0052] Parallel decoding module: The main thread performs feature extraction and calculation of the common part on the audio input, and the remaining multiple threads are respectively calculated by the neural networks of the corresponding models;

[0053] Recognition result acquisition module: When only one model recognizes the command word, filter out the misrecognition to obtain the final recognition result; when multiple models detect the command word at the same time, the one with the highest score of the command word is used as the final recognition result.

[0054] Another embodiment of the present invention also provides an electronic device, including: one or more processors; a memory for storing one or more programs; wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the method described above.

[0055] Another embodiment of the present invention also provides a readable storage medium, in which a computer program is stored, wherein the computer program is set to execute the recognition method of the multi-model voice command words described above when running.

[0056] The above are only specific application examples of the present invention and do not constitute any limitation to the protection scope of the present invention. Any technical solutions formed by equivalent transformation or equivalent substitution fall within the scope of the present invention's rights protection.

Claims

1. A recognition method for multi-model voice command words, characterized in that, It includes the following steps: Based on the speech command words to be supported, divide and construct the corresponding number of models and the command words supported by each model; among them, the speech command words are divided according to function or category and the length of the command words; the number of command words for each model does not exceed 10; According to the division result, obtain the maximum number of models that need to run in parallel during system operation. Create a thread pool according to the maximum number of models, and load the models that need to run. Each model is allocated a separate thread from the thread pool; The main thread performs feature extraction and calculation of the common part on the audio input, and the remaining multiple threads are respectively calculated by the neural networks of the corresponding models; When only one model recognizes the command word, filter out the misrecognition to obtain the final recognition result; when multiple models detect the command word at the same time, the one with the highest score of the command word is used as the final recognition result; Among them, the calculation method of the score of the command word is as follows: , Where N is the number of frames of the selected command word segment, represents the set of selected command word segments, is the probability of a certain command word, that is, the output of the corresponding node of the neural network, is the probability of a non-command word.

2. The recognition method of multi-model voice command words according to claim 1, characterized in that: Set different length limits for different models. Since command words with similar lengths are grouped together, when the duration of the selected command word segment does not reach the set length, it is considered a misrecognition.

3. The recognition method of multi-model voice command words according to claim 1, characterized in that: The speech command words include main command words, and secondary command words are hierarchically divided under each main command word; Among them, the models constructed by each secondary command word are command words for the specific functions of the main command word or the model constructed by the previous secondary command word in the usage scenario.

4. A recognition system for multi-model voice command words, characterized in that, It includes: Sub - model building module: Based on the speech command words to be supported, divide and construct the corresponding number of models and the command words supported by each model; Thread pool building module: Create a thread pool based on the maximum number of models that need to run in parallel during system operation, load the models that need to run, and each model is allocated a separate thread from the thread pool; Parallel decoding module: The main thread performs feature extraction and calculation of the common part on the audio input, and the remaining multiple threads are respectively calculated by the neural networks of the corresponding models; Recognition result acquisition module: When only one model recognizes the command word, filter out the misrecognition to obtain the final recognition result; When multiple models detect the command word at the same time, the one with the highest score of the command word is used as the final recognition result.

5. A processing device, characterized in that, It includes: One or more processors; A memory for storing one or more programs; among them, when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any one of claims 1 - 3.

6. A readable storage medium, characterized in that: A computer program is stored in the storage medium, and the computer program is set to execute the recognition method of multi - model speech command words described in any one of claims 1 - 3 during operation.

Citation Information

Patent Citations

  • Incremental voice command word recognition method

    CN110808036A

  • Artificial intelligence device for providing voice recognition function and method for operating artificial intelligence device

    WO2020226213A1