Speech Recognition Method and Apparatus, and Computer-Readable Storage Medium
By using a specific speech model and domain speech model to calculate the probability and threshold of speech instructions under a small amount of training data, determine whether the speech instruction is a preset instruction, the problem of training speech recognition model under a small amount of training data is solved, and better speech instruction recognition effect and recognition results are achieved.
Patent Information
- Application Number
- CN202111399172.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-19
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2041-11-19
AI Technical Summary
How to train a speech recognition model with a good convergence effect with a small amount of training data and achieve a better speech instruction recognition effect.
By receiving voice commands and inputting them into a specific voice model and a domain voice model, the current probability and probability threshold of the voice command as a preset command are calculated respectively, and whether the voice command is a preset command is determined based on the relationship between the two. A specific speech model is trained in the first sample set, and a specific domain speech model is trained in the training set and the first sample set, and the training set contains samples of preset instructions and non-preset instructions.
With a small amount of training data, a recognition model with good convergence effect can be trained to achieve better voice command recognition effects, and maintain targeted recognition results.
Smart Images

Figure CN114049883B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of speech recognition, and in particular, to a speech recognition method, an apparatus, and a computer-readable storage medium. Background Art
[0002] With the development of computer technology and big data technology, various electronic data have been widely used. As an important part of electronic data, how to intelligently recognize speech data has become a common problem currently faced.
[0003] Traditionally, generally, a large number of speech samples are labeled and used as training data for a model to train a speech recognition model. However, generally, a large number of training samples are required to converge to obtain a speech recognition model with good performance.
[0004] In summary, how to train a speech recognition model with good convergence performance and achieve good speech command recognition effect with a small amount of training data has become an urgent problem to be solved. Summary of the Invention
[0005] The technical problem solved by this application is how to provide a speech recognition method that can train a speech recognition model with good convergence performance and achieve good speech command recognition effect with a small amount of training data.
[0006] To solve the above problems, an embodiment of this application provides a speech recognition method, which includes: receiving a speech command; inputting the speech command into a specific speech model to obtain a current probability that the speech command is a preset command; inputting the speech command into a specific domain speech model to obtain a probability threshold for determining whether the speech command is the preset command; determining whether the speech command is a preset command according to the relationship between the current probability and the probability threshold; where the specific speech model is trained with a first sample set as samples, the specific domain speech model is trained with a training set and the first sample set as samples, and the training set includes the first sample set corresponding to the preset command and a second sample set corresponding to a non-preset command.
[0007] Optionally, both the first sample set and the training set include a plurality of instructions, and each instruction includes a plurality of speech signal frames; the specific speech model includes a first clustering module and an identification module. The first clustering module is used to perform frame-by-frame analysis on the speech signal frames in the input instruction, and the identification module is used to calculate the probability that the input instruction is a preset instruction according to the frame-by-frame analysis results of each speech signal frame in the input instruction. The training step of the first clustering module includes: using the speech signal frames of the plurality of instructions in the first sample set as training samples to perform model training on the initial first clustering module to obtain the trained first clustering module; the training step of the identification module includes: using the clustering results obtained by the first sample set passing through the first clustering module as training samples to perform model training on the initial identification module to obtain the trained identification module; the specific domain speech model includes a second clustering module and the identification module, and the second clustering module is used to perform frame-by-frame analysis on the speech signal frames in the input instruction; the training step of the second clustering module includes: using the speech signal frames of each instruction in the training set as training samples to perform model training on the initial second clustering module to obtain the trained second clustering module.
[0008] Optionally, the first clustering module and / or the second clustering module includes a Gaussian mixture clustering model, and the identification module includes a hidden Markov model.
[0009] Optionally, when the probability threshold is less than the current probability, and / or, the absolute value of the difference between the current probability and the probability threshold is greater than or equal to a first preset value, it is determined that the recognition result of the voice instruction is valid.
[0010] Optionally, determining whether the voice instruction is a preset instruction according to the relationship between the current probability and the probability threshold includes: when the value of the current probability is greater than or equal to a second preset value and the recognition result of the voice instruction is valid, the voice instruction is the preset instruction.
[0011] Optionally, determining whether the voice instruction is a preset instruction according to the relationship between the current probability and the probability threshold further includes: when the value of the current probability is less than the second preset value and the recognition result of the voice instruction is valid, the voice instruction is the non-preset instruction.
[0012] Optionally, the method further includes: if it is determined that the recognition result of the voice instruction is invalid, then output a message indicating that the recognition of the voice instruction fails.
[0013] An embodiment of the present application further provides a voice recognition device, which includes: an instruction receiving module for receiving voice instructions; a probability calculation unit for inputting the voice instructions into a specific voice model to obtain the current probability that the voice instructions are preset instructions; a threshold obtaining module for inputting the voice instructions into a specific domain voice model to obtain a probability threshold for determining whether the voice instructions are the preset instructions; an instruction determining module for determining whether the voice instructions are preset instructions according to the relationship between the current probability and the probability threshold; wherein, the specific voice model is trained with a first sample set as a sample, and the specific domain voice model is trained with a training set and the first sample set as samples, and the training set includes the first sample set corresponding to the preset instructions and a second sample set corresponding to non-preset instructions.
[0014] An embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, the steps of the above method are implemented.
[0015] An embodiment of the present application further provides a voice recognition device, including a memory and a processor. A computer program that can run on the processor is stored on the memory. When the processor runs the computer program, the steps of the above method are executed.
[0016] Compared with the prior art, the technical solution of the embodiment of the present application has the following beneficial effects:
[0017] An embodiment of the present application provides a voice recognition method, which includes: receiving voice instructions; inputting the voice instructions into a specific voice model to obtain the current probability that the voice instructions are preset instructions; inputting the voice instructions into a specific domain voice model to obtain a probability threshold for determining whether the voice instructions are the preset instructions; determining whether the voice instructions are preset instructions according to the relationship between the current probability and the probability threshold; wherein, the specific voice model is trained with a first sample set as a sample, and the specific domain voice model is trained with a training set and the first sample set as samples, and the training set includes the first sample set corresponding to the preset instructions and a second sample set corresponding to non-preset instructions. Compared with the prior art, the voice recognition method of the embodiment of the present application can train a specific voice model and a specific domain voice model based on a training set with a small amount of data, generate the current probability and probability threshold dynamically corresponding to each input voice instruction, and comprehensively identify the current voice instruction based on the current probability and the probability threshold. Thus, a recognition model with good convergence effect can be trained even with a small amount of training data, and a good voice instruction recognition effect can be achieved.
[0018] Furthermore, the specific speech model is trained based on the first sample set. It has a small amount of training data and higher pertinence in the recognition of speech commands. The clustering module of the specific-domain speech model is trained based on a training set with a larger amount of data, and can introduce the analysis ability learned from the frame-by-frame analysis results of more command samples. If the recognition module in the specific-domain speech model is trained based on the same training samples as the second clustering module, it will introduce interference factors of non-preset commands, which will instead reduce the pertinence of the recognition result. Therefore, in this embodiment, the parameters of the recognition module in the specific-domain speech model are kept consistent with those of the recognition module in the specific speech model, so that while the specific-domain speech model introduces higher analysis ability, it maintains the pertinence of the recognition result, thus having a better speech recognition effect.
[0019] Furthermore, the embodiment of the present application creates a dynamic threshold independent of the fixed training threshold generated during the model training stage, that is, the probability threshold for each speech command, to implement a soft decision scheme. Furthermore, by combining the two decision forms of soft decision and hard decision, decisions can be made for more complex speech recognition situations. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 is a schematic flowchart of the first speech recognition method according to the embodiment of the present application;
[0021] Figure 2 is a schematic diagram of a data set for model training according to the embodiment of the present application;
[0022] Figure 3 is a simplified training diagram of the specific speech model and the specific-domain speech model according to the embodiment of the present application;
[0023] Figure 4 is a partial flowchart of a specific speech recognition method according to the embodiment of the present application
[0024] Figure 5 is a schematic structural diagram of a speech recognition device according to the embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0025] As mentioned in the background art, the current problem is: how to train a speech recognition model with good convergence effect and achieve good speech command recognition effect under the condition of a small amount of training data.
[0026] Specifically, the time - series based statistical models provide a good internal aspect of the data distribution, making them suitable for classification, decision - making, and prediction. Therefore, they are widely used in many artificial recognition / prediction tasks in modern mobile devices. For example, in intelligent speech recognition tasks, time - series based statistical models are widely applied. In most cases, these models identify the input speech data and return the observations of the corresponding category, direction, or trend probability of the input speech data.
[0027] However, these models always require some auxiliary information (such as the threshold corresponding to the calculated probability) to help obtain the final recognition result. When the speech recognition task runs on a device with limited resources, there are mainly the following two difficulties: (1) These statistical models need to be trained by as much training data as possible to ensure their effectiveness in practical applications. However, the training data does not always well represent the distribution of real - world speech data. This makes the recognition results of the model easily fall into local maxima / minima, that is, overfitting occurs, resulting in unclear recognition results. (2) In a device with limited computing or storage resources, the number of statistical models that can be run at one time is very small. This will greatly limit the auxiliary information that can be provided and also reduce the accuracy of the recognition results.
[0028] Currently, there is a sequence statistical model using the Gaussian Mixture Model - Hidden Markov Model (GMM - HMM) system, which has great potential applications in artificial recognition / prediction tasks. Because the GMM - HMM system can well simulate the acoustic model in the Automatic Speech Recognition (ASR) system. Although such models require good sample training data and multiple carefully associated prediction models, it has been the standard method for ASR for decades. To further improve the robustness of the system, a garbage model has also been proposed, which is used to collect and model any unwanted samples to improve accuracy and reduce false alarms. The idea can be referred to in the US patent (Patent No.: US5895448A) "Methods and Apparatus for Generating and Using Speaker - Independent Garbage Models for Speaker - Dependent Speech Recognition Purposes" and the works of Rose, R.C., Paul, D.B, such as "A Hidden - Markov - Model - Based Keyword Recognition System".
[0029] Based on the above - mentioned system, identify the data sequence of a voice command and output a probability or likelihood value to determine whether the data sequence comes from the voice command of the real world simulated by the system. When the output probability / likelihood value is greater than the fixed training threshold generated in the training phase, this data sequence can be regarded as an example of a voice command of a real - world phenomenon, and this judgment can be called a hard decision.
[0030] However, the training thresholds are based on limited training samples, which cannot reflect the distribution of all voice commands in the real world. In many cases, hard decisions may lead to recognition errors, with low recognition accuracy and high warning problems. Even after applying the junk model, it is still limited in practical applications, prone to bias, and unable to overcome the limitations of training data.
[0031] In order to solve the above problems, an embodiment of the present application provides a speech recognition method. In order to make the present application clearer, the specific scheme is introduced below in combination with the schematic diagrams of each embodiment.
[0032] See also Figure 1 , Figure 1 The following is a flow chart of a speech recognition method according to an embodiment of the present application. The speech recognition method can be executed by a terminal, and the terminal can include a mobile phone, a computer, a tablet computer, a smart watch, an intelligent robot, a server, and a server cluster. The method can include the following steps S101 to S104, which are described in detail as follows.
[0033] Step S101, receiving a voice command.
[0034] The voice command is a piece of speech to be recognized. It can be speech received by the terminal from other devices, or it can be speech received by the terminal through a built-in or external recording device. After receiving the voice command, the terminal determines whether it is a preset command that needs to perform a corresponding operation.
[0035] The preset instruction includes one or more preset instructions. For example, the preset instruction includes "start the search interface". After the terminal recognizes that the received voice instruction is the preset instruction of "start the search interface", the terminal calls the search interface of the browser or the terminal system, and displays the search interface on the display interface such as the screen. It should be noted that the preset instruction can be one or more instructions set based on demand, including but not limited to the aforementioned examples, and the preset instruction can be stored in the terminal in the form of text or voice.
[0036] Step S102: input the voice instruction into a specific voice model to obtain a current probability that the voice instruction is a preset instruction.
[0037] The specific speech model is used to calculate the probability that the speech instruction is a preset instruction (referred to as the current probability), and the current probability is used to determine whether the speech instruction is a preset instruction. Optionally, when the current probability of the speech instruction is higher than a preset threshold, the speech instruction is determined to be a preset instruction.
[0038] A specific speech model takes the collected speech data as samples and learns through big data training the ability to recognize the current probability that the input speech command is a preset command. Optionally, the specific speech model may include a model for Automatic Speech Recognition (ASR) trained based on big data samples. It can be a model trained by different modeling methods, such as an end-to-end model based on Recurrent Neural Network Transducer (RNN-T), an end-to-end model based on Google's Transformer model framework, a speech model based on Weighted Finite State Transducers (WFST), and so on.
[0039] Step S103: Input the speech command into the specific-domain speech model to obtain a probability threshold for determining whether the speech command is the preset command.
[0040] Step S104: Determine whether the speech command is a preset command according to the relationship between the current probability and the probability threshold.
[0041] Among them, the specific speech model is trained with a first sample set as samples, the specific-domain speech model is trained with a training set and the first sample set as samples, and the training set includes the first sample set corresponding to the preset command and a second sample set corresponding to non-preset commands.
[0042] Optionally, the specific-domain speech model is used to determine the probability that the input speech command belongs to the speech commands in the preset domain set. The type of the specific-domain speech model may be the same as or different from that of the specific speech model. That is, the specific-domain speech model may include a model for ASR trained based on big data samples. It can be a model trained by different modeling methods, such as an end-to-end model of RNN-T, an end-to-end model based on Google's Transformer model framework, a speech model based on WFST, and so on.
[0043] Among them, the preset domain set is a data set of the fields that need to be concerned about when performing speech recognition in the embodiments of the present application, and it can correspond to the common situations of our speech command recognition. For example, when the current speech recognition is for Chinese speech commands, the preset domain set may include various Chinese speech commands. Another example is that when the current speech recognition is for e-commerce speech commands, the preset domain set may include various speech commands in the e-commerce field.
[0044] In addition, the embodiments of the present application may further include a preset set (which may also be referred to as a specific set), and the preset set is a voice data set corresponding to a preset instruction. Further, for one or more preset instructions, voice data obtained by issuing the same preset instruction in different expressions can be collected as data in the preset domain set. Different expressions of the same preset instruction may include speaking the same preset instruction in multiple languages (such as Chinese, English, French, and multiple dialects), and different expressions of the same preset instruction may also include. Or, voice data obtained by different people issuing the same preset instruction (which may include people of different ages and genders) can be collected as data in the preset domain set.
[0045] The training data for model training usually consists of a limited number of data samples from the real world. The embodiments of the present application design a selection process for training data to ensure that a data set (abbreviated as the training set) of training data with a relatively small amount of selected data can still provide good information for the subsequent modeling process. For example, this training set can be a corpus composed of dozens of hours of voice data.
[0046] Please refer to Figure 2 , Figure 2 FIG. is a schematic diagram of a data set for model training according to an embodiment of the present application. The training set 201 is a sample data set collected for training a model (including a specific domain voice model and a specific voice model). The preset domain set (also referred to as the specific domain set Specific domain set) 202 is a data set of the domain that needs to be concerned when performing voice recognition in the embodiments of the present application, and the preset set 203 is a data set corresponding to the preset instruction of the embodiments of the present application. The preset set 203 is a subset of the preset domain set 202. There is an intersection between the training set 201 and the preset set 203. The data in the training set 201, the preset set 202, and the preset domain set 202 are all data collected from the real world, so they are all subsets of the real world data set 204. In actual situations, the data in the preset domain set 202 and the preset set 202 are almost infinite. In the embodiments of the present application, the collected training set 201 is used as a part of the data for model training.
[0047] In a specific embodiment, the generation process of the training set 201 includes: collecting some data in the preset domain set 202 as the data in the training set 201. Mark the voice data corresponding to the preset instruction in the training set 201, and the marked data is the intersection of the training set 201 and the preset set 203. Thus, using the intersection of the training set 201 and the preset set 203 as the first sample set, and using the other data in the training set 201 except the first sample set as the data of the second sample set. The second sample set corresponds to a non-preset instruction, and the non-preset instruction is other instructions except the preset instruction.
[0048] The first sample set is used to train a specific speech model so that the specific speech model can determine the probability that the input speech command is a preset command, that is, the current probability. The training set 201 is used to train a specific-domain speech model so that the specific speech model can determine the probability that the input speech command belongs to a preset domain set, and use it as the probability threshold for determining the speech command. When the relationship between the current probability of the speech command recognized this time and its probability threshold meets the preset determination condition, it can be determined that the speech command recognized this time is a preset command.
[0049] By Figure 1 the method described above, a specific speech model and a specific-domain speech model can be trained based on a training set with a small amount of data, generating the current probability and probability threshold dynamically corresponding to each input speech command, and comprehensively using the current probability and probability threshold to identify the speech command this time. Thus, even with a small amount of training data, a recognition model with good convergence effect can be trained, achieving a good speech command recognition effect.
[0050] In one embodiment, please refer to Figure 3 , Figure 3 is a training schematic diagram of the specific speech model 31 and the specific-domain speech model 32; both the first sample set and the training set include multiple commands, and each command includes multiple speech signal frames; the specific speech model 31 includes a first clustering module 311 and a recognition module 312. The first clustering module 311 is used to perform frame-by-frame analysis on the speech signal frames in the input command, and the recognition module 312 is used to calculate the probability that the input command is a preset command according to the frame-by-frame analysis results of each speech signal frame in the input command.
[0051] The training steps of the first clustering module 311 may include: using the speech signal frames of multiple commands in the first sample set as training samples to perform model training on the initial first clustering module 311 to obtain the trained first clustering module 311. Among them, the parameters of the trained first clustering module 311 (such as Figure 3 shown) are obtained through this training step.
[0052] Optionally, the trained first clustering module 311 can perform frame-by-frame analysis on the speech signal frames in the input command (that is, Figure 1 the speech command in). Optionally, the clustering result of the speech signal frames of each command in the first sample set may refer to the clustering result of one or more phonemes in the speech signal frames, and the trained first clustering module 311 can recognize multiple phonemes included in the input command.
[0053] The training steps of the recognition module 312 may include: using the clustering results obtained by the first clustering module for the speech signal frames in multiple first sample instructions as training samples to perform model training on the initial recognition module to obtain the trained recognition module 312. Among them, the parameters of the trained recognition module 312 (such as Figure 3 shown in) are obtained through this training step.
[0054] Optionally, the recognition module 312 learns the phoneme features corresponding to the preset instructions according to the frame-by-frame analysis results of the speech signal frames of each instruction in the first sample set and the corresponding preset instructions.
[0055] After inputting the voice instruction into the specific voice model, first obtain the frame-by-frame analysis result of the voice instruction through the trained first clustering module 311, and then input the frame-by-frame analysis result into the trained recognition module 312, so that the trained recognition module 312 calculates the current probability corresponding to the frame-by-frame analysis result.
[0056] The specific domain voice model 32 includes a second clustering module 321 and a recognition module 312. The second clustering module is used to perform frame-by-frame analysis on the speech signal frames in the input instruction. The training steps of the second clustering module 321 may include: using the speech signal frames of each instruction in the training set as training samples to perform model training on the initial second clustering module 321 to obtain the trained second clustering module 321. Among them, the parameters of the trained second clustering module 321 (such as Figure 3 shown in) are obtained through this training step.
[0057] Optionally, the trained second clustering module 321 can perform frame-by-frame analysis on the speech signal frames in the input instruction (that is, Figure 1 the voice instruction in) based on the clustering results of the speech signal frames of each instruction in the training set. Optionally, the clustering results of the speech signal frames of each instruction in the training set may refer to the results of clustering one or more phonemes in these speech signal frames.
[0058] After inputting the voice instruction into the specific domain voice model, first identify multiple phonemes included in the input instruction through the trained second clustering module 321 to obtain the frame-by-frame analysis result of the voice instruction, and then calculate the probability that the voice instruction is a preset instruction through the trained recognition module 312, denoted as the probability threshold.
[0059] In this embodiment, both the specific speech model and the specific domain speech model are composed of a clustering module and a recognition module. The parameters of the clustering modules of the two models (i.e., the first clustering module 311 and the second clustering module 321) are trained based on different training samples, and different frame-by-frame analysis results of the same speech command can be obtained according to their respective parameters. The parameters of the recognition modules of the two models are all trained based on the first sample set corresponding to the preset command, and the probability that the same speech command is the preset command can be determined based on the unified recognition target from the different frame-by-frame analysis results, which are denoted as the current probability and the probability threshold respectively. Based on the relationship between the current probability and the probability threshold, it is determined whether the speech command is the preset command.
[0060] The specific speech model is trained based on the first sample set (which is a subset of the training set), and its training data volume is small, and it has higher pertinence in the recognition of speech commands. The clustering module of the specific domain speech model (i.e., the second clustering module) is trained based on a training set with a larger data volume, and the analysis ability learned from the frame-by-frame analysis results of more command samples can be introduced. If the recognition module in the specific domain speech model is trained based on the same training samples as the second clustering module (i.e., the data in the training set), the interference factors of non-preset commands will be introduced, which will instead reduce the pertinence of the recognition result. Therefore, in this embodiment, the parameters of the recognition module in the specific domain speech model are kept consistent with the parameters of the recognition module in the specific speech model, so that while introducing higher analysis ability, the specific domain speech model maintains the pertinence of the recognition result, and thus has a better speech recognition effect.
[0061] In a specific embodiment, please refer to again Figure 3 , the first clustering module 311 and / or the second clustering module 321 includes a Gaussian Mixture clustering Model (GMM for short), and the recognition module 312 includes a Hidden Markov Model (HMM for short). Further, both the first clustering module 311 and the second clustering module 321 include GMM, the recognition module 312 includes HMM, the specific domain model 31 is constructed based on the GMM-HMM system, and the specific domain speech model 32 is also constructed based on the GMM-HMM system. For the structure of the GMM-HMM system, reference can be made to the relevant descriptions of the existing GMM-HMM system, which will not be elaborated here.
[0062] It should be noted that the first clustering module 311 and / or the second clustering module 321 may also include models or computing modules using other clustering algorithms, such as the K-means clustering algorithm, etc. The recognition module 312 may also include other models or computing modules for speech recognition detection, such as Markov models, neural network models, etc.
[0063] Optionally, the hidden Markov model (HMM) can use algorithms such as the Baum-Welch estimation algorithm or the Viterbi algorithm to obtain the parameters of the trained recognition module.
[0064] In one embodiment, when condition one and / or condition two are satisfied, the recognition result of the voice command is determined to be valid; wherein, condition one includes: the probability threshold is less than the current probability; condition two includes: the absolute value of the difference between the current probability and the probability threshold is greater than or equal to a first preset value.
[0065] Specifically, since the clustering module of the specific domain voice model (i.e., the second clustering module) is trained based on a training set with a larger data volume compared to the clustering module of the specific voice model (i.e., the first clustering module), it can introduce the analysis ability learned from the frame-by-frame analysis results of more command samples. Therefore, when recognizing the same pair of voice commands, the value of the probability threshold should be less than the current probability. At this time, the difference in the recognition abilities of the two models (the specific voice model and the specific domain voice model) after sample training can be reflected. Therefore, the recognition result of this voice command is valid at this time.
[0066] The first preset value is a probability value obtained based on experiments or experience. After recognizing the same voice command by the two models, if the absolute value of the difference between the current probability and the probability threshold is greater than or equal to the first preset value, it indicates that the recognition of this voice command can reflect the difference in the recognition abilities of the two models after sample training. Therefore, the recognition result of this voice command is valid at this time.
[0067] Optionally, when the above condition one or condition two is not satisfied, the recognition result of the voice command is determined to be invalid. Further, the terminal can output a message indicating that the recognition of the voice command fails to prompt the user that the recognition of this voice command exceeds the recognition ability of the trained model and the corresponding recognition result cannot be obtained.
[0068] Optionally, Figure 1 In step S104, determining whether the voice command is a preset command according to the relationship between the current probability and the probability threshold may include: when the value of the current probability is greater than or equal to a second preset value and the recognition result of the voice command is valid, the voice command is the preset command.
[0069] Among them, the second preset value is a probability value obtained based on experiments or experience. When the recognition result of the current voice command is valid, if the current probability is greater than or equal to the first preset value, it can be determined that the voice command is a preset command. Optionally, the second preset value can be a fixed training threshold generated by a specific voice model during the training phase.
[0070] Furthermore, when the value of the current probability is less than the second preset value and the recognition result of the voice command is valid, the voice command is the non-preset command.
[0071] Among them, whether the terminal determines that the voice command is a preset command or a non-preset command, the recognition of the current voice command does not exceed the recognition ability of the trained model.
[0072] In a specific embodiment, please refer to Figure 4 , Figure 4 which is a partial flowchart of a specific voice recognition method according to an embodiment of the present application. After obtaining the current probability and the probability threshold according to Figure 1 steps S102 and S103, the method may further include:
[0073] Step S401, determining whether the recognition result of the voice command is valid. Among them, when condition one and / or condition two are satisfied, it is determined that the recognition result of the voice execution is valid. If it is determined in step S401 that the recognition result of the voice execution is valid, then jump to step S402 to determine whether the value of the current probability is greater than or equal to the second preset value. If the determination result is yes, then jump to step S403 to determine that the voice command is a preset command. If the calculation result of step S402 is no, then jump to step S404 to determine that the voice command is a non-preset command. Additionally, if it is determined in step S401 that the recognition result of the voice execution is invalid, then jump to step S405 to output a message indicating that the recognition of the voice command fails.
[0074] Based on the above specific voice model and specific domain voice model (both of which can be constructed based on the GMM-HMM system), if the input voice command satisfies the characteristics of the data in the first sample set, the value of the current probability should be greater than the second preset value. If the input voice command also satisfies the characteristics of the data in the training set and can reflect the difference in the recognition capabilities of the two models, the probability threshold should be lower than the current probability. If the input voice command exceeds the recognition ability of the model, a message indicating that the recognition of the voice command fails is output.
[0075] Accordingly, the embodiments of the present application create a dynamic threshold independent of the fixed training threshold generated during the model training phase, that is, the probability threshold for each voice command, to implement a soft decision scheme. Furthermore, by combining the two decision forms of soft decision and hard decision, decisions can be made for more complex speech recognition situations.
[0076] In one embodiment, the present application further provides a speech recognition device 50. Please refer to Figure 5 , Figure 5 which is a schematic structural diagram of a speech recognition device 50. The speech recognition device 50 may include:
[0077] A command receiving module 501, configured to receive a voice command;
[0078] A probability calculation module 502, configured to input the voice command into a specific speech model to obtain the current probability that the voice command is a preset command;
[0079] A threshold obtaining module 503, configured to input the voice command into a specific domain speech model to obtain a probability threshold for determining whether the voice command is the preset command;
[0080] A command determination module 504, configured to determine whether the voice command is a preset command according to the relationship between the current probability and the probability threshold;
[0081] Wherein, the specific speech model is trained with a first sample set as a sample, and the specific domain speech model is trained with a training set and the first sample set as samples. The training set includes the first sample set corresponding to the preset command and a second sample set corresponding to a non-preset command.
[0082] Optionally, both the first sample set and the training set include multiple instructions, and each instruction includes multiple speech signal frames; the specific speech model includes a first clustering module and an identification module. The first clustering module is configured to perform frame-by-frame analysis on the speech signal frames in the input instruction, and the identification module is configured to calculate the probability that the input instruction is a preset instruction according to the frame-by-frame analysis results of each speech signal frame in the input instruction. The training step of the first clustering module includes: using the speech signal frames of multiple instructions in the first sample set as training samples to perform model training on the initial first clustering module to obtain the trained first clustering module; the training step of the identification module includes: using the clustering results obtained by the first sample set passing through the first clustering module as training samples to perform model training on the initial identification module to obtain the trained identification module; the specific domain speech model includes a second clustering module and the identification module, and the second clustering module is configured to perform frame-by-frame analysis on the speech signal frames in the input instruction; the training step of the second clustering module includes: using the speech signal frames of each instruction in the training set as training samples to perform model training on the initial second clustering module to obtain the trained second clustering module.
[0083] Optionally, the first clustering module and / or the second clustering module includes a Gaussian mixture clustering model, and the identification module includes a hidden Markov model.
[0084] In one embodiment, the speech recognition device 50 may further include:
[0085] An effectiveness determination module, configured to determine that the recognition result of the speech instruction is effective when the probability threshold is less than the current probability, and / or the absolute value of the difference between the current probability and the probability threshold is greater than or equal to a first preset value.
[0086] In one embodiment, when the value of the current probability is greater than or equal to a second preset value and the recognition result of the speech instruction is effective, the instruction determination module 504 is further configured to determine that the speech instruction is the preset instruction.
[0087] In one embodiment, when the value of the current probability is less than the second preset value and the recognition result of the speech instruction is effective, the instruction determination module 504 is further configured to determine that the speech instruction is the non-preset instruction.
[0088] In one embodiment, the speech recognition device 50 may further include:
[0089] A message output module, configured to output a message indicating that the recognition of the speech instruction fails when it is determined that the recognition result of the speech instruction is invalid.
[0090] For more information about the working principle and mode of the voice recognition device 50, reference can be made to Figures 1 to 4 any relevant description of the voice recognition method, which will not be elaborated here.
[0091] In a specific implementation, the above-mentioned voice recognition device 50 may correspond to a chip with computing functions in the terminal, or a chip with data processing functions, such as a System-On-a-Chip (SOC), a radio frequency chip, etc.; or a chip module including a chip with computing functions in the terminal; or a chip module with a chip having data processing functions, or a terminal.
[0092] Regarding each module / unit included in the various devices and products described in the above embodiments, it may be a software module / unit, a hardware module / unit, or may also be partly a software module / unit and partly a hardware module / unit. For example, for each device and product applied to or integrated into a chip, each module / unit included therein may be implemented in a hardware manner such as a circuit, or at least part of the module / unit may be implemented in the form of a software program that runs on a processor integrated inside the chip, and the remaining (if any) part of the module / unit may be implemented in a hardware manner such as a circuit; for each device and product applied to or integrated into a chip module, each module / unit included therein may be implemented in a hardware manner such as a circuit, and different modules / units may be located in the same component (such as a chip, a circuit module, etc.) or different components of the chip module, or at least part of the module / unit may be implemented in the form of a software program that runs on a processor integrated inside the chip module, and the remaining (if any) part of the module / unit may be implemented in a hardware manner such as a circuit; for each device and product applied to or integrated into a terminal, each module / unit included therein may be implemented in a hardware manner such as a circuit, and different modules / units may be located in the same component (such as a chip, a circuit module, etc.) or different components inside the terminal, or at least part of the module / unit may be implemented in the form of a software program that runs on a processor integrated inside the terminal, and the remaining (if any) part of the module / unit may be implemented in a hardware manner such as a circuit.
[0093] The embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is run by a processor, it executes Figures 1 to 4 the steps of any voice recognition method. The computer-readable storage medium may include non-volatile memory or non-transitory memory, and may also include optical discs, mechanical hard disks, solid-state drives, etc.
[0094] An embodiment of the present application further provides a voice recognition device, including a memory and a processor. A computer program that can run on the processor is stored on the memory. When the processor runs the computer program, the steps of Figures 1 to 4 any voice recognition method are implemented.
[0095] An embodiment of the present application further provides a computer program product, on which a computer program is stored. When the computer program is run by a processor, the steps of Figures 1 to 4 any voice recognition method are implemented.
[0096] It should be understood that the term "and / or" in this article is merely a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article indicates that the associated objects before and after are in an "or" relationship.
[0097] In the embodiments of the present application, "a plurality of" refers to two or more.
[0098] In the embodiments of the present application, the descriptions such as first and second are only for illustration and distinguishing the described objects, without an order, nor do they represent a special limitation on the number of devices in the embodiments of the present application, and cannot constitute any limitation on the embodiments of the present application.
[0099] In the embodiments of the present application, "connection" refers to various connection methods such as direct connection or indirect connection to achieve communication between devices. The embodiments of the present application do not make any limitation on this.
[0100] It should be understood that in the embodiments of the present application, the processor may be a central processing unit (CPU for short), and this processor may also be other general-purpose processors, digital signal processors (DSP for short), application specific integrated circuits (ASIC for short), field programmable gate arrays (FPGA for short), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or this processor may also be any conventional processor, etc.
[0101] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions according to the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired or wireless manner. It should be understood that in various embodiments of the present application, the sequence numbers of the above processes do not indicate the order of execution, and the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0102] In several embodiments provided in the present application, it should be understood that the disclosed methods, devices, and systems can be implemented in other ways. For example, the device embodiments described above are merely illustrative; for example, the division of units is only a logical function division, and there can be other division methods in actual implementation; for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the couplings, direct couplings, or communication connections shown or discussed with each other can be through some interfaces, and the indirect couplings or communication connections of devices or units can be in electrical, mechanical, or other forms.
[0103] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0104] In addition, the functional units in various embodiments of the present application can be integrated into one processing unit, or each unit can be physically included separately, or two or more units can be integrated into one unit. The above integrated units can be implemented in the form of hardware or in the form of a combination of hardware and software functional units.
[0105] The integrated unit implemented in the form of software functional units can be stored in a computer-readable storage medium. The above-mentioned software functional units are stored in a storage medium and include several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute some steps of the methods according to the embodiments of the present application.
[0106] Although the present application is disclosed as above, the present application is not limited thereto. Any person skilled in the art can make various changes and modifications without departing from the spirit and scope of the present application. Therefore, the protection scope of the present application should be subject to the scope defined by the claims.
Claims
1. A voice recognition method, characterized in that, The method includes: Receiving a voice command; Inputting the voice command into a specific voice model to obtain the current probability that the voice command is a preset command; Inputting the voice command into a specific domain voice model to obtain a probability threshold for determining whether the voice command is the preset command; Determining whether the voice command is a preset command according to the relationship between the current probability and the probability threshold; Wherein, the specific voice model is trained with a first sample set as samples, the specific domain voice model is trained with a training set and the first sample set as samples, and the training set includes the first sample set corresponding to the preset command and a second sample set corresponding to non-preset commands; The determining whether the voice command is a preset command according to the relationship between the current probability and the probability threshold includes: When the probability threshold is less than the current probability, and / or, the absolute value of the difference between the current probability and the probability threshold is greater than or equal to a first preset value, determining that the recognition result of the voice command is valid; When the value of the current probability is greater than or equal to a second preset value and the recognition result of the voice command is valid, the voice command is the preset command.
2. The method according to claim 1, characterized in that, Both the first sample set and the training set include multiple commands, and each command includes multiple voice signal frames; The specific voice model includes a first clustering module and a recognition module. The first clustering module is used to perform frame-by-frame analysis on the voice signal frames in the input command, and the recognition module is used to calculate the probability that the input command is a preset command according to the frame-by-frame analysis results of each voice signal frame in the input command; The training step of the first clustering module includes: using the voice signal frames of multiple commands in the first sample set as training samples to perform model training on the initial first clustering module to obtain the trained first clustering module; The training step of the recognition module includes: using the clustering result obtained by the first sample set passing through the first clustering module as a training sample to perform model training on the initial recognition module to obtain the trained recognition module; The specific domain voice model includes a second clustering module and the recognition module. The second clustering module is used to perform frame-by-frame analysis on the voice signal frames in the input command; The training step of the second clustering module includes: using the voice signal frames of each command in the training set as training samples to perform model training on the initial second clustering module to obtain the trained second clustering module.
3. The method according to claim 2, wherein The first clustering module and / or the second clustering module includes a Gaussian mixture clustering model, and the recognition module includes a hidden Markov model.
4. The method according to claim 1, wherein The determining whether the voice command is a preset command according to the relationship between the current probability and the probability threshold further includes: When the value of the current probability is less than the second preset value and the recognition result of the voice command is valid, the voice command is the non-preset command.
5. The method according to claim 1, wherein The method further includes: If it is determined that the recognition result of the voice command is invalid, outputting a message indicating that the recognition of the voice command fails.
6. A voice recognition device, characterized in that, The device includes: A command receiving module for receiving voice commands; A probability calculation unit, configured to input the voice command into a specific voice model to obtain the current probability that the voice command is a preset command; A threshold acquisition module, configured to input the voice command into a specific domain voice model to obtain a probability threshold for determining whether the voice command is the preset command; An instruction determination module, configured to determine whether the voice command is a preset command according to the relationship between the current probability and the probability threshold; Wherein, the specific voice model is trained with a first sample set as samples, the specific domain voice model is trained with a training set and the first sample set as samples, and the training set includes the first sample set corresponding to the preset command and a second sample set corresponding to a non-preset command; Wherein, the instruction determination module includes: A sub-module configured to determine that the recognition result of the voice command is valid when the probability threshold is less than the current probability, and / or, the absolute value of the difference between the current probability and the probability threshold is greater than or equal to a first preset value; A sub-module configured to determine that the voice command is the preset command when the value of the current probability is greater than or equal to a second preset value and the recognition result of the voice command is valid.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is run by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.
8. A voice recognition device, comprising a memory and a processor, wherein a computer program capable of running on the processor is stored on the memory, and is characterized in that, When the processor runs the computer program, the steps of the method according to any one of claims 1 to 5 are executed.
Citation Information
Patent Citations
Methods and apparatus for generating and using speaker independent garbage models for speaker dependent speech recognition purpose
US5895448A
Voice recognition method, device and apparatus and storage medium
CN111508479A
Model training method, voice recognition method, device, equipment and storage medium
CN112435656A