An Online Instruction Word Speech Recognition Method and System in a Noisy Environment
Through the CNN classification model and activation vector judgment, the conversion speech recognition is used to image recognition, which solves the accuracy and open set recognition problems of intelligent mining robot command word speech recognition in a noisy environment, and realizes efficient command word recognition and non-instruction word rejection.
Patent Information
- Application Number
- CN202111023192.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-08-25
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2041-08-25
AI Technical Summary
The existing sequence reasoning model has poor voice recognition effect in noisy environments, especially in intelligent mining robots, the noise is unstable and variable, and the interference recognition effect is.
The CNN classification model is used to convert speech recognition into image recognition problem, and the MFCC feature vector and activation vector are used for processing, combined with the CNN binary classification network and multi-classification network to achieve the distinction between speech and noise, and judge unknown category speech through the activation vector.
It improves the accuracy of the speech recognition of command words in a noisy environment, solves the problems of online voice endpoint detection and open set recognition, and realizes effective rejection of non-command words.
Smart Images

Figure CN113921000B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of speech recognition, and particularly relates to a method and system for online command word speech recognition in a noisy environment. Background Art
[0002] Speech recognition is a technology that enables a machine to recognize or understand external speech input. This technology is the basis for further interaction between humans and machines. Only when the machine can recognize the input speech is it possible to further understand the recognized content and make corresponding feedback. At present, the development prospect of speech recognition technology is very broad. As a direction for the development of intelligent mining robots, it can assist drivers in operating intelligent mining robots, reducing the operation difficulty of intelligent mining robots, and has a wide range of application prospects.
[0003] In recent years, the rapid development of deep learning has attracted many researchers to devote their energy to research. Related technologies have been applied to the field of speech recognition, which has greatly improved the speech recognition accuracy rate and calculation speed. However, in recent years, less work has been done on isolated words in speech recognition, and the methods used are all based on sequence inference. This method can handle speech recognition in a noise-free environment, but the effect is not ideal in a noisy environment. In the speech recognition task of intelligent mining robots, the noise is non-stationary, may appear at any time, and different noises may appear. For example, in the previous time period, the intelligent mining robot rotates the cockpit, and in the next time period, the intelligent mining robot starts to extend the boom. Such noise interference added to the command word speech signal can greatly interfere with the sequence inference model. If the noise of the current frame is relatively large, it will probably interfere with the prediction output of the current frame, and then cause all subsequent inferences to fail. Summary of the Invention
[0004] In view of the above problems, the present invention proposes a method and system for online command word speech recognition in a noisy environment to solve the problem that the existing sequence inference model has a poor effect on command word speech recognition in a noisy environment.
[0005] According to one aspect of the present invention, a method for online command word speech recognition in a noisy environment is proposed. The method includes the following steps:
[0006] Step 1: Preprocess the real-time input unknown sound signal, and the preprocessing includes high-frequency compensation and frame segmentation processing;
[0007] Step 2: Extract features from the preprocessed unknown sound signal to obtain MFCC feature vectors;
[0008] Step 3: Input the MFCC feature vectors into a trained CNN binary classification network model to identify whether the unknown sound signal is speech or noise;
[0009] Step 4: When it is recognized that the unknown sound signal is speech, splice the MFCC feature vectors corresponding to the speech, and calculate to obtain an MFCC feature map; input the MFCC feature map into the trained CNN multi-classification network model to recognize the speech, and obtain the recognition result of the speech.
[0010] Further, in Step 1, the high-frequency compensation is implemented by using a first-order FIR high-pass filter; the frame segmentation process is to divide the unknown sound signal into multiple short-time speech segments, and each speech segment is called a frame signal.
[0011] Further, the specific process of Step 3 includes:
[0012] Establish a data buffer to store the MFCC feature vectors corresponding to multiple frames of signals;
[0013] Each time the MFCC feature vector corresponding to a new frame of signal enters the data buffer, discard the MFCC feature vector corresponding to the first frame of signal that entered;
[0014] Splice the MFCC feature vectors corresponding to the multiple frames of signals stored in the data buffer, and calculate to obtain an MFCC feature map;
[0015] Input the MFCC feature map into the trained CNN binary-classification network model to recognize whether the multiple frames of signals are speech or noise. If it is speech, continue to process the MFCC feature vectors corresponding to the real-time input and pre-processed unknown sound signals according to the above steps; if it is noise, continue to wait.
[0016] Further, during the training process of the CNN multi-classification network model in Step 4, an activation vector is introduced as the input vector for the softmax operation of the last layer in the CNN multi-classification network model structure. The training process includes: after training the CNN multi-classification network model according to the training set data, input the training set data into the model to calculate the activation vector of each category, and for each category, use the activation vector of this category to fit the corresponding GMM function.
[0017] Further, the specific process of inputting the MFCC feature map into the trained CNN multi-classification network model to recognize the speech and obtaining the recognition result of the speech includes: inputting the MFCC feature map into the trained CNN multi-classification network model, first calculating the known class K with the highest probability to which the MFCC feature map belongs, then calculating the activation vector corresponding to the MFCC feature map, and then substituting the activation vector into the GMM function fitted during the training process of the known class K for calculation to determine whether the speech corresponding to the MFCC feature map belongs to an unknown class or the known class K. When the calculated function value is greater than the preset hyperparameter threshold, it is determined that the speech corresponding to the MFCC feature map belongs to the above-mentioned known class K, otherwise it belongs to an unknown class.
[0018] According to another aspect of the present invention, an online command word speech recognition system in a noise environment is proposed, and the system includes:
[0019] A preprocessing module for preprocessing a real-time input unknown sound signal, and the preprocessing includes high-frequency compensation and framing processing; wherein, the framing processing is to divide the unknown sound signal into multiple short-time speech segments, and each speech segment is called a frame signal;
[0020] A feature extraction module for extracting features from the preprocessed unknown sound signal to obtain MFCC feature vectors;
[0021] A binary classification module for inputting the MFCC feature vectors into a trained CNN binary classification network model to recognize whether the unknown sound signal is speech or noise;
[0022] A speech recognition module for, when the binary classification module recognizes that the unknown sound signal is speech, splicing the MFCC feature vectors corresponding to the speech, calculating to obtain an MFCC feature map; inputting the MFCC feature map into a trained CNN multi-classification network model to recognize the speech and obtaining the recognition result of the speech.
[0023] Further, a first-order FIR high-pass filter is used in the preprocessing module to implement high-frequency compensation.
[0024] Further, the specific process of recognizing whether the unknown sound signal is speech or noise in the binary classification module includes:
[0025] Establishing a data buffer for storing MFCC feature vectors corresponding to multiple frames of signals;
[0026] Each time the MFCC feature vector corresponding to a new frame of signal enters the data buffer, the MFCC feature vector corresponding to the earliest entered frame of signal is discarded;
[0027] Concatenate the MFCC feature vectors corresponding to multiple frames of signals stored in the data buffer, and calculate to obtain an MFCC feature map;
[0028] Input the MFCC feature map into the trained CNN binary classification network model to identify whether the multiple frames of signals are speech or noise. If it is speech, continue to process the MFCC feature vectors corresponding to the preprocessed unknown sound signals input in real time according to the above steps; if it is noise, continue to wait.
[0029] Further, during the training process of the CNN multi-classification network model in the speech recognition module, an activation vector is introduced as the input vector for the softmax operation in the last layer of the CNN multi-classification network model structure. The training process includes: after training the CNN multi-classification network model according to the training set data, input the training set data into the model to calculate the activation vector for each category, and for each category, use the activation vector of that category to fit the corresponding GMM function.
[0030] Further, the specific process of inputting the MFCC feature map into the trained CNN multi-classification network model in the speech recognition module to recognize the speech and obtain the recognition result of the speech includes: input the MFCC feature map into the trained CNN multi-classification network model, first calculate the known category K with the highest probability to which the MFCC feature map belongs, then calculate the activation vector corresponding to the MFCC feature map, and then substitute the activation vector into the GMM function fitted during the training process of the known category K for calculation to determine whether the speech corresponding to the MFCC feature map belongs to an unknown category or the known category K. When the calculated function value is greater than the preset hyperparameter threshold, it is determined that the speech corresponding to the MFCC feature map belongs to the above-known category K, otherwise it belongs to an unknown category.
[0031] The beneficial technical effects of the present invention are:
[0032] The present invention can effectively reject non-instruction speech and achieve accurate recognition of instruction speech in a noisy environment. For the recognition of instruction speech in a noisy environment, the present invention uses a CNN classification model and creatively converts the speech recognition problem into an image recognition problem for processing, thus abandoning the sequence inference used in general speech recognition methods and only recognizing the features of the image, effectively improving the recognition accuracy.
[0033] Further, for online use, that is, for the problem of needing to detect in real time whether there is speech input, the present invention designs a binary classification network model based on CNN, effectively solving the problem of online speech endpoint detection, and being able to accurately recognize the speech signal and use it as an input signal for subsequent recognition.
[0034] Furthermore, for the problem of open-set recognition, the present invention proposes a classification and judgment method based on the input of activation vectors. This method uses the activation vectors of the CNN network as the judgment basis, accurately realizes the classification of unknown-class speech and command-word speech, and well solves the problem of open-set recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] The present invention can be better understood by referring to the description given below in conjunction with the accompanying drawings. The accompanying drawings, together with the following detailed description, are included in this specification and form a part of this specification, and are used to further illustrate the preferred embodiments of the present invention and explain the principles and advantages of the present invention.
[0036] Figure 1 It is a flowchart of an online command-word speech recognition method in a noisy environment according to the present invention;
[0037] Figure 2 It is the MFCC heat (feature) map in the present invention; wherein, Figure (a) is the MFCC heat map of different command-word speeches, (a1) is forward, (a2) is backward, and (a3) is left turn; Figure (b) is the MFCC heat map of the same command-word speech, (b1) is left turn command 1, and (b2) is left turn command 2;
[0038] Figure 3 It is a flowchart of online speech endpoint detection in the present invention;
[0039] Figure 4 It is a schematic diagram of online speech endpoint detection by the CNN binary classification network in the present invention;
[0040] Figure 5 It is a schematic diagram of model training based on activation vectors in the present invention;
[0041] Figure 6 It is a process diagram of input classification and judgment based on activation vectors in the present invention;
[0042] Figure 7 It is a comparison diagram of the effects of different methods in the embodiment of the present invention;
[0043] Figure 8 It is a schematic diagram of the structure of an online command-word speech recognition system in a noisy environment according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0044] To enable those skilled in the art to better understand the solution of the present invention, the exemplary embodiments or examples of the present invention will be described below in conjunction with the accompanying drawings. Obviously, the described embodiments or examples are only a part of the embodiments or examples of the present invention, rather than all of them. All other embodiments or examples obtained by those of ordinary skill in the art based on the embodiments or examples in the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0045] The CNN does not rely on the idea of sequence reasoning and has a certain tolerance ability for noise. Based on this characteristic, the present invention uses a CNN classification model to transform the problem of command word speech recognition into an image recognition problem for processing, effectively solving the drawbacks of the general sequence reasoning method and achieving good command word speech recognition effect under noise. In addition, for the application of model mining robots, considering online speech endpoint detection and open set recognition, the present invention proposes to use a CNN to establish a binary classification network to distinguish speech and noise and a distribution function based on the activation vector to judge the rejection of unknown classes, solving the open set recognition problem.
[0046] As Figure 1 shown, the specific steps of an online command word speech recognition method based on CNN in a noisy environment are as follows:
[0047] Step 1: Preprocessing of speech signals
[0048] When the speech is input into the speech acquisition device, the high-frequency components will be attenuated through the air transmission. Therefore, it is necessary to first perform high-frequency compensation on the speech signal. Secondly, the speech emitted by humans is a non-stationary signal, but it can be considered stationary within a short time, that is, the statistical parameters of the speech signal can be guaranteed to be stable within a very short time range. Therefore, it is also necessary to perform frame segmentation on the speech sequence.
[0049] For high-frequency compensation, the present invention uses a first-order FIR high-pass filter to implement, and the filter expression is:
[0050] y(n) = x(n) - ax(n - 1)
[0051] where x(n) is the speech sampling value at the nth moment in the input speech sequence, y(n) is the speech signal after passing through the FIR high-pass filter, and a is the pre-emphasis coefficient, generally selected as 0.98.
[0052] For frame segmentation, it is necessary to artificially divide the speech signal into multiple short-time speech segments, and these speech segments are called analysis frames. The speech signal used in the present invention is 1 second, and each frame signal is 25 milliseconds.
[0053] Step 2: Feature extraction of speech signals
[0054] Feature extraction is performed on the preprocessed voice signal. In the present invention, Mel-scale Frequency Cepstral Coefficient (MFCC) features of the voice signal are selected for extraction, and the MFCC feature vectors after frame processing are concatenated into an MFCC heat map. As Figure 2 shown, Figure (a) is the MFCC heat map of voice with different command words; Figure (b) is the MFCC heat map of voice with the same command word. From Figure 2 this, it can be seen that the MFCC feature maps of voices with different command words vary greatly, while the MFCC feature maps of voices with the same command word are very similar, which can be understood as pictures of objects of the same category in image recognition problems. Therefore, the command word voice recognition problem can be converted into an image recognition problem.
[0055] Step 3: Online voice endpoint detection
[0056] Since the command word voice recognition algorithm of the intelligent mining robot needs to be used online, it is necessary to continuously detect whether there is voice input during online voice recognition. Until it is judged that there is voice input, then the voice signal is collected, and finally it is substituted into the trained model for recognition. In a noisy environment, the present invention uses a binary classification network implemented by CNN for recognition to judge whether the input is voice or noise. If it is voice, the subsequent signal is collected as voice input for subsequent voice recognition; if it is noise, it continues to wait until there is voice input. The implementation process is as Figure 3 shown. A data buffer is established. The size of the data buffer can store 4 frames of voice data. The MFCC feature map is calculated for the 4 frames of data, and then the MFCC features calculated from the 4 frames of data are used as pictures and input into the CNN network. The corresponding two outputs are "voice" and "noise". After the binary classification network is trained, for external signal input, it is substituted into the binary classification network. When it is judged as "voice", 1 second of microphone input data is read and substituted into the trained model for recognition. The more detailed process is as Figure 4 shown. The data buffer always stores 4 frames of data. Each time a new frame of data enters, the first frame in the data buffer is discarded, and then the MFCC feature map is calculated and substituted into the binary classification network to determine whether it is voice. If it is voice, the input data and the signal in the data buffer are continued to be collected to form an input of 1 second in length. The advantage of doing this is to prevent the voice signal in the data buffer from being lost, resulting in inaccurate subsequent command word voice recognition. It should be noted that one frame of data here is 25 milliseconds.
[0057] Step 4: Open-set recognition
[0058] For a specific recognition problem, the commonly used recognition method is closed-set recognition, that is, it is assumed that the input sample to be tested must belong to a known database. However, in practice, the test samples often contain unknown samples. If the closed-set recognition system is still used, the system will wrongly recognize the test samples from unknown classes as belonging to one of the known closed-set classes, resulting in a decrease in the accuracy rate. Therefore, open-set recognition needs to be introduced in the online speech recognition of the present invention. The goal of open-set recognition is to correctly classify and correctly reject other unknown classes.
[0059] The noisy speech input after online speech endpoint detection is subjected to feature extraction to obtain an MFCC feature map, and then it is regarded as an image recognition problem and substituted into the CNN network for recognition. In order to reject unknown instruction inputs, the present invention introduces an activation vector as a judgment basis. The activation vector is the input vector of the last softmax operation in the CNN structure. Through the activation vector, it can be judged whether the input speech belongs to an instruction word or an unknown input, thus solving the open-set recognition problem. The duration of each instruction word speech is 1 second, and the sampling frequency of the speech data is 44100 Hz.
[0060] The model structure for processing the open-set recognition problem using the activation vector is as Figure 5 shown. After the training set trains the model, the training set is substituted into the trained model to calculate the activation vector of each class. Then, for each class, the GMM corresponding to the activation vector of that class is fitted. Here, it is equivalent to using the GMM to fit the probability density function of the activation vector distribution of each class. The overall process is as Figure 6 shown. First, calculate that the probability of the input belonging to class k in the CNN model is the largest, and then use the activation vector calculated by the input in the model to judge whether the input is an unknown class. The basis here is that the activation vector mapped by the unknown class input after passing through the CNN model is far from the activation vectors of the known classes. The specific method is to substitute the calculated activation vector into the GMM corresponding to class k and calculate the result. There is a hyperparameter that needs to be set, that is, the Figure 6 threshold shown in. When the function value calculated by substituting the activation vector into the GMM is greater than the threshold, the input is judged to be class k, otherwise it is judged to be an unknown class.
[0061] The technical effect of the present invention is verified by experiments.
[0062] The dataset used in the experiment is divided into a training set and a test set. The training set consists entirely of noisy command word voice samples. There are three types of noisy command word voice samples, and their corresponding command words are "forward", "backward", and "left turn". There are 30 noisy command word voice samples for each type, a total of 90: The test set also has three types of noisy command word voice samples, namely "forward", "backward", and "left turn". There are 6 voice samples for each command word, a total of 18. In addition, there are 50 test samples that are not the above command words. After the training model is completed, the training set is re-substituted into the model to calculate the activation vectors. The set of activation vectors corresponding to class k is initially an empty set. When the true label of the training data is k and the prediction after substituting into the model is also class k, the activation vector corresponding to this training data is added to the set of activation vectors corresponding to class k. In this way, three sets of activation vectors corresponding to the three classes of "forward", "backward", and "left turn" can be calculated. Then, the corresponding GMM is calculated according to the set of activation vectors corresponding to each class. After training, the test set is used to detect the training results.
[0063] Regarding the parameters used in the CNN model, the input of the CNN model is the MFCC feature map after calculating the command word voice. The dimension of the feature map is 13×99. The number of channels in the first convolutional layer is 16, and the number of channels in the second convolutional layer is 32. The size of the convolutional window in both convolutional layers is 5×5, so as to capture features in a larger range. The moving step of the convolutional window in both convolutional layers is 1, and the size of the pooling window in both convolutional layers is 2×2. The moving step of the pooling window is 1. After the operations of the two convolutional layers, the column vector obtained after splicing has a dimension of 32×11×97. The loss function for the entire training is the cross-entropy loss function.
[0064] The experimental results are as Figure 7 shown. Compared with other methods, such as the threshold method and the unknown class method, the method based on activation vectors in the present invention has a recognition accuracy rate of 0.94 for command words and a rejection rate of 0.98 for non-command words; for the threshold method, the recognition accuracy rate of command words is 0.7, and the rejection rate of non-command words is 0.42; for the unknown class method, the recognition accuracy rates of both command words and non-command words are 0.96. Although in open-set recognition, the unknown class method also has excellent recognition rates, this method requires a large number of unknown commands as "unknown classes" for training, resulting in a large increase in the number of model parameters and a long training time for the model used. Therefore, the method proposed in the present invention has an ideal effect, a high recognition accuracy rate, and can effectively realize the recognition of command words and the rejection of non-command words.
[0065] The present invention proposes an online command word speech recognition method in a noisy environment based on CNN, which can effectively solve the problems of the accuracy of command word speech recognition and open set recognition in a noisy environment. It has the following characteristics: the device used is an intelligent mining robot; the noise to be processed is various noises during the operation of the intelligent mining robot; the command words used are "forward", "backward", "turn left", "turn right" and "stop"; the problems to be solved are command word speech recognition and open set recognition in a noisy environment.
[0066] Another embodiment of the present invention provides an online command word speech recognition system in a noisy environment, as Figure 8 shown. The system includes:
[0067] A preprocessing module 110, which is used to preprocess the real-time input unknown sound signal. The preprocessing includes high-frequency compensation and frame segmentation processing. Among them, the frame segmentation processing is to divide the unknown sound signal into multiple short-time speech segments, and each speech segment is called a frame signal;
[0068] A feature extraction module 120, which is used to extract features from the preprocessed unknown sound signal to obtain MFCC feature vectors;
[0069] A binary classification module 130, which is used to input the MFCC feature vectors into a trained CNN binary classification network model to identify whether the unknown sound signal is speech or noise;
[0070] A speech recognition module 140, which is used to splice the MFCC feature vectors corresponding to the speech when the binary classification module identifies the unknown sound signal as speech, calculate to obtain an MFCC feature map, and input the MFCC feature map into a trained CNN multi-classification network model to recognize the speech and obtain the recognition result of the speech.
[0071] Among them, high-frequency compensation is realized by using a first-order FIR high-pass filter in the preprocessing module 110. The specific process of identifying whether the unknown sound signal is speech or noise in the binary classification module 130 includes: establishing a data buffer area for storing MFCC feature vectors corresponding to multiple frames of signals; each time the MFCC feature vector corresponding to a new frame of signal enters the data buffer area, the MFCC feature vector corresponding to the first frame of signal that entered the buffer area is discarded; splicing the MFCC feature vectors corresponding to multiple frames of signals stored in the data buffer area, calculating to obtain an MFCC feature map; inputting the MFCC feature map into a trained CNN binary classification network model to identify whether multiple frames of signals are speech or noise. If it is speech, continue to process the MFCC feature vectors corresponding to the real-time input preprocessed unknown sound signal according to the above steps; if it is noise, continue to wait.
[0072] Among them, during the training process of the CNN multi-classification network model in the speech recognition module 140, an activation vector is introduced as the input vector of the softmax operation in the last layer of the CNN multi-classification network model structure. The training process includes: after obtaining the CNN multi-classification network model according to the training set data, inputting the training set data into the model to calculate the activation vector of each category, and for each category, fitting the corresponding GMM function using the activation vector of this category.
[0073] The specific process of inputting the MFCC feature map into the trained CNN multi-classification network model in the speech recognition module 140 to recognize the speech and obtain the recognition result of the speech includes: inputting the MFCC feature map into the trained CNN multi-classification network model, first calculating the known category K with the highest probability to which the MFCC feature map belongs, then calculating the activation vector corresponding to the MFCC feature map, and then substituting the activation vector into the GMM function fitted during the training process of the known category K for calculation to determine whether the speech corresponding to the MFCC feature map belongs to an unknown category or the known category K. When the calculated function value is greater than the preset hyperparameter threshold, it is determined that the speech corresponding to the MFCC feature map belongs to the above-mentioned known category K, otherwise it belongs to an unknown category.
[0074] The functions of the online command word speech recognition system in a noise environment described in this embodiment can be illustrated by the aforementioned online command word speech recognition method in a noise environment. Therefore, for the parts not described in detail in this embodiment, reference can be made to the above method embodiments, which will not be elaborated here.
[0075] Although the present invention has been described based on a limited number of embodiments, those skilled in the art in this technical field understand that other embodiments can be envisioned within the scope of the present invention thus described. Regarding the scope of the present invention, the disclosure of the present invention is illustrative rather than restrictive, and the scope of the present invention is defined by the appended claims.
Claims
1. An online command word speech recognition method in a noisy environment, characterized in that, It includes the following steps: Step 1: Preprocess the real-time input unknown sound signal, where the preprocessing includes high-frequency compensation and frame segmentation; Step 2: Extract features from the preprocessed unknown sound signal to obtain MFCC feature vectors; Step 3: Input the MFCC feature vectors into the trained CNN binary classification network model to identify whether the unknown sound signal is speech or noise; Step 4: When it is identified that the unknown sound signal is speech, splice the MFCC feature vectors corresponding to the speech and calculate to obtain an MFCC feature map; Input the MFCC feature map into the trained CNN multi-classification network model to identify the speech and obtain the recognition result of the speech; where, during the training process of the CNN multi-classification network model, an activation vector is introduced as the input vector for the softmax operation in the last layer of the CNN multi-classification network model structure, and the training process includes: after training the CNN multi-classification network model according to the training set data, input the training set data into the model to calculate the activation vectors of each category, and for each category, use the activation vector of this category to fit the corresponding GMM function; The specific process of inputting the MFCC feature map into the trained CNN multi-classification network model to identify the speech and obtain the recognition result of the speech includes: input the MFCC feature map into the trained CNN multi-classification network model, first calculate the known category K with the highest probability to which the MFCC feature map belongs, then calculate the activation vector corresponding to the MFCC feature map, and then substitute the activation vector into the GMM function fitted during the training process of the known category K for calculation to determine whether the speech corresponding to the MFCC feature map belongs to an unknown category or the known category K. When the calculated function value is greater than the preset hyperparameter threshold, it is determined that the speech corresponding to the MFCC feature map belongs to the above-known category K, otherwise it belongs to an unknown category.
2. The method for online command word speech recognition in a noisy environment according to claim 1, wherein In Step 1, the high-frequency compensation is implemented using a first-order FIR high-pass filter; the frame segmentation is to divide the unknown sound signal into multiple short-time speech segments, and each speech segment is called a frame signal.
3. The method for online command word speech recognition in a noisy environment according to claim 2, wherein, The specific process of Step 3 includes: Establish a data buffer for storing MFCC feature vectors corresponding to multiple frames of signals; Each time the MFCC feature vector corresponding to a new frame of signal enters the data buffer, the MFCC feature vector corresponding to the first frame of signal that entered is discarded; Splice the MFCC feature vectors corresponding to the multiple frames of signals stored in the data buffer and calculate to obtain an MFCC feature map; Input the MFCC feature map into the trained CNN binary classification network model to identify whether the multiple frames of signals are speech or noise. If it is speech, continue to process the MFCC feature vectors corresponding to the real-time input preprocessed unknown sound signal according to the above steps; if it is noise, continue to wait.
4. An online command word speech recognition system in a noisy environment, characterized in that, It includes: A preprocessing module for preprocessing a real-time input unknown sound signal, where the preprocessing includes high-frequency compensation and frame segmentation; wherein, the frame segmentation is to divide the unknown sound signal into multiple short-time speech segments, and each speech segment is called a frame signal; A feature extraction module for extracting features from the preprocessed unknown sound signal to obtain MFCC feature vectors; A binary classification module for inputting the MFCC feature vectors into a trained CNN binary classification network model to identify whether the unknown sound signal is speech or noise; A speech recognition module, when the binary classification module identifies that the unknown sound signal is speech, splicing the MFCC feature vectors corresponding to the speech, calculating to obtain an MFCC feature map; inputting the MFCC feature map into a trained CNN multi-classification network model to recognize the speech and obtain the recognition result of the speech; wherein, during the training process of the CNN multi-classification network model, an activation vector is introduced as the input vector of the softmax operation in the last layer of the CNN multi-classification network model structure, and the training process includes: after training the CNN multi-classification network model according to the training set data, inputting the training set data into the model to calculate the activation vectors of each category, and for each category, fitting the corresponding GMM function using the activation vector of that category; The specific process of inputting the MFCC feature map into a trained CNN multi-classification network model to recognize the speech and obtain the recognition result of the speech includes: inputting the MFCC feature map into a trained CNN multi-classification network model, first calculating the known category K with the highest probability to which the MFCC feature map belongs, then calculating the activation vector corresponding to the MFCC feature map, and then substituting the activation vector into the GMM function fitted during the training process of the known category K for calculation to determine whether the speech corresponding to the MFCC feature map belongs to an unknown category or the known category K. When the calculated function value is greater than a preset hyperparameter threshold, it is determined that the speech corresponding to the MFCC feature map belongs to the above-known category K, otherwise it belongs to an unknown category.
5. An online command word speech recognition system in a noisy environment according to claim 4, characterized in that, In the preprocessing module, a first-order FIR high-pass filter is used to achieve high-frequency compensation.
6. The online command word speech recognition system under a noisy environment according to claim 5, characterized in that, The specific process of identifying whether the unknown sound signal is speech or noise in the binary classification module includes: Establishing a data buffer for storing MFCC feature vectors corresponding to multiple frames of signals; Each time the MFCC feature vector corresponding to a new frame of signal enters the data buffer, the MFCC feature vector corresponding to the first frame of signal that entered is discarded; Splicing the MFCC feature vectors corresponding to multiple frames of signals stored in the data buffer, and calculating to obtain an MFCC feature map; Inputting the MFCC feature map into a trained CNN binary classification network model to identify whether the multiple frames of signals are speech or noise. If it is speech, continue to process the MFCC feature vectors corresponding to the real-time input preprocessed unknown sound signal according to the above steps; if it is noise, continue to wait.
Citation Information
Patent Citations
Speech activity detection method, speech recognition method and system
CN110808073A
Counterfeit voice recognition method and device and computer readable storage medium
CN111933154A