A speech recognition method and system
By introducing deep learning technology into the speech recognition system, extracting the spectrum characteristics, timing information and contextual connections of speech signals, the shortcomings of traditional methods in complex speech environments and unknown vocabulary processing are solved, and higher recognition accuracy and generalization capabilities are achieved.
Patent Information
- Application Number
- CN202510271777.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-03-10
AI Technical Summary
Traditional speech recognition methods perform poorly when dealing with unknown vocabulary or variable and complex speech environments, and there are limitations in speech feature extraction and model training, resulting in limited recognition accuracy and generalization capabilities.
By introducing feature extraction modules, speech modules, language modules and decoding modules into the speech recognition system, deep learning technology is used to extract spectrum features, timing information and contextual connections from speech signals, so as to map speech feature representations and generate language texts.
It improves the accuracy and generalization ability of speech recognition, can handle complex speech environments and unknown vocabulary more effectively, and enhances the robustness and recognition accuracy of the model.
Smart Images

Figure CN119785774B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of speech recognition, and particularly to a speech recognition method and system. Background Art
[0002] Traditional speech recognition methods often struggle to achieve ideal recognition effects in the face of diverse and complex speech environments (such as telephone communication, smart home, intelligent vehicle, etc.) due to environmental noise, which limits the wide application of speech recognition technology. Or, due to user habits (non-standard Mandarin, dialects, and daily communication with obvious oral characteristics), there are unknown words, which will affect the accuracy of speech recognition and often fail to achieve ideal recognition effects, thus restricting the wide application of speech recognition technology.
[0003] In the process of implementing the present invention, the applicant found that there are at least the following problems in the prior art:
[0004] In traditional speech recognition methods, the system relies on predefined rules and dictionaries for matching, and this method performs poorly when dealing with unknown words or diverse and complex speech environments. At the same time, traditional methods also have limitations in the extraction of speech features and the training of models, resulting in limited recognition accuracy and generalization ability. Summary of the Invention
[0005] Embodiments of the present invention provide a speech recognition method and system, which can solve the technical problem that traditional methods in the prior art also have limitations in the extraction of speech features and the training of models, resulting in limited recognition accuracy and generalization ability.
[0006] To achieve the above object, in a first aspect, embodiments of the present invention provide a speech recognition method, including:
[0007] Step 11: For each preprocessed sample speech, respectively extract a corresponding sample speech feature vector from the preprocessed sample speech through a to-be-trained feature extraction module. The sample speech feature vector retains the speech features of the speech producer, and the sample speech feature vector includes the spectral features of the speech, the timing information of the speech, and the context connection of the speech;
[0008] Step 12: Extract a feature representation and the timing dependence relationship of the speech from the sample speech feature vector through a to-be-trained speech module. The feature representation includes the mel cepstral coefficients of the speech, the timing features of the speech, and the timing dependence relationship of the speech; map the spectral features of the speech to corresponding phoneme labels or word labels in sequence according to the feature representation to obtain a mapping representation;
[0009] Step 13: Input the mapping representation and the temporal dependency relationship of the speech into the language module to be trained. Through the language module to be trained, according to the temporal dependency relationship of the speech, phonemes or words are combined and sorted according to grammar rules and semantic relationships to form multiple probability factor sequences;
[0010] Step 14: Input the temporal dependency relationship of the speech and the multiple probability factor sequences into the decoding module to be trained. For each probability factor sequence, judge the matching degree between the probability factor sequence and the sample speech signal according to the temporal dependency relationship of the speech through the dynamic programming algorithm in the decoding module to be trained, and take the probability factor sequence with the highest matching degree as the language text of the sample speech signal and output it; the sample speech signal is the original speech of the preprocessed sample speech; the sample speech signal refers to the speech stream generated by the speech producer using any language that can express meaning;
[0011] Step 15: According to the matching degree between the output language text and the sample speech signal, adjust the parameters of the modules in each step from Step 11 to Step 14, and continue training with the preprocessed sample speech until the output language text conforms to the labeled text corresponding to the sample speech signal, and obtain the trained speech recognition model; the speech recognition model is used to receive the speech signal to be recognized and output the corresponding language text.
[0012] In a second aspect, an embodiment of the present invention provides a speech recognition system, including:
[0013] A feature extraction module training unit, configured to, for each preprocessed sample speech, respectively extract the corresponding sample speech feature vector from the preprocessed sample speech through the feature extraction module to be trained. The sample speech feature vector retains the speech features of the speech producer, and the sample speech feature vector includes the spectral features of the speech, the temporal information of the speech, and the context connection of the speech;
[0014] A speech module training unit, configured to extract a feature representation and a temporal dependency relationship of the speech from the sample speech feature vector through the speech module to be trained. The feature representation includes the mel cepstral coefficients of the speech, the temporal features of the speech, and the temporal dependency relationship of the speech; according to the feature representation, map the spectral features of the speech to the corresponding phoneme labels or word labels in sequence to obtain a mapping representation;
[0015] A language module training unit, configured to input the mapping representation and the temporal dependency relationship of the speech into the language module to be trained. Through the language module to be trained, according to the temporal dependency relationship of the speech, phonemes or words are combined and sorted according to grammar rules and semantic relationships to form multiple probability factor sequences;
[0016] A decoding module training unit, configured to input the temporal dependency relationship of speech and multiple said probability factor sequences into a decoding module to be trained. For each probability factor sequence, determine the matching degree between the probability factor sequence and the sample speech signal according to the temporal dependency relationship of speech through a dynamic programming algorithm in the decoding module to be trained, and use the probability factor sequence with the highest matching degree as the language text of the sample speech signal and output it; the preprocessed sample speech is obtained by preprocessing the sample speech signal;
[0017] Wherein, according to the matching degree between the output language text and the sample speech signal, adjust the parameters of each module to be trained, and use the preprocessed sample speech to sequentially pass through a feature extraction module training unit, a speech module training unit, a language module training unit, and a decoding module training unit to continue training until the output language text conforms to the labeled text corresponding to the sample speech signal, and obtain a trained speech recognition model; the speech recognition model is used to receive a speech signal to be recognized and output a corresponding language text.
[0018] In a third aspect, an embodiment of the present invention provides a computer device, including:
[0019] A processor; and a memory arranged to store computer-executable instructions, where the executable instructions, when executed, cause the processor to execute the foregoing speech recognition method.
[0020] The above technical solution has the following beneficial effects: The sample speech feature vector not only contains the spectral characteristics of the sample speech signal, but also contains the temporal information and context relationship of the sample speech signal, providing a basis for the modeling of subsequent speech modules. At the same time, the sample speech feature vector also contains user feature information, accents, dialects, etc., providing a basis for personalized training. It can automatically learn the speech features of different speakers and different speech environments, thereby learning the complex features in the speech signal, enhancing the generalization ability of the model. And in the language module to be trained, it can be mapped to the corresponding phoneme or word label, so as to achieve accurate recognition of the speech signal. It can combine and sort the recognized phonemes or words according to grammar rules and semantic relationships to form meaningful sentences or paragraphs. Thus, the recognition result can be optimized and corrected, improving the accuracy, accuracy and fluency of speech recognition. The process of decoding and predicting the probability factor sequence to generate a language text model training reduces the computational complexity, improves the training efficiency, and improves the generalization ability. The speech recognition method of the embodiment of the present invention can adapt to the speech recognition problem in a complex speech environment. Description of the Drawings
[0021] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0022] Figure 1 is a flowchart of a speech recognition method according to an embodiment of the present invention;
[0023] Figure 2 is a logical structure diagram of a speech recognition system according to an embodiment of the present invention;
[0024] Figure 3 is a flowchart of the application of a speech recognition model according to an embodiment of the present invention. Detailed implementation manners
[0025] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0026] As Figure 1 shown, in combination with the embodiments of the present invention, a speech recognition method is provided, including:
[0027] Step 11: For each preprocessed sample speech, respectively extract the corresponding sample speech feature vector from the preprocessed sample speech through the to-be-trained feature extraction module. The sample speech feature vector retains the speech features of the speech producer, and the sample speech feature vector includes the spectral features of the speech, the timing information of the speech, and the context connection of the speech;
[0028] Step 12: Extract the feature representation and the timing dependence relationship of the speech from the sample speech feature vector through the to-be-trained speech module. The feature representation includes the mel cepstral coefficients of the speech, the timing features of the speech, and the timing dependence relationship of the speech; map the spectral features of the speech to the corresponding phoneme labels or word labels in sequence according to the feature representation to obtain a mapping representation;
[0029] Step 13: Input the mapping representation and the timing dependence relationship of the speech into the to-be-trained language module. Through the to-be-trained language module, combine and sort the phonemes or words according to the grammar rules and semantic relationships according to the timing dependence relationship of the speech to form a plurality of probability factor sequences;
[0030] Step 14: Input the temporal dependency relationship of the speech and the multiple probability factor sequences into the decoding module to be trained. For each probability factor sequence, use the dynamic programming algorithm within the decoding module to be trained to determine the matching degree between the probability factor sequence and the sample speech signal according to the temporal dependency relationship of the speech. Output the probability factor sequence with the highest matching degree as the language text of the sample speech signal; the sample speech signal is the original speech of the preprocessed sample speech; the sample speech signal refers to the speech stream generated by the speech producer using any language that can express meaning.
[0031] Step 15: According to the matching degree between the output language text and the sample speech signal, adjust the parameters of the modules in each of Steps 11 to 14, and continue training using the preprocessed sample speech until the output language text conforms to the labeled text corresponding to the sample speech signal, obtaining a trained speech recognition model; the speech recognition model is used to receive the speech signal to be recognized and output the corresponding language text.
[0032] The feature vector not only contains the spectral characteristics of the sample speech signal, but also contains the temporal information and context relationship of the sample speech signal, providing a basis for the subsequent modeling of the speech module. At the same time, the feature vector also contains user feature information, accents, dialects, etc., providing a basis for personalized training.
[0033] It can automatically learn the speech features of different speakers and different speech environments, thereby learning the complex features in the speech signal and enhancing the generalization ability of the model. And in the language module to be trained, it can map them to the corresponding phoneme or word labels, thereby achieving accurate recognition of the speech signal.
[0034] It can combine and sort the recognized phonemes or words according to grammar rules and semantic relationships to form meaningful sentences or paragraphs. Thereby realizing the optimization and correction of the recognition results, and improving the accuracy, accuracy and fluency of speech recognition.
[0035] The process of decoding and predicting the probability factor sequence to generate the language text model training reduces the computational complexity, improves the training efficiency, and improves the generalization ability.
[0036] The speech recognition method of the embodiment of the present invention can adapt to the speech recognition problem in a complex speech environment. In practical applications, speech signals are often interfered by various factors, such as background noise, echo, multi-speaker interference, etc., and these factors will seriously affect the accuracy of speech recognition.
[0037] Preferably, Step 11 includes:
[0038] Extract spectral features from the preprocessed sample speech through a hybrid Mel Frequency Cepstral Coefficients feature extraction algorithm;
[0039] Extract the temporal information of the speech from the preprocessed sample speech through a first deep learning model. The temporal information of the speech refers to the law of how the previous phoneme evolves into the next phoneme over time. The first deep learning model is a Convolutional Neural Network (CNN) with an attention mechanism; the temporal information refers to the characteristics of the speech signal changing over time. For example, in a continuous speech recognition task, understanding how one phoneme evolves into another over time can help the speech recognition model better predict the next possible sound.
[0040] Extract the context connection of the speech (context information, especially long-distance context information) from the preprocessed sample speech through a second deep learning model. The context connection of the speech refers to the connection between the current speech frame and its adjacent speech frames. The second deep learning model is a Long Short-Term Memory network (LSTM);
[0041] Form a sample speech feature vector corresponding to the sample speech signal through the spectral features, the temporal information of the speech, and the context connection of the speech. The context relationship refers to the connection between the current speech frame (phoneme) and its adjacent frames (phonemes) before and after. Since speech is continuous streaming media data, the current speech feature (phoneme) is often affected by the speech features (phonemes) before and after it.
[0042] Extract a feature vector containing rich speech information from the preprocessed sample speech by means of a hybrid Mel Frequency Cepstral Coefficients (MFCC) feature extraction algorithm and a first deep learning model (Convolutional Neural Network CNN and Long Short-Term Memory network LSTM), so as to be able to automatically learn rich feature vectors from the original sample speech signal. Among them, the CNN is responsible for extracting static features. Static features are local features of the speech signal, such as spectral features. The LSTM is used to extract dynamic features. Dynamic features refer to features that change over time, temporal information (temporal features), that is, the changes of the speech signal over time. When learning, the static features and dynamics will be mixed for learning. The feature vector not only contains the spectral characteristics of the sample speech signal (the spectral characteristics are the manifestations of the speech signal in the frequency domain. The spectral diagram shows the energy distribution of the speech signal at different frequencies and can reflect information such as the pitch and timbre of the speech), but also contains the temporal information and context relationship of the sample speech signal, providing a basis for the subsequent modeling of the speech module. At the same time, the feature vector also contains user feature information and accents and dialects, etc., providing a basis for personalized training and being able to process complex speech signals (preprocessing and feature extraction) more effectively.
[0043] Preferably, step 12 includes:
[0044] Constructing a speech module to be trained, the speech module to be trained including a third deep learning model and a fourth deep learning model;
[0045] Extracting a feature representation from the sample speech feature vectors by the third deep learning model, the feature representation including Mel cepstral coefficients of the speech and temporal features of the speech, the temporal features of the speech representing the frequency features of the speech varying with time, and mapping the spectral features of the speech to corresponding phoneme labels or word labels in sequence according to the feature representation, the third deep learning model being a convolutional neural network CNN with an attention mechanism;
[0046] Capturing the temporal dependence relationship of the speech in the sample speech feature vectors by the fourth deep learning model, the fourth deep learning model being a long short-term memory network LSTM.
[0047] The deep learning model includes: speech modeling (acoustic modeling) and language modeling. Speech modeling (acoustic modeling) adopts a deep neural network structure. By introducing a convolutional neural network (CNN) with an attention mechanism and a long short-term memory network (LSTM), etc., a speech module is constructed. The CNN network has a powerful feature learning ability and can extract richer feature representations from the input feature vectors: spectral features (mainly MFCC, the abbreviation of Mel Frequency Cepstral Coefficients), temporal features (time-frequency features), and map the feature representations to corresponding phoneme or word labels, that is, convert the phonemes or words corresponding to the feature representations into their respective corresponding labels; while the LSTM network can capture the temporal dependence relationship in the speech signal, is particularly effective for processing continuous speech signals, can learn the probability distributions of various speech features, and further reduce the probabilities of misrecognition and missed recognition.
[0048] During training, the feature representation and the temporal relationship information are used for the text corresponding to the segment of speech. After training, the feature representation and the temporal relationship information are used as model inputs to predict the corresponding text. For example, a voice is "The weather is really nice today". During training, the information features obtained for this speech are ABCDEF (not necessarily one-to-one). When the speech module is trained and can perform recognition, the ABCDEF features correspond to "The weather is really nice today"; if the speech signal to be recognized is "The weather is really nice tomorrow" with a bit of noise; the extracted features should be zBCDEF because the feature representation of "tomorrow" in the speech recognition model is zB, and finally the language text "The weather is really nice tomorrow" will be extracted.
[0049] By performing speech modeling on the feature vectors extracted by the feature extraction module to be trained, accurate classification of speech signals is achieved. Classification is to map each time segment (usually a short-time frame) in the speech signal to a specific category. For example, in Chinese, "b", "p", "m", etc. Classification is to convert continuous speech signals into discrete language symbols (phonemes) so that the computer can understand and process human language. Through classification, the system can gradually build a complete speech transcription result and finally achieve accurate recognition of speech content.
[0050] These deep learning models can automatically learn the speech features in different speakers and different speech environments through a large number of sample language signals, thus learning the complex features in the speech signal. For example, the timbre and pitch are different, and the complexity of the language itself enhances the generalization ability of the model. And in the language module to be trained, it can map them to the corresponding phoneme or word labels, thus achieving accurate recognition of speech signals. In the process of deep learning, words are not directly processed as information, but in a digital way, similar to the form of labels and characters.
[0051] Preferably, step 13 includes:
[0052] Construct a language module to be trained. The language module to be trained includes an N-gram model (N-gram model) and a fifth deep learning model, and the fifth deep learning model is a Transformer model;
[0053] Based on the context grammar rules by the N-gram, according to the temporal dependency relationship of the speech with a grammar strength lower than the preset value in the temporal dependency relationship of the speech, the corresponding phonemes or words are combined and sorted according to the grammar rules and semantic relationships. Through the self-attention mechanism of the fifth deep learning model, the temporal dependency relationship of the speech with a grammar strength not lower than the preset value in the temporal dependency relationship of the speech is captured, and the corresponding phonemes or words are combined and sorted according to the grammar rules and semantic relationships to form multiple probability factor sequences. Each probability factor sequence includes one of the following: meaningful words, meaningful paragraphs, and meaningful sentences.
[0054] In language modeling, the output of the speech module to be trained is used as the input of the language model to be trained. The language model to be trained adopts the traditional N-gram model and the more advanced deep learning model Transformer. N-gram is a language model based on probability statistics. Through probability distribution, it captures basic local context grammar rules. It is suitable for simple and short grammars, captures simple grammars, has low computational complexity and fast speed. The Transformer model can capture the grammar rules and semantic relationships of the global context in a sentence through the self-attention mechanism. It has strong grammar understanding ability and is suitable for capturing complex grammars, but has high computational complexity and slow speed. By combining the two and taking the best results of both, the overall accuracy and fluency are improved.
[0055] Since both of these language models can learn the grammar rules and semantic relationships of the language, they can combine and sort the recognized phonemes or words according to the grammar rules and semantic relationships to form meaningful sentences or paragraphs. Thus, the recognition results are optimized and corrected, and the accuracy of speech recognition and the fluency of the language are improved. Because there are multiple possibilities (noise, polyphonic characters, accents, etc.) between the speech signal and the text, the speech module to be trained outputs a probability sequence and there are multiple such sequences.
[0056] By simultaneously learning the static features and dynamic features of the speech signal, the accuracy of speech recognition is improved. By introducing technologies such as the attention mechanism, the robustness and recognition accuracy of the speech module to be trained are further improved.
[0057] Preferably, step 14 includes:
[0058] For each probability factor sequence, a dynamic programming decoding method based on the Viterbi algorithm is used to judge the matching degree between the probability factor sequence and the sample speech signal according to the temporal dependence relationship of the speech, and the probability factor sequence with the highest matching degree is used as the matching probability factor sequence of the sample speech signal;
[0059] A pre-trained language model is used to capture the grammar and semantic structures in the matching probability factor sequence to form a language text.
[0060] The language module can capture the probability factor sequence that best conforms to the language rules (grammar, semantic structure) and context rules, providing additional constraint conditions for the dynamic programming decoding method based on the Viterbi algorithm to help it select a text sequence that better conforms to the language rules. The dynamic programming decoding method based on the Viterbi algorithm will comprehensively consider the probabilities of the speech module and the language module, find the probability factor sequence with the largest joint probability, and use the result with the highest probability to ensure the accuracy of recognition.
[0061] The pre-trained language model is a model that has been trained in the relevant field and used as a pre-trained model, which can further accelerate the training process of the speech recognition model and improve the performance of the speech recognition model. The pre-trained language model can be existing pre-trained speech modules (such as HuBert) and language modules (such as Bert, GPT). By training and fine-tuning the data model on the training dataset, the purpose of quickly generating a new model can be achieved.
[0062] Preferably, the speech recognition method further includes:
[0063] Step 16: Before step 12, for each of the sample speech signals, the sample speech signal stream is sequentially subjected to noise reduction processing, echo removal processing, and pre-emphasis processing through a preprocessing module to be trained, and a pre-processed speech signal is obtained; the pre-emphasis processing is to reduce the influence of low-frequency noise and make the spectrum of the sample speech signal flatter.
[0064] Endpoint detection is performed on the pre-processed speech signal to detect the start position and end position of the pre-processed speech signal; the pre-processed speech signal is subjected to speech pre-emphasis processing from the start position to the end position to obtain the corresponding pre-processed sample speech, so as to improve the quality of the speech signal and provide a sufficiently clear speech signal for the subsequent process.
[0065] Through preprocessing, complex speech signals can be processed more effectively.
[0066] Preferably, the speech recognition method further includes:
[0067] Step 17: Obtain multiple original speech signals, transform and expand each original speech signal to obtain new speech signals. The ways of transformation and expansion include at least one of the following: adding noise, changing the playback speed, changing the pitch of the speech signal, adding echo, and perturbing the filter characteristics of the speech signal. The sample speech signal refers to the speech stream generated by a speech producer using any language that can express meaning. The data augmentation technology is used to transform and expand the original speech data to generate more training data to improve the generalization ability of the model.
[0068] Both the original speech signal and the new speech signal are used as sample speech signals to construct a sample speech signal set; the scale of the sample speech signal set is relatively large to ensure that the speech recognition model can learn sufficiently rich speech features.
[0069] Among them, step 17 is executed before step 16.
[0070] Preferably, in step 17 of the speech recognition method, it further includes:
[0071] For application scenarios with user feature requirements, the original voice signal is formed by obtaining the natural language recording of the text read by the user for a certain duration, so as to meet the requirements of special application scenarios. The accuracy and adaptability of speech recognition are comprehensively improved to meet more extensive and complex application requirements.
[0072] In summary, the speech recognition method of the embodiment of the present invention can adapt to the speech recognition problem in a complex speech environment. In practical applications, speech signals are often interfered by various factors, such as background noise, echo, multi-speaker interference, etc. These factors will seriously affect the accuracy of speech recognition. By introducing deep learning technology, this embodiment can automatically learn the deep features in the speech signal, effectively distinguish the target speech from the background noise, and thus significantly improve the recognition effect in a complex environment. The realization of this goal will greatly promote the application of speech recognition technology in complex environments such as noisy environments and outdoor scenes, bringing a more convenient and efficient interaction experience to users. Among them, the deep features refer to: phoneme-level features, which are the smallest pronunciation units in the speech signal; word-level features, which represent how these phonemes are combined into words and capture the grammatical and semantic relationships between words; sentence-level features, which can understand the semantics of the entire sentence.
[0073] Because in practical applications, speech signals are often interfered by various environmental noises, such as background noise, wind noise, echo, etc. The pronunciation habits, speech rates, and volumes of different users will also affect the speech recognition effect. Through the adaptive ability of deep learning, these changes can be learned during the training process, and the model parameters can be adjusted to adapt to different speech environments, maintaining a high recognition accuracy and enhancing the robustness of speech recognition.
[0074] By retaining user feature information during preprocessing, such as personal accents, dialect features, etc., and including this part of feature information during feature extraction, personalized training of the speech recognition model is realized. In this way, the model can not only better understand the user's speech commands, but also make intelligent adjustments according to the user's habits and needs, improving the recognition accuracy and user experience. Thus, the personalization level of speech recognition is improved, and the unique needs of different users can be met. Therefore, it can promote the popularization of speech recognition technology in personalized application scenarios such as smart homes and intelligent vehicles, bringing more considerate and personalized services to users.
[0075] Because it can handle complex language structures, pronunciation rules, and vocabulary changes, accurate recognition of multiple languages and dialects can be achieved. For example, in scenarios such as multinational enterprises, international tourism, or cultural exchanges, users communicate in different languages or dialects. The multi-language and multi-dialect recognition ability of this embodiment enables the system to automatically adapt to and understand the user's speech commands, thus providing a more personalized and intelligent service experience.
[0076] By optimizing the network structure and training strategy of the deep learning model, fast processing and accurate recognition of continuous speech signals are achieved. Meanwhile, an efficient decoding algorithm further improves the real-time performance and efficiency of speech recognition, meeting the requirements of efficient interaction. It can promote the development of speech recognition technology in high-efficiency application scenarios such as real-time interaction and online recognition, bringing a smoother and more efficient interaction experience to users.
[0077] Through deep learning methods, fine feature extraction and complex modeling processes are carried out on speech signals. The deep neural network can learn high-level features in speech signals, which are crucial for distinguishing different speech contents, improving the accuracy of speech recognition. For application scenarios that require high-precision speech recognition, such as smart home control and in-vehicle voice assistants, it can bring fewer misoperations, higher user satisfaction, and more reliable system performance.
[0078] A lightweight network structure (LSTM and Transformer can be set relatively small) and an efficient computing algorithm are adopted, reducing the occupation of computing resources and storage space. In the inference stage, it can quickly process the input speech signal and output the recognition result, improving the real-time performance and response speed of speech recognition. This characteristic of low resource consumption and high efficiency gives this patent outstanding advantages in resource-constrained environments (such as mobile devices, embedded systems, etc.).
[0079] By dividing and encoding the feature sequence of the sample speech signal, a feature representation (target speech feature block) is obtained, multiple probability factor sequences are generated, and the probability factor sequences are decoded and predicted to generate language text. This processing method simplifies the model training process, reduces the computational complexity, and improves the training efficiency. At the same time, the model can be iteratively optimized according to the feedback of the predicted text sequence. This processing method enables the speech recognition model to no longer rely on whole-sentence input, but can process streaming speech data in real time, and is effectively applied to scenarios that require instant feedback such as real-time communication and online speech transcription. Enterprises and developers can use the tools and frameworks provided in this application to quickly build and deploy speech recognition systems.
[0080] Preferably, the speech recognition method further includes:
[0081] Step 21: Preprocess the speech signal to be recognized through a preprocessing module to obtain the corresponding speech to be recognized; wherein, the speech signal to be recognized refers to the speech stream generated by a speech producer using any language that can express meaning.
[0082] Step 22: Extract the speech feature vector to be recognized from the speech to be recognized through the feature extraction module. The speech feature vector to be recognized includes the spectral feature of the speech, the timing information of the speech, and the context connection of the speech. The spectral feature retains the speech feature of the speech generator.
[0083] Step 23: Extract the feature representation and the timing dependence relationship of the speech from the speech feature vector to be recognized through the speech module. The feature representation includes the mel cepstral coefficients of the speech, the timing feature of the speech, and the timing dependence relationship of the speech. Map the spectral feature of the speech to the corresponding phoneme label or word label in sequence according to the feature representation to obtain a mapping representation.
[0084] Step 24: Input the mapping representation and the timing dependence relationship of the speech into the language module. Through the language module, according to the timing dependence relationship of the speech, combine and sort the phonemes or words according to the grammar rules and semantic relationships to form multiple probability factor sequences.
[0085] Step 25: Input the timing dependence relationship of the speech and the multiple probability factor sequences into the module to be decoded. For each probability factor sequence, judge the matching degree between the probability factor sequence and the speech signal to be recognized according to the timing dependence relationship of the speech through the dynamic programming algorithm in the decoding module, and output the probability factor sequence with the highest matching degree as the language text of the speech signal to be recognized.
[0086] As Figure 2 shown, in combination with the embodiments of the present invention, a speech recognition system is provided, including:
[0087] The training unit 301 of the feature extraction module is used to extract the corresponding sample speech feature vector from each preprocessed sample speech through the feature extraction module to be trained. The sample speech feature vector retains the speech feature of the speech generator. The sample speech feature vector includes the spectral feature of the speech, the timing information of the speech, and the context connection of the speech.
[0088] The training unit 302 of the speech module is used to extract the feature representation and the timing dependence relationship of the speech from the sample speech feature vector through the speech module to be trained. The feature representation includes the mel cepstral coefficients of the speech, the timing feature of the speech, and the timing dependence relationship of the speech. Map the spectral feature of the speech to the corresponding phoneme label or word label in sequence according to the feature representation to obtain a mapping representation.
[0089] The training unit 303 of the language module is used to input the mapping representation and the timing dependence relationship of the speech into the language module to be trained. Through the language module to be trained, according to the timing dependence relationship of the speech, combine and sort the phonemes or words according to the grammar rules and semantic relationships to form multiple probability factor sequences.
[0090] The decoding module training unit 304 is configured to input the temporal dependence relationship of the speech and the multiple probability factor sequences into the decoding module to be trained. For each probability factor sequence, the dynamic programming algorithm in the decoding module to be trained is used to determine the matching degree between the probability factor sequence and the sample speech signal according to the temporal dependence relationship of the speech, and the probability factor sequence with the highest matching degree is used as the language text of the sample speech signal and output; the preprocessed sample speech is obtained by preprocessing the sample speech signal.
[0091] Among them, according to the matching degree between the output language text and the sample speech signal, the parameters of each module to be trained are adjusted, and the preprocessed sample speech is used to pass through the feature extraction module training unit, the speech module training unit, the language module training unit, and the decoding module training unit in sequence to continue training until the output language text conforms to the labeled text corresponding to the sample speech signal, and a trained speech recognition model is obtained; the speech recognition model is used to receive the speech signal to be recognized and output the corresponding language text.
[0092] The feature vector not only contains the spectral characteristics of the sample speech signal, but also contains the temporal information and context relationship of the sample speech signal, providing a basis for the subsequent modeling of the speech module. At the same time, the feature vector also contains user feature information, accents, dialects, etc., providing a basis for personalized training.
[0093] It can automatically learn the speech features of different speakers and different speech environments, thereby learning the complex features in the speech signal and enhancing the generalization ability of the model. And in the language module to be trained, it can map them to the corresponding phoneme or word labels, thereby realizing the accurate recognition of the speech signal.
[0094] It can combine and sort the recognized phonemes or words according to grammar rules and semantic relationships to form meaningful sentences or paragraphs. Thereby realizing the optimization and correction of the recognition results, and improving the accuracy, accuracy and fluency of speech recognition.
[0095] The process of decoding and predicting the probability factor sequence to generate the language text model training reduces the computational complexity, improves the training efficiency, and improves the generalization ability.
[0096] The speech recognition method according to the embodiment of the present invention can adapt to the speech recognition problem in a complex speech environment. In practical applications, speech signals are often interfered by various factors, such as background noise, echo, multi-speaker interference, etc., and these factors will seriously affect the accuracy of speech recognition.
[0097] Preferably, the feature extraction module training unit 301 is specifically configured to:
[0098] Extract spectral features from the preprocessed sample speech through a hybrid Mel-frequency cepstral coefficient feature extraction algorithm;
[0099] Extract the temporal information of the speech from the preprocessed sample speech through a first deep learning model. The temporal information of the speech refers to the law of the evolution of the previous phoneme to the next phoneme over time. The first deep learning model is a convolutional neural network CNN with an attention mechanism; the temporal information refers to the characteristics of the speech signal changing over time. For example, in a continuous speech recognition task, understanding how a phoneme evolves over time to another phoneme can help the speech module better predict the next possible sound.
[0100] Extract the context connection (context information, especially long-distance context information) of the speech from the preprocessed sample speech through a second deep learning model. The context relationship of the speech refers to the connection between the current speech frame and its adjacent speech frames. The second deep learning model is a long short-term memory network LSTM;
[0101] Form a sample speech feature vector corresponding to the sample speech signal through the spectral features, the temporal information of the speech, and the context connection of the speech. The context connection refers to the connection between the current speech frame (phoneme) and its adjacent front and back frames (phonemes). Since speech is continuous streaming data, the current speech feature (phoneme) is often affected by its adjacent front and back speech features (phonemes).
[0102] Through a hybrid Mel-frequency cepstral coefficient (MFCC) feature extraction algorithm and a first deep learning model (CNN and LSTM), extract a feature vector containing rich speech information from the preprocessed sample speech, so as to be able to automatically learn rich feature vectors from the original sample speech signal. Among them, CNN is responsible for extracting static features, and static features are local features of the speech signal, such as spectral features. LSTM is responsible for extracting dynamic features, and dynamic features refer to features that change over time, temporal information (temporal features), that is, the changes of the speech signal over time. When learning, the static features and dynamics will be mixed for learning. The feature vector not only contains the spectral characteristics of the sample speech signal (the spectral characteristics are the performance of the speech signal in the frequency domain, and the energy distribution of the speech signal at different frequencies is shown through the spectrogram, which can reflect information such as the pitch and timbre of the speech), but also contains the temporal information and context relationship of the sample speech signal, providing a basis for the subsequent modeling of the speech module. At the same time, the feature vector also contains user feature information and accents and dialects, etc., providing a basis for personalized training and being able to more effectively process complex speech signals (preprocessing and feature extraction).
[0103] Preferably, the speech module training unit 302 is specifically configured to:
[0104] Construct a voice module to be trained, where the voice module to be trained includes a third deep learning model and a fourth deep learning model;
[0105] Extract a feature representation from the sample voice feature vectors through the third deep learning model. The feature representation includes the Mel cepstral coefficients of the voice and the temporal features of the voice. The temporal features of the voice represent the frequency features of the voice changing over time. Map the spectral features of the voice to the corresponding phoneme labels or word labels in sequence according to the feature representation. The third deep learning model is a convolutional neural network CNN with an attention mechanism;
[0106] Capture the temporal dependence relationship of the voice in the sample voice feature vectors through the fourth deep learning model. The fourth deep learning model is a long short-term memory network LSTM.
[0107] The deep learning model includes: voice modeling (acoustic modeling) and language modeling, which are the core parts. The voice modeling (acoustic modeling) mainly adopts an advanced deep neural network structure. By introducing a convolutional neural network (CNN) with an attention mechanism and a long short-term memory network (LSTM), etc., to construct the voice module. The CNN network has a powerful feature learning ability and can extract richer feature representations from the input feature vectors: spectral features (mainly MFCC, the abbreviation of Mel Frequency Cepstral Coefficients), temporal features (time-frequency features), and map the feature representations to the corresponding phoneme or word labels, that is, convert the phonemes or words corresponding to the feature representations into their respective corresponding labels; while the LSTM network can capture the temporal dependence relationship in the voice signal, is particularly effective for processing continuous voice signals, can learn the probability distributions of various voice features, and further reduce the probabilities of misrecognition and missed recognition.
[0108] During training, the feature representation and the temporal relationship information are used for the text corresponding to that segment of voice. After training, the feature representation and the temporal relationship information are used as model inputs to predict the corresponding text. For example, a voice is "The weather is nice today". The information features obtained during training for this voice are ABCDEF (not necessarily one-to-one). When the voice module can be recognized after training, the text corresponding to the ABCDEF features is "The weather is nice today"; if the voice signal to be recognized is "The weather is nice tomorrow" with a little noise, the extracted features should be zBCDEF, because the feature representation of "tomorrow" in the voice recognition model is zB, and finally the text "The weather is nice tomorrow" will be extracted.
[0109] By performing speech modeling on the feature vectors extracted by the feature extraction module to be trained, accurate classification of speech signals is achieved. Classification maps each time segment (usually a short-time frame) in the speech signal to a specific category. For example, in Chinese, "b", "p", "m", etc. Classification is to convert continuous speech signals into discrete language symbols (phonemes) so that the computer can understand and process human language. Through classification, the system can gradually build a complete speech transcription result and finally achieve accurate recognition of speech content.
[0110] These deep learning models can automatically learn the speech features of different speakers and different speech environments through a large number of sample language signals, thus learning the complex features in the speech signals. For example, the timbre and pitch are different, and the complexity of the language itself enhances the generalization ability of the model. And in the language module to be trained, it can map them to the corresponding phoneme or word labels, thus achieving accurate recognition of speech signals. In the process of deep learning, words are not directly processed as information, but in a digital way, similar to the form of labels and characters.
[0111] Preferably, the language module training unit 303 is specifically used for:
[0112] Construct a language module to be trained, where the language module to be trained includes an N-gram model (N-gram model) and a fifth deep learning model, and the fifth deep learning model is a Transformer model;
[0113] Based on the context grammar rules of the N-gram, according to the temporal dependence relationship of the speech whose grammar strength in the temporal dependence relationship of the speech is lower than the preset value, the corresponding phonemes or words are combined and sorted according to the grammar rules and semantic relationships. Through the self-attention mechanism of the fifth deep learning model, the temporal dependence relationship of the speech whose grammar strength in the temporal dependence relationship of the speech is not lower than the preset value is captured, and the corresponding phonemes or words are combined and sorted according to the grammar rules and semantic relationships to form multiple probability factor sequences, and each probability factor sequence includes one of the following: meaningful words, meaningful paragraphs, and meaningful sentences.
[0114] In language modeling, the output of the speech module to be trained is used as the input of the language model to be trained. The language model to be trained adopts the N-gram model and the deep learning model Transformer. N-gram is a language model based on probability statistics that captures basic local context grammar rules through probability distributions. It is suitable for simple and short grammars, capturing simple grammars, with low computational complexity and high speed. The Transformer model can capture the grammar rules and semantic relationships of the global context in a sentence through the self-attention mechanism, with strong grammar understanding ability, suitable for capturing complex grammars, but with high computational complexity and low speed. By combining the two and taking the optimal results of both, the overall accuracy and fluency can be improved.
[0115] Since both of these language models can learn the grammar rules and semantic relationships of the language, they can combine and sort the recognized phonemes or words according to the grammar rules and semantic relationships to form meaningful sentences or paragraphs. Thus, the recognition results can be optimized and corrected, improving the accuracy of speech recognition and the fluency of the language. Because there are multiple possibilities (noise, polyphonic characters, accents, etc.) between the speech signal and the text, the speech module to be trained outputs a probability sequence, and there are multiple such sequences.
[0116] By simultaneously learning the static features and dynamic features of the speech signal, the accuracy of speech recognition is improved. By introducing technologies such as the attention mechanism, the robustness and recognition accuracy of the speech module to be trained are further improved.
[0117] Preferably, the decoding module training unit 304 is specifically used for:
[0118] For each probability factor sequence, a dynamic programming decoding method based on the Viterbi algorithm is adopted to judge the matching degree between the probability factor sequence and the sample speech signal according to the temporal dependence relationship of the speech, and the probability factor sequence with the highest matching degree is used as the matching probability factor sequence of the sample speech signal;
[0119] A pre-trained language model is used to capture the grammar and semantic structures in the matching probability factor sequence to form a language text.
[0120] The language module can capture the probability factor sequence that best conforms to the language rules (grammar, semantic structure) and context rules, providing additional constraint conditions for the dynamic programming decoding method based on the Viterbi algorithm to help it select a text sequence that better conforms to the language rules. The dynamic programming decoding method based on the Viterbi algorithm will comprehensively consider the probabilities of the speech module and the language module to find the probability factor sequence with the largest joint probability, and use the result with the highest probability to ensure the accuracy of recognition.
[0121] The pre-trained language model is a model that has been trained in the relevant field and used as a pre-training model, which can further accelerate the training process of the speech recognition model and improve the performance of the speech recognition model. The pre-trained language model can be existing pre-trained speech modules (such as HuBert) and language modules (such as Bert, GPT). By training and fine-tuning the data model on the training data set, the purpose of quickly generating a new model can be achieved.
[0122] Preferably, the speech recognition system further includes:
[0123] A preprocessing unit 306, configured to perform noise reduction processing, echo removal processing, and pre-emphasis processing on each of the sample speech signals in sequence through a preprocessing module to be trained, so as to obtain a pre-processed speech signal; the pre-emphasis processing is to reduce the influence of low-frequency noise and make the spectrum of the sample speech signal flatter.
[0124] Perform endpoint detection on the pre-processed speech signal to detect the start position and end position of the pre-processed speech signal; perform speech pre-emphasis processing on the pre-processed speech signal from the start position to the end position to obtain a corresponding pre-processed sample speech, so as to improve the quality of the speech signal and provide a sufficiently clear speech signal for the subsequent process.
[0125] Through preprocessing, complex speech signals can be processed more effectively.
[0126] Preferably, the speech recognition method further includes:
[0127] An expansion unit 307, configured to obtain multiple original speech signals, perform transformation and expansion on each original speech signal to obtain a new speech signal, and the ways of transformation and expansion include at least one of the following: adding noise, changing the playback speed, changing the pitch of the speech signal, adding echo, and disturbing the filter characteristics of the speech signal. The sample speech signal refers to the speech stream generated by a speech producer using any language that can express meaning. The data augmentation technology is used to transform and expand the original speech data to generate more training data to improve the generalization ability of the model.
[0128] Both the original speech signal and the new speech signal are used as sample speech signals to construct a sample speech signal set; the scale of the sample speech signal set is relatively large to ensure that the speech recognition model can learn sufficiently rich speech features.
[0129] Preferably, for the speech recognition method, the expansion unit 307 is further configured to:
[0130] For application scenarios with user feature requirements, the original voice signal is formed by obtaining the natural language recording of the text read by the user for a certain period of time, so as to meet the needs of special application scenarios. The accuracy and adaptability of speech recognition are comprehensively improved to meet more extensive and complex application requirements.
[0131] Preferably, the speech recognition system further includes a speech recognition unit to be processed, and the speech recognition unit to be processed includes:
[0132] A preprocessing module for preprocessing the speech signal to be recognized to obtain the corresponding speech to be recognized; wherein, the speech signal to be recognized refers to the speech stream generated by the speech producer using any language that can express meaning;
[0133] A feature extraction module for extracting the speech feature vector to be recognized from the speech to be recognized. The speech feature vector to be recognized includes the spectral feature of the speech, the timing information of the speech, and the context connection of the speech. The spectral feature retains the speech feature of the speech producer;
[0134] A speech module for extracting the feature representation and the timing dependence relationship of the speech from the speech feature vector to be recognized. The feature representation includes the mel cepstral coefficients of the speech, the timing features of the speech, and the timing dependence relationship of the speech; mapping the spectral features to the corresponding phoneme labels or word labels in sequence according to the feature representation to obtain a mapping representation;
[0135] A language module for inputting the mapping representation and the timing dependence relationship of the speech into the language module. The language module combines and sorts the phonemes or words according to the grammar rules and semantic relationships according to the timing dependence relationship of the speech to form multiple probability factor sequences;
[0136] A module to be decoded for inputting the timing dependence relationship of the speech and multiple probability factor sequences into the module to be decoded. For each probability factor sequence, the dynamic programming algorithm in the decoding module determines the matching degree between the probability factor sequence and the speech signal to be recognized according to the timing dependence relationship of the speech, and outputs the probability factor sequence with the highest matching degree as the language text of the speech signal to be recognized.
[0137] In addition, a speech recognition device is also provided. The speech recognition device includes the aforementioned speech recognition system and a microphone array. The microphone array is used to collect speech signals. In application scenarios with user feature requirements, it can be obtained by having the user read the natural language recording formed by the text for a certain period of time.
[0138] As Figure 3 shown, it is the training and application flow chart of the speech recognition model. The brief process is as follows:
[0139] 1. Collect and input the voice signal using a microphone array.
[0140] 2. Preprocess the sample voice signal, which includes: echo cancellation processing, noise reduction processing, pre-emphasis processing, and endpoint detection to increase the amplitude of high-frequency signals and improve the spectral characteristics of the voice signal. Among them, the echo cancellation processing step processes the echo components in the voice signal to reduce its impact on speech recognition. Perform noise reduction and other processing on the voice signal. Through noise reduction processing, the noise in the signal can be further reduced to make the voice signal clearer. Then perform filtering processing to improve the signal quality. Enter the endpoint detection step to determine the start and end positions of the voice signal and obtain the preprocessed sample voice.
[0141] 3. Enter the feature extraction step to extract the features useful for speech recognition, that is, the feature vectors, from the preprocessed sample voice.
[0142] 4. For each feature vector corresponding to the sample voice signals in the sample voice signal set, perform model training through a neural network model to obtain a speech recognition model.
[0143] Among them, use a test data set to test the trained speech recognition model so that the speech recognition model can recognize the text content in the voice signal.
[0144] 5. In the step of performing speech recognition on the voice signal to be recognized, obtain multiple probability factor sequences. The probability factor sequences are calculated based on a probability distribution to evaluate the accuracy of the recognition result. Determine the most matching probability factor sequence through a decoding module (also known as a decoding unit) and present it as a language text, completing the work process of recognizing the sample voice signal.
[0145] It should be understood that the specific order or hierarchy of steps in the disclosed process is an example of an exemplary method. Based on design preferences, it should be understood that the specific order or hierarchy of steps in the process can be rearranged without departing from the scope of protection of the present disclosure. The appended method claims present the elements of the various steps in an exemplary order and are not intended to be limited to the specific order or hierarchy described.
[0146] In the above detailed description, various features are combined in a single embodiment to simplify the present disclosure. This disclosure method should not be interpreted as reflecting the intention that the embodiments of the claimed subject matter require more features than those clearly stated in each claim. On the contrary, as reflected in the appended claims, the present invention is in a state with fewer features than all the features of the disclosed single embodiment. Therefore, the appended claims are hereby clearly incorporated into the detailed description, where each claim alone serves as a separate preferred embodiment of the present invention.
[0147] The above-described embodiments have been presented for the purpose of enabling any person skilled in the art to make or use the present invention. For those skilled in the art, various modifications to these embodiments will be apparent, and the general principles defined herein may be applied to other embodiments without departing from the spirit and scope of the present disclosure. Thus, the present disclosure is not limited to the embodiments given herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
[0148] The foregoing description includes examples of one or more embodiments. Of course, it is not possible to describe all possible combinations of components or methods for the purpose of describing the above embodiments, but those of ordinary skill in the art should recognize that the various embodiments may be further combined and arranged. Thus, the embodiments described herein are intended to cover all such changes, modifications, and variations that fall within the scope of the appended claims. Additionally, with respect to the term "comprising" used in the specification or claims, this term is inclusive in a manner similar to the term "including" as interpreted when used as a transitional word in a claim. Moreover, any use of the term "or" in the specification or claims is to be meant "non-exclusive or".
[0149] Those skilled in the art will also appreciate that the various illustrative logical blocks, units, and steps listed in the embodiments of the present invention may be implemented by electronic hardware, computer software, or a combination of both. To clearly show the interchangeability of hardware and software, the various illustrative components, units, and steps have been generally described in terms of their functionality. Whether such functionality is implemented by hardware or software depends upon the particular application and design constraints of the overall system. For each particular application, those skilled in the art may use various means to implement the described functionality, but such implementation should not be construed as departing from the scope of the embodiments of the present invention.
[0150] In the embodiments of the present invention, the various illustrative logical blocks or units described can be implemented or operate the described functions by a general-purpose processor, a digital signal processor, an application specific integrated circuit (ASIC), a field programmable gate array or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination of the above designs. The general-purpose processor can be a microprocessor, and optionally, the general-purpose processor can also be any conventional processor, controller, microcontroller or state machine. The processor can also be implemented by a combination of computing devices, such as a digital signal processor and a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with a digital signal processor core, or any other similar configuration.
[0151] The steps of the methods or algorithms described in the embodiments of the present invention can be directly embedded in hardware, software modules executed by a processor, or a combination of both. The software modules can be stored in a RAM memory, a flash memory, a ROM memory, an EPROM memory, an EEPROM memory, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium in the art. Exemplarily, the storage medium can be connected to the processor so that the processor can read information from the storage medium and write information to the storage medium. Optionally, the storage medium can also be integrated into the processor. The processor and the storage medium can be disposed in an ASIC, and the ASIC can be disposed in a user terminal. Optionally, the processor and the storage medium can also be disposed in different components of the user terminal.
[0152] In one or more exemplary designs, the functions described in embodiments of the present invention may be implemented in hardware, software, firmware, or any combination of the three. If implemented in software, these functions may be stored on a computer-readable medium or transmitted as one or more instructions or code on a computer-readable medium. A computer-readable medium includes both computer storage media and communication media that facilitate transfer of a computer program from one place to another. The storage media may be any available media that can be accessed by a general-purpose or special-purpose computer. By way of example, and not limitation, such computer-readable media may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to carry or store program code in the form of instructions or data structures and that can be accessed by a general-purpose or special-purpose computer, or a general-purpose or special-purpose processor. In addition, any connection is properly termed a computer-readable medium. For example, if the software is transmitted from a website, server, or other remote source using a coaxial cable, fiber optic cable, twisted pair, DSL, or wireless means such as infrared, radio, and microwave, it is included in the definition of computer-readable medium. Disk and disc include compact disk, laser disk, optical disk, DVD, floppy disk, and Blu-ray disc, where disks usually reproduce data magnetically, while discs usually reproduce data optically with lasers. Combinations of the above should also be included within the scope of computer-readable media.
[0153] The specific embodiments described above further elaborate on the object, technical solution, and beneficial effects of the present invention. It should be understood that the above description is only for the specific embodiments of the present invention and is not intended to limit the protection scope of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A speech recognition method, characterized in that: include: Step 11: for each preprocessed sample speech, extract a corresponding sample speech feature vector from the preprocessed sample speech through a feature extraction module to be trained, wherein the sample speech feature vector retains the speech features of the speech generator, and the sample speech feature vector includes the spectral features of the speech, the time sequence information of the speech, and the contextual connection of the speech; Step 12: extracting a feature representation and a temporal dependency of speech from the sample speech feature vector through the speech module to be trained, wherein the feature representation includes Mel-frequency cepstral coefficients of speech and temporal features of speech; and mapping the spectral features of speech to corresponding phoneme labels or word labels in sequence according to the feature representation to obtain a mapping representation; Step 13: input the mapping representation and the temporal dependency of the speech into a language module to be trained, and the language module to be trained combines and sorts the phonemes or words according to grammatical rules and semantic relationships according to the temporal dependency of the speech to form a plurality of probability factor sequences; Step 14: input the temporal dependency of the speech and the plurality of probability factor sequences into the decoding module to be trained; for each probability factor sequence, determine the degree of matching between the probability factor sequence and the sample speech signal according to the temporal dependency of the speech through a dynamic programming algorithm in the decoding module to be trained; and output the probability factor sequence with the highest matching degree as the language text of the sample speech signal; the sample speech signal is the original speech of the preprocessed sample speech; the sample speech signal refers to a speech stream generated by a speech generator in any language that can express meaning; Step 15: According to the matching degree between the output language text and the sample speech signal, adjust the parameters of the modules in each step from step 11 to step 14, and continue training using the preprocessed sample speech until the output language text matches the marked text corresponding to the sample speech signal, thereby obtaining a trained speech recognition model; the speech recognition model is used to receive the speech signal to be recognized and output the corresponding language text.
2. The speech recognition method according to claim 1, characterized in that: Step 11 includes: Extracting spectral features from the preprocessed sample speech by using a mixed Mel-frequency cepstral coefficient feature extraction algorithm; Extracting speech timing information from the preprocessed sample speech through a first deep learning model, wherein the speech timing information refers to the law of evolution of a previous phoneme to a subsequent phoneme over time, and the first deep learning model is a convolutional neural network with an attention mechanism; Extracting a speech contextual connection from the preprocessed sample speech through a second deep learning model, wherein the speech contextual connection refers to a connection between a current speech frame and its adjacent speech frames, and the second deep learning model is a long short-term memory network; A sample speech feature vector corresponding to the sample speech signal is formed through the frequency spectrum feature, the time sequence information of the speech and the context of the speech.
3. The speech recognition method according to claim 1, characterized in that: Step 12 includes: Constructing a speech module to be trained, wherein the speech module to be trained includes a third deep learning model and a fourth deep learning model; Extracting a feature representation from the sample speech feature vector through the third deep learning model, the feature representation includes Mel-frequency cepstral coefficients of the speech and the time series features of the speech, the time series features of the speech represent the frequency features of the speech changing over time, and mapping the spectral features of the speech to corresponding phoneme labels or word labels in sequence according to the feature representation, wherein the third deep learning model is a convolutional neural network with an attention mechanism; The temporal dependency of speech in the sample speech feature vector is captured by the fourth deep learning model, and the fourth deep learning model is a long short-term memory network.
4. The speech recognition method according to claim 1, characterized in that: Step 13 includes: Constructing a language module to be trained, wherein the language module to be trained includes an N-gram model N-gram and a fifth deep learning model, wherein the fifth deep learning model is a transformer model Transformer; Through the N-gram based on the context grammatical rules, according to the temporal dependencies of speech whose grammatical strength is lower than the preset value in the temporal dependencies of speech, the corresponding phonemes or words are combined and sorted according to the grammatical rules and semantic relationships, and the temporal dependencies of speech whose grammatical strength is not lower than the preset value in the temporal dependencies of speech are captured through the self-attention mechanism of the fifth deep learning model, and the corresponding phonemes or words are combined and sorted according to the grammatical rules and semantic relationships to form a plurality of probability factor sequences, each probability factor sequence including one of the following: meaningful words, meaningful paragraphs, and meaningful sentences.
5. The speech recognition method according to claim 1, characterized in that: Step 14 includes: For each probability factor sequence, a dynamic programming decoding method based on the Viterbi algorithm is used to determine the degree of matching between the probability factor sequence and the sample speech signal according to the temporal dependency of the speech, and the probability factor sequence with the highest matching degree is used as the probability factor sequence matching the sample speech signal; A pre-trained language model is used to capture the grammatical and semantic structures in the matched sequence of probability factors to form a language text.
6. The speech recognition method according to claim 1, characterized in that: Also includes: Step 16, before step 12, for each of the sample speech signals, the sample speech signals are subjected to noise reduction processing and echo removal processing in sequence by the preprocessing module to be trained to obtain an initial processed speech signal; Endpoint detection is performed on the initially processed speech signal to detect the starting position and the ending position of the initially processed speech signal; speech pre-emphasis processing is performed on the initially processed speech signal from the starting position to the ending position to obtain a corresponding pre-processed sample speech.
7. The speech recognition method according to claim 1, characterized in that: Also includes: Step 17, obtaining multiple original voice signals, transforming and expanding each original voice signal to obtain a new voice signal, wherein the transformation and expansion method includes at least one of the following: adding noise, changing the playback speed, changing the pitch of the voice signal, adding echo, and disturbing the filter characteristics of the voice signal, wherein the sample voice signal refers to a voice stream generated by a voice generator in any language that can express meaning; Taking the original speech signal and the new speech signal as sample speech signals, a sample speech signal set is constructed; Wherein, step 17 is performed before step 16.
8. The speech recognition method according to claim 7, characterized in that: Step 17 also includes: For application scenarios with user feature requirements, the original voice signal is formed by obtaining the natural language recording formed by the user reading a text for a certain length of time.
9. A speech recognition system, characterized in that: include: A feature extraction module training unit is used to extract corresponding sample speech feature vectors from the preprocessed sample speech through the feature extraction module to be trained for each preprocessed sample speech, wherein the sample speech feature vector retains the speech features of the speech generator, and the sample speech feature vector includes the spectral features of the speech, the time sequence information of the speech, and the contextual connection of the speech; A speech module training unit, used for extracting a feature representation and a temporal dependency of speech from the sample speech feature vector through the speech module to be trained, wherein the feature representation includes a Mel-frequency cepstral coefficient of speech and a temporal feature of speech; and mapping the spectral features of speech to corresponding phoneme labels or word labels in sequence according to the feature representation to obtain a mapping representation; A language module training unit, used for inputting the mapping representation and the temporal dependency of the speech into the language module to be trained, and combining and sorting the phonemes or words according to grammatical rules and semantic relationships through the language module to be trained according to the temporal dependency of the speech to form a plurality of probability factor sequences; A decoding module training unit, used for inputting the temporal dependency of the speech and the plurality of probability factor sequences into a decoding module to be trained, and for each probability factor sequence, judging the degree of matching between the probability factor sequence and the sample speech signal according to the temporal dependency of the speech through a dynamic programming algorithm in the decoding module to be trained, and taking the probability factor sequence with the highest matching degree as the language text of the sample speech signal and outputting it; The preprocessed sample speech is obtained based on preprocessing the sample speech signal; Among them, according to the matching degree between the output language text and the sample voice signal, the parameters of each module to be trained are adjusted, and the pre-processed sample voice is used to continue training through the feature extraction module training unit, the voice module training unit, the language module training unit, and the decoding module training unit in sequence until the output language text matches the marked text corresponding to the sample voice signal, thereby obtaining a trained voice recognition model; the voice recognition model is used to receive the voice signal to be recognized and output the corresponding language text.
10. A computer device, characterized in that: include: processor; And, a memory arranged to store computer executable instructions, which, when executed, cause the processor to perform the speech recognition method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Automatic speech recognition method and automatic speech recognition system based on artificial intelligence
CN110827801A
Novel multi-task combination based speech recognition training framework and method
CN110875035A