A speech recognition method, device and computer equipment
By aligning phonemes and word units of speech data, generating target phonemes and word collections, and adjusting the initial speech recognition text, the problem of low recognition accuracy in existing systems when processing speech weakened languages is solved, and higher speech recognition accuracy is achieved.
Patent Information
- Application Number
- CN202110815555.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-07-19
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2041-07-19
AI Technical Summary
When existing speech recognition systems deal with languages with speech weakening, they cannot effectively model and recognize, resulting in a decrease in recognition accuracy.
By obtaining the speech feature frames of the speech data in the target language, phoneme alignment and word unit alignment are performed, target phoneme collection and target word collection are generated, and the initial speech recognition text is adjusted based on these collections to output more accurate speech recognition text.
The accuracy of speech recognition is improved, especially when dealing with languages with speech weakening, and the system's understanding and recognition ability of different language features is enhanced.
Smart Images

Figure CN113823265B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and particularly to a voice recognition method, apparatus, and computer device. Background Art
[0002] In recent years, with the rapid development of information science and technology, voice recognition technology has also developed rapidly and gradually changed our way of life and work. For example, products such as voice-controlled voice dialing systems, voice-controlled intelligent toys, and smart home appliances can make human-computer interaction simple and easy.
[0003] However, there are currently various and diverse languages. For example, Chinese, English, Russian, Arabic, etc. all belong to different languages, and each language has its own characteristics. For example, some languages have the phenomenon of phonetic weakening. In existing voice recognition systems, a multi-pronunciation dictionary is generally used to model such phenomena, but the phenomenon of phonetic weakening cannot be exhausted during the modeling process. If an existing voice recognition system is used to recognize speech with the phenomenon of phonetic weakening, the accuracy of voice recognition will be reduced. Summary of the Invention
[0004] Embodiments of this application propose a voice recognition method, apparatus, and computer device, which can improve the accuracy of voice recognition.
[0005] Embodiments of this application provide a voice recognition method, including:
[0006] Obtain at least one voice feature frame of voice data in a target language;
[0007] Perform phoneme alignment on the at least one voice feature frame to obtain a target phoneme set of the voice data in the target language;
[0008] Perform word unit alignment on the at least one voice feature frame to obtain a target word set of the voice data in the target language, where the target word set includes a word unit corresponding to each voice feature frame;
[0009] Perform text mapping on the at least one voice feature frame to obtain an initial voice recognition text of the voice data in the target language;
[0010] Adjust the initial voice recognition text according to the target phoneme set and the target word set, and obtain and output the voice recognition text of the voice data.
[0011] Correspondingly, embodiments of this application also provide a voice recognition apparatus, including:
[0012] An acquisition unit, configured to acquire at least one speech feature frame of speech data in a target language;
[0013] A phoneme alignment unit, configured to perform phoneme alignment on the at least one speech feature frame to obtain a target phoneme set of the speech data in the target language;
[0014] A word unit alignment unit, configured to perform word unit alignment on the at least one speech feature frame to obtain a target word set of the speech data in the target language, where the target word set includes word units corresponding to each speech feature frame;
[0015] A text mapping unit, configured to perform text mapping on the at least one speech feature frame to obtain an initial speech recognition text of the speech data in the target language;
[0016] An adjustment unit, configured to adjust the initial speech recognition text according to the target phoneme set and the target word set, and obtain and output the speech recognition text of the speech data.
[0017] In one embodiment, the phoneme alignment unit includes:
[0018] A path search subunit, configured to perform path search on each speech feature frame in a preset phoneme search space to obtain at least one phoneme search path;
[0019] A calculation subunit, configured to calculate the cumulative probability of the speech feature frame on each phoneme search path;
[0020] A determination subunit, configured to determine the target phoneme set of the speech data according to the cumulative probability.
[0021] In one embodiment, the path search subunit includes:
[0022] A feature enhancement module, configured to perform feature enhancement on the speech feature frame at the phoneme granularity to obtain the phoneme feature of the speech feature frame;
[0023] A screening module, configured to screen out a target phoneme set from the multiple phoneme sets according to the phoneme feature;
[0024] A phoneme search module, configured to perform phoneme search in the target phoneme set according to the phoneme feature, and generate at least one phoneme search path according to the search result.
[0025] In one embodiment, the phoneme search module includes:
[0026] A calculation sub-module, configured to calculate the matching probabilities between the phoneme feature and the multiple phoneme nodes respectively;
[0027] A determination sub-module, configured to determine at least one target phoneme node from the multiple phoneme nodes according to the matching probability;
[0028] An association sub-module, configured to associate the target phoneme nodes of each phoneme feature to obtain at least one target search path.
[0029] In one embodiment, the word unit alignment unit includes:
[0030] A path search sub-unit, configured to perform path search on each speech feature frame in a preset dictionary search space to obtain at least one word unit search path;
[0031] A calculation sub-unit, configured to calculate the cumulative probability of the speech feature frame on each word unit search path;
[0032] A determination sub-unit, configured to determine a target word set of the speech feature frame according to the cumulative probability.
[0033] In one embodiment, the text mapping unit includes:
[0034] An attention feature extraction sub-unit, configured to perform attention feature extraction on the speech feature frame in multiple attention dimensions to obtain attention features of the speech feature frame in each attention dimension;
[0035] A decoding sub-unit, configured to decode the attention features in each attention dimension to obtain an initial speech recognition text of the speech data in the target language.
[0036] In one embodiment, the adjustment unit includes:
[0037] An identification sub-unit, configured to identify phoneme information of the initial speech recognition text;
[0038] A first adjustment sub-unit, configured to adjust the phoneme information by using the target phoneme set to obtain a phoneme-adjusted speech recognition text;
[0039] A second adjustment sub-unit, configured to adjust the phoneme-adjusted speech recognition text by using the target word set to obtain and output the speech recognition text of the speech data.
[0040] An embodiment of the present application further provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in various alternative manners in the above-mentioned aspect.
[0041] Correspondingly, an embodiment of the present application further provides a storage medium, which stores instructions that, when executed by a processor, implement the voice recognition method provided in any one of the embodiments of the present application.
[0042] Embodiments of the present application can obtain at least one voice feature frame of voice data in a target language; perform phoneme alignment on at least one voice feature frame respectively to obtain a target phoneme set of the voice data in the target language; perform word unit alignment on at least one voice feature frame respectively to obtain a target word set of the voice data in the target language, where the target word set includes the word unit corresponding to each voice feature frame; perform text mapping on at least one voice feature frame respectively to obtain an initial voice recognition text of the voice data in the target language; and adjust the initial voice recognition text according to the target phoneme set and the target word set to obtain and output the voice recognition text of the voice data, thereby improving the accuracy of voice recognition. Description of the Drawings
[0043] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those skilled in the art, without creative efforts, other drawings can be obtained according to these drawings.
[0044] Figure 1 is a schematic diagram of the scenario of the voice recognition method provided by the embodiment of the present application;
[0045] Figure 2 is a schematic flowchart of the voice recognition method provided by the embodiment of the present application;
[0046] Figure 3 is a schematic diagram of the scenario of windowing and sliding for voice data provided by the embodiment of the present application;
[0047] Figure 4 is a schematic structural diagram of a preset voice recognition model provided by the embodiment of the present application;
[0048] Figure 5 is a schematic flowchart of generating a phoneme identification frame provided by the embodiment of the present application;
[0049] Figure 6 is a schematic diagram of the scenario of phoneme masking for a phoneme annotation frame provided by the embodiment of the present application;
[0050] Figure 7 is a schematic diagram of the scenario of training a preset voice recognition model to be trained provided by the embodiment of the present application;
[0051] Figure 8It is a schematic diagram of a scenario of a preset phoneme search space provided by an embodiment of the present application;
[0052] Figure 9 It is a schematic diagram of a scenario of path search provided by an embodiment of the present application;
[0053] Figure 10 It is another flowchart of a speech recognition method provided by an embodiment of the present application;
[0054] Figure 11 It is a schematic structural diagram of a speech recognition device provided by an embodiment of the present application;
[0055] Figure 12 It is a schematic structural diagram of a terminal provided by an embodiment of the present application. Detailed implementation manners
[0056] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. However, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.
[0057] An embodiment of the present application proposes a speech recognition method. The speech recognition method can be executed by a speech recognition device based thereon, and the speech recognition device based thereon can be integrated in a computer device. Among them, the computer device can include a terminal, a server, and so on.
[0058] Among them, the terminal can be a notebook computer, a personal computer (PC), an in-vehicle computer, and so on.
[0059] Among them, the server can be an intercommunication server or a background server between multiple heterogeneous systems, and can also be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, and big data and artificial intelligence platforms, and so on.
[0060] In one embodiment, as Figure 1As described above, the speech recognition device can be integrated into computer devices such as terminals or servers to implement the speech recognition method proposed in the embodiments of the present application. Specifically, the computer device can obtain at least one speech feature frame of the speech data in the target language; perform phoneme alignment on the at least one speech feature frame to obtain the target phoneme set of the speech data in the target language; perform word unit alignment on the at least one speech feature frame to obtain the target word set of the speech data in the target language, where the target word set includes the word unit corresponding to each speech feature frame; perform text mapping on the at least one speech feature frame to obtain the initial speech recognition text of the speech data in the target language; and adjust the initial speech recognition text according to the target phoneme set and the target word set to obtain and output the speech recognition text of the speech data.
[0061] The following will be described in detail respectively. It should be noted that the description order of the following embodiments does not limit the preferred order of the embodiments.
[0062] The embodiments of the present application will be described from the perspective of the speech recognition device. The speech recognition device can be integrated into a computer device, which can be a server or a terminal or other devices.
[0063] As Figure 2 described above, a speech recognition method is provided. The specific process includes:
[0064] 101. Obtain at least one speech feature frame of the speech data in the target language.
[0065] Among them, the target language can include various languages with special phenomena in daily use. For example, the target language can include languages with weakening phenomena in daily use. For another example, the target language can include languages with trill phenomena in daily use. For another example, the target language can include languages with elision phenomena in daily use, and so on.
[0066] Among them, languages with weakening phenomena can include languages with weakening phenomena in vowels and consonants.
[0067] In one embodiment, the speech feature frame includes a frame that can identify the features of the speech data.
[0068] For example, sound is actually a kind of wave. What the speech recognition task faces is a sequence of samples after several signal processes, which is also called a waveform. Among them, the waveform can be speech data. And the speech feature frame is the data frame obtained after feature extraction of the speech data.
[0069] In one embodiment, according to different feature extraction methods, the speech feature frame also has different expression forms.
[0070] For example, when extracting features from speech data using Mel-scale Frequency Cepstral Coefficients (MFCC), the speech feature frame can be MFCC. Another example is that when extracting features from speech data using FilterBank (FBank), the speech feature frame can be FBank. Still another example is that when extracting features from speech data frames using Linear Prediction Coefficient (LPC), the speech feature frame can be LPC.
[0071] In one embodiment, when extracting features from speech data using MFCC, the speech data can first be windowed by sliding, thereby dividing the speech data into frames. For example, as Figure 3 shown, Figure 3 001 in it can be the speech data, and the speech data can be divided into frames by the sliding window method. Among them, when windowing the speech data by sliding, the frame length is usually set to 25 ms and the frame shift is set to 10 ms, which can ensure the stationarity of the signal within the frame and make the frames overlap, improving the reliability of the frames.
[0072] Next, a Fast Fourier Transform (FFT) can be performed on each frame, and the power spectrum can be calculated. Then, the Mel filter bank is applied to the power spectrum to obtain the logarithmic energy within each filter as coefficients. Finally, a Discrete Cosine Transform (DCT) can be performed on the obtained Mel filter logarithmic energy vector to obtain the speech data frame.
[0073] In one embodiment, the speech recognition device proposed in the embodiments of the present application can be integrated into various computer devices. For example, the speech recognition device proposed in the embodiments of the present application can be integrated into a mobile phone so that when people control the mobile phone by voice, the mobile phone can recognize the corresponding text information in the voice through the speech recognition device. Another example is that the speech recognition device proposed in the embodiments of the present application can be integrated into various smart homes so that when people control the smart home by voice, the smart home can recognize the corresponding text information in the voice through the speech recognition device.
[0074] In one embodiment, to more conveniently implement the speech recognition method proposed in the embodiments of the present application, a preset speech recognition model is proposed in the embodiments of the present application. Among them, the preset speech recognition model can be an end-to-end speech recognition model, and through this preset speech recognition model, the speech data can be directly converted into a speech recognition text, thereby improving the accuracy and efficiency of speech recognition.
[0075] In one embodiment, the model architecture of the preset speech recognition model may include an encoding layer, a phoneme alignment layer, a word unit alignment layer, a decoding layer, and an attention layer. For example, as Figure 4 shown.
[0076] Among them, the encoding layer can obtain speech data and extract features from the speech data.
[0077] Among them, the phoneme alignment layer can be used to align phonemes for at least one speech feature frame to obtain a set of target phonemes of the speech data in the target language.
[0078] Among them, the word unit alignment layer can be used to align word units for at least one speech feature frame to obtain a set of target words of the speech data in the target language.
[0079] Among them, the attention layer and the decoding layer can perform text mapping on at least one speech feature frame to obtain an initial speech recognition text of the speech data in the target language.
[0080] In addition, the preset speech recognition model can also adjust the initial speech recognition text according to the set of target phonemes and the set of target words, and obtain and output the speech recognition text of the speech data.
[0081] Among them, the encoding layer can be a machine learning network or a deep learning network. For example, the encoding layer can be any one of a convolutional neural network (CNN), a recurrent neural network (RNN), a de-convolutional network (DN), a deep neural network (DNN), a deep convolutional inverse graphics network (DCIGN), a region-based convolutional network (RCNN), a faster region-based convolutional network (Faster RCNN), and a bidirectional encoder representations from transformers (BERT) model, etc.
[0082] Among them, the decoding layer can also be a machine learning network or a deep learning network. For example, the decoding layer can be one of the networks such as CNN, RNN, DN, DNN, etc.
[0083] Among them, the attention layer can include a machine learning network or a deep learning network with an attention mechanism. Among them, the attention mechanism stems from the research on human vision. In cognitive science, due to the bottleneck of information processing, humans will selectively focus on a part of all information while ignoring other visible information. The above mechanism is usually called the attention mechanism. Different parts of the human retina have different degrees of information processing capabilities, that is, acuity, and only the fovea of the retina has the strongest acuity. In order to make reasonable use of limited visual information processing resources, humans need to select specific parts of the visual area and then focus on it. For example, when people are reading, usually only a small number of words to be read will be focused on and processed. In summary, the attention mechanism mainly has two aspects: determining which part of the input needs to be focused on; allocating limited information processing resources to important parts.
[0084] Among them, the phoneme alignment layer can have a machine learning network or a deep learning network with the Connectionist Temporal Classification (CTC) algorithm. For example, the phoneme alignment layer can be an RNN improved based on CTC.
[0085] Among them, the word unit alignment layer can also be a machine learning network or a deep learning network with CTC. For example, the word unit alignment layer can be an RNN improved based on CTC.
[0086] Among them, the phoneme alignment layer aligns the speech feature frames at the phoneme granularity, while the word unit alignment layer aligns the speech feature frames at the word granularity, and their focuses are different.
[0087] In one embodiment, the preset speech recognition model proposed in the embodiments of the present application forms an end-to-end speech recognition hybrid model by combining CTC and the attention layer, so that the preset speech recognition model can learn more granular alignment information, so that the model can better find the alignment of speech features to multiple modeling unit sequences (the modeling units in the embodiments of the present application are phoneme units and word units), and finally improve the performance of the speech recognition model.
[0088] In one embodiment, before using the preset speech recognition model for speech recognition, the speech recognition model to be trained can be trained to obtain the preset speech recognition model. Specifically, the steps of training the speech recognition model to be trained can include:
[0089] Obtain multiple phoneme identification frames and the speech recognition model to be trained;
[0090] Perform phoneme masking processing on multiple phoneme identification frames to obtain masked phoneme identification frames;
[0091] Use the masked phoneme identification frames to train the speech recognition model to be trained, and obtain a preset speech recognition model.
[0092] Among them, the speech recognition model to be trained includes a model with poor speech recognition performance that still needs to be trained.
[0093] Among them, a phoneme identification frame is an audio frame with phoneme identification information, and through the phoneme identification information in the phoneme identification frame, it is possible to know what phoneme the phoneme frame corresponds to.
[0094] By using the phoneme identification frames to train the speech recognition model to be trained, the training process belongs to supervised learning, so that developers can control the training process of the model, improving the reliability and robustness of the training process.
[0095] In one embodiment, since the phoneme identification frame has phoneme identification information, before obtaining the phoneme identification frame, it is necessary to obtain training data and generate phoneme identification frames according to the training data. Among them, the process of generating phoneme identification frames according to the training data can be as Figure 5 shown. Specifically, the training data is first processed by a Hidden Markov Model - Gaussian Mixture Model to obtain first processed data. Then, the first processed data is further processed by a Hidden Markov Model - Deep Neural Network to obtain second processed data. Next, the second processed data is subjected to alignment processing to obtain first aligned data. In addition, the training data is also directly subjected to alignment processing to obtain second aligned data. Finally, the first aligned data and the second aligned data are combined to obtain phoneme identification frames.
[0096] Among them, the training data may include various speech data in the target language.
[0097] Among them, the Hidden Markov Model (HMM) is a statistical model that is used to describe a Markov process containing hidden unknown parameters, determine the hidden parameters of the process from the observable parameters, and then use these parameters for further analysis. Among them, in the field of speech recognition, HMM uses two stochastic processes, namely the state transition process and the observation sampling process, to model the conversion process from voice features to pronunciation units as a probability problem, and trains the parameters of the Hidden Markov Model through the existing speech data. During decoding, using the corresponding parameters, estimate the probability of converting the input acoustic features into a specific sequence of pronunciation units, and then obtain the probability of outputting specific text, so as to select the text that is most likely to represent a certain section of sound.
[0098] Among them, the Gaussian Mixture Model (GMM) quantifies things precisely using the Gaussian probability density function (normal distribution curve). It is a model that decomposes things into several models formed based on the Gaussian probability density function (normal distribution curve). In the field of speech recognition, in a standard Hidden Markov Model, when observing quantities are output from the hidden pronunciation states, it is necessary to model the output probability distribution. In a classic speech recognition system based on the Hidden Markov Model, this process is generally modeled using the Gaussian Mixture Model (Gaussian Mixture Model).
[0099] Among them, the HMM-GMM model is a classic speech recognition system. This model can convert the input speech data into text information using the knowledge of probability theory and statistics.
[0100] Among them, the Deep Neural Networks (DNN) is actually also a type of probability model.
[0101] In recent years, with the development of artificial intelligence technology, some researchers have begun to attempt to combine machine learning or deep learning with the HMM-GMM model to improve the performance of speech recognition. The HMM-DNN is one of the variants of the HMM-GMM model.
[0102] The HMM-DNN model can also convert the input speech data into text information. However, different from the HMM-GMM, when the HMM-DNN processes speech data, it requires a frame-level alignment information as a basis, while when the HMM-GMM processes speech data, it does not require a frame-level alignment information, and the HMM-GMM can also generate frame-level alignment information.
[0103] Therefore, when generating the phoneme identification frame, the phoneme identification information of the training data can be first generated using the HMM-GMM, and then the HMM-DNN can use the phoneme identification information as a basis to generate the first phoneme identification frame of the training data.
[0104] In one example, in order to improve the reliability of the phoneme identification frame, a model with CTC can also be used to perform phoneme alignment on the training data to obtain the second phoneme identification frame of the speech data. Then, the first phoneme identification frame and the second phoneme identification frame are combined to obtain the phoneme identification frame.
[0105] In one embodiment, in order to improve the accuracy of speech recognition by a preset speech recognition, phoneme masking processing can be performed on multiple phoneme identification frames to obtain masked phoneme identification frames. Then, the masked phoneme identification frames are used to train a speech recognition model to be trained. By using the masked phoneme identification frames to train the speech recognition model to be trained, during the training process, the speech recognition model to be trained can autonomously learn more speech context knowledge, thereby improving the recognition performance of the preset speech recognition model.
[0106] In one embodiment, when performing phoneme masking processing on a phoneme identification frame, conversion processing can be performed on some phoneme information in the phoneme identification frame, so that some information of the phoneme identification frame is masked. Specifically, the step of "performing phoneme masking processing on multiple phoneme identification frames respectively to obtain masked phoneme identification frames" may include:
[0107] Screen out target phoneme information from the phoneme information of the phoneme identification frame;
[0108] Perform information conversion processing on the target phoneme information to obtain converted phoneme information;
[0109] Add the converted phoneme information to the phoneme identification frame to obtain masked phoneme identification information.
[0110] Among them, the phoneme information includes the information constituting the phoneme annotation frame. For example, as Figure 3 shown, Figure 3 frames n and n + 1 in are phoneme identification frames, where the waveforms in the phoneme identification frames can be phoneme information.
[0111] In one embodiment, several pieces of phoneme information can be screened out from the phoneme information as target phoneme information. For example, as Figure 6 shown, the phoneme information in the phoneme annotation frame includes "n", "A", "d", "i", "va", "s", "i", "y", "A", "w", "a", and "H". Then, "i" and "a" can be screened out from these phoneme information as target phoneme information.
[0112] In one embodiment, after screening out the target phoneme information, information conversion processing can be performed on the target phoneme information to obtain converted phoneme information. For example, the information of phonemes "i" and "a" can be set to 0. For example, the waveforms corresponding to phonemes "i" and "a" can be set to zero. For another example, the information of phonemes "i" and "a" can also be added to obtain converted phoneme information. For example, the waveforms of phonemes "i" and "a" can be superimposed, and the superimposed waveform is used as the converted phoneme information. For another example, the information of phonemes "i" and "a" can also be added and then averaged to obtain converted phoneme information, and so on.
[0113] In one embodiment, after obtaining the converted phoneme information, the converted phoneme information can be added to the phoneme identification frame, so as to obtain the masked phoneme identification information.
[0114] For example, the converted phoneme information can be used to replace the target phoneme information, so as to obtain the masked phoneme identification information.
[0115] For example, when the information of phonemes "i" and "a" is set to 0, the waveforms corresponding to phonemes "i" and "a" in the phoneme identification frame can be replaced with no waveforms, so as to obtain the masked phoneme identification information.
[0116] In one embodiment, after obtaining the masked phoneme identification frame, the masked phoneme identification frame can be used to train the speech recognition model to be trained, so as to obtain the preset speech recognition model. Specifically, the step of "using the masked phoneme identification frame to train the speech recognition model to be trained to obtain the preset speech recognition model" may include:
[0117] Using the speech recognition model to be trained to extract features from the masked phoneme identification frame to obtain the feature information of the masked phoneme identification frame;
[0118] Using the speech recognition model to be trained to perform phoneme alignment and word unit alignment on the feature information respectively to obtain the target phoneme and target word unit of the masked phoneme identification frame;
[0119] Using the speech recognition model to be trained to perform text mapping on the feature information to obtain the speech recognition text of the masked phoneme identification frame;
[0120] Performing joint operations on the target phoneme, target word unit and speech recognition text to obtain joint loss information;
[0121] Adjusting the model parameters of the speech recognition model to be trained according to the joint loss information to obtain the preset speech recognition model.
[0122] Among them, training the model may include enabling the model to learn from a large amount of data, so that the model can summarize the rules from the large amount of data, and can process any data input into the model according to the rules.
[0123] Among them, training the speech recognition model to be trained may be to make the speech model to be trained continuously perform speech recognition on the masked phoneme identification frame, so that the speech recognition model to be trained learns how to convert speech data into text information through the masked phoneme identification frame.
[0124] In one embodiment, the flowchart of the speech recognition model to be trained can be as Figure 7As shown. Among them, the encoding layer in the speech recognition model to be trained can be used to extract features from the masked phoneme identification frames, obtaining the feature information of the masked phoneme identification frames.
[0125] In one embodiment, the training process of the encoding layer is to let the encoding layer learn acoustic knowledge, that is, to learn which modeling unit the current masked phoneme identification frame is more similar to and represent it with a vector. Different from the conventional training method, in the embodiment of the present application, by using the masked phoneme identification frames to train the encoding layer, the encoding layer can be encouraged to learn to predict the identification of the current masked phoneme identification frame based on the surrounding speech frames.
[0126] In one embodiment, in an end-to-end speech recognition hybrid model, the role of CTC is generally to make the model training converge faster and, in addition, to assist the encoding layer, thereby improving the recognition performance of the model. In the preset speech recognition model proposed in the embodiment of the present application, CTC can also enable the encoding layer to learn more knowledge from the masked phoneme identification frames. Therefore, the phoneme alignment layer in the speech recognition model to be trained can be used to perform phoneme alignment on the feature information, and the word unit alignment layer can be used to perform word unit alignment on the feature information, thereby obtaining the target phoneme and target word unit of the masked phoneme identification frame. In addition, the attention layer and decoding layer in the speech recognition model to be trained can be used to perform text information on the feature information to obtain the speech recognition text of the masked phoneme identification frame.
[0127] Among them, there is no restriction on the execution order of the steps of "using the speech recognition model to be trained to perform phoneme alignment and word unit alignment on the feature information respectively to obtain the target phoneme and target word unit of the masked phoneme identification frame" and "using the speech recognition model to be trained to perform text mapping on the feature information to obtain the speech recognition text of the masked phoneme identification frame". Either the step of "using the speech recognition model to be trained to perform phoneme alignment and word unit alignment on the feature information respectively to obtain the target phoneme and target word unit of the masked phoneme identification frame" can be executed first, or the step of "using the speech recognition model to be trained to perform text mapping on the feature information to obtain the speech recognition text of the masked phoneme identification frame" can be executed first, or the two steps can be executed in parallel.
[0128] Next, joint operations can be performed on the target phoneme, target word unit, and speech recognition text to obtain joint loss information, and the model parameters of the speech recognition model to be trained can be adjusted according to the joint loss information to obtain the preset speech recognition model.
[0129] In one embodiment, when performing a joint operation on the target phoneme, the target word unit, and the speech recognition text to obtain joint loss information, the alignment loss information between the target phoneme and the target word unit can be calculated, and the text loss information between the speech recognition text and the preset identification text can be calculated. Then, the alignment loss information and the text loss information are fused to obtain the joint loss information. Specifically, the step of "performing a joint operation on the target phoneme, the target word unit, and the speech recognition text to obtain joint loss information" may include:
[0130] Calculating the alignment loss information between the target phoneme and the target word unit;
[0131] Calculating the text loss information between the speech recognition text and the preset identification text;
[0132] Fusing the alignment loss information and the text loss information to obtain the joint loss information.
[0133] Among them, the alignment loss information includes the phoneme loss information between the target phoneme and the preset phoneme and the word unit loss information between the target word unit and the preset word unit.
[0134] In one embodiment, since the phoneme identification frame has phoneme identification information, the preset phoneme and the preset word unit of the phoneme identification frame can be included in the speech recognition model to be trained. That is, through the phoneme identification information, the speech recognition model to be trained already knows which phonemes the phoneme identification frame corresponds to and which word units. Therefore, when calculating the alignment loss information, the phoneme loss information between the target phoneme and the preset phoneme and the word unit loss information between the target word unit and the preset word unit can be calculated respectively.
[0135] Among them, the CTC function can be used to calculate the phoneme loss information between the target phoneme and the preset phoneme. Similarly, the CTC function can also be used to calculate the word unit loss information between the target word unit and the preset word unit.
[0136] In one embodiment, after obtaining the phoneme loss information and the word unit loss information, the phoneme loss information and the word unit loss information can be fused to obtain the alignment loss information.
[0137] For example, the alignment loss information can be represented as Loss_CTC, the phoneme loss information can be represented as Loss_grapheme_CTC, and the word unit loss information can be represented as Loss_word-piece _CTC. Then, the phoneme loss information and the word unit loss information can be fused according to the following formula to obtain the alignment loss information:
[0138] Loss_CTC = Loss_word-piece _CTC + Theta Loss_grapheme_CTC
[0139] In one embodiment, when calculating the text loss information between the speech recognition text and the preset identification text, the cross-entropy function (Cross-Entropy loss, CE loss) can be used to calculate the text loss information between the speech recognition text and the preset identification text. Then, the alignment loss information and the text loss information can be fused to obtain the joint loss information.
[0140] For example, the text loss information can be expressed as Loss_CE, and the joint loss information can be expressed as Loss_Joint. Then, the text loss information and the joint loss information can be fused according to the following formula to obtain the joint loss information:
[0141] Loss_Joint = Loss_CE + Alpha Loss_CTC
[0142] In one embodiment, after obtaining the joint loss information, the model parameters of the speech recognition model to be trained can be adjusted according to the joint loss information, so as to obtain the preset speech recognition model.
[0143] 102. Align the phonemes of at least one speech feature frame to obtain the target phoneme set of the speech data in the target language.
[0144] In one embodiment, after obtaining at least one speech feature frame of the speech data, the speech feature frame can be phoneme-aligned to obtain the target phoneme set of the speech data in the target language.
[0145] Among them, a phoneme includes the smallest unit in speech and is the smallest speech unit divided from the perspective of timbre.
[0146] For example, there are 48 phonemes in English, among which there are 20 vowel phonemes and 28 consonant phonemes. Another example is that there are 32 phonemes in Chinese, among which there are 10 vowel phonemes and 22 consonant phonemes, and so on.
[0147] For example, taking Chinese as an example, the Chinese syllable "ā" has only one phoneme, "ài" has two phonemes, "dài" has three phonemes, etc.
[0148] Among them, phoneme alignment refers to determining what the corresponding phoneme in each speech feature frame is.
[0149] Among them, the target phoneme set includes the set of phonemes corresponding to each speech feature frame. Through the target phoneme set, the computer device can obtain what the pronunciation of the speech data is.
[0150] In one embodiment, when performing phoneme alignment on speech feature frames, path search can be performed on each speech feature frame in a preset phoneme search space to obtain at least one phoneme search path, and a target phoneme set can be generated in the phoneme search path. Specifically, the step of "performing phoneme alignment on at least one speech feature frame respectively to obtain the target phoneme set of the speech data in the target language" may include:
[0151] Performing path search on each speech feature frame in the preset phoneme search space to obtain at least one phoneme search path;
[0152] Calculating the cumulative probability of the speech feature frame on each phoneme search path;
[0153] Determining the target phoneme set of the speech data according to the cumulative probability.
[0154] Among them, the preset phoneme search space may include a space composed of acoustic knowledge in the target language. In the preset phoneme search space, what features each phoneme in the target language has, and the relationships between various phonemes, etc. are defined.
[0155] For example, the preset phoneme search space may include the MFCC corresponding to each phoneme in the target language. Another example is that some phonemes are always used together, then the distance between these phonemes in the preset phoneme search space will be relatively close. On the contrary, if some phonemes are not used together, then the distance between these phonemes in the preset phoneme search space will be relatively far.
[0156] In one embodiment, the preset phoneme search space can have various forms of representation. For example, the preset phoneme search space can be a matrix. Another example is that the preset phoneme search space can be a graph structure. Another example is that the preset phoneme search space can be a tree structure, etc.
[0157] In one embodiment, path search can be performed on each phoneme feature frame in the preset phoneme search space to obtain at least one phoneme search path. Among them, the phoneme search space includes multiple phoneme sets. Therefore, when performing path search on each phoneme feature frame in the preset phoneme search space, the phoneme features can be searched for paths according to the phoneme sets. Specifically, the step of "performing path search on each speech feature frame in the preset phoneme search space to obtain at least one phoneme search path" may include:
[0158] Enhancing the features of the speech feature frame at the phoneme granularity to obtain the phoneme features of the speech feature frame;
[0159] Filtering out the target phoneme set from multiple phoneme sets according to the phoneme features;
[0160] Perform phoneme search in a target phoneme set according to phoneme features, and generate at least one phoneme search path based on the search results.
[0161] In one embodiment, to improve efficiency, phonemes with similar phoneme features can be grouped together to form a phoneme feature set. And multiple phoneme feature sets grouped together constitute a preset phoneme search space.
[0162] In one embodiment, to obtain a phoneme search path more precisely, the speech feature frames can be enhanced in terms of features at the phoneme granularity to obtain the phoneme features of the speech feature frames. For example, the speech feature frames can be spectrally enhanced to obtain the phoneme features of the speech feature frames.
[0163] Next, the target phoneme combination can be selected from multiple phoneme sets according to the phoneme features. In one embodiment, each phoneme set can include at least one preset phoneme. Therefore, the phoneme features of each preset phoneme can be normalized as the identifying phoneme features of the phoneme set. Then, the phoneme features can be matched with the identifying phoneme features on each phoneme set, and the target phoneme set can be selected from them according to the matching results.
[0164] Among them, when the matching degrees between the phoneme features and the preset phoneme features on multiple phoneme sets are the same, all these multiple phoneme sets can be regarded as the target phoneme sets.
[0165] In one embodiment, after the target phoneme set is selected, phoneme search can be performed in the target phoneme set according to the phoneme features, and at least one phoneme search path can be generated based on the search results. Specifically, the step "perform phoneme search in the target phoneme set according to the phoneme features, and generate at least one phoneme search path based on the search results" can include:
[0166] Calculate the matching probabilities between the phoneme features and at least one preset phoneme respectively;
[0167] Determine the target phoneme among at least one preset phoneme according to the matching probabilities;
[0168] Associate the target phonemes of each phoneme feature to obtain a phoneme search path.
[0169] Among them, when calculating the matching probabilities between the phoneme features and each preset phoneme, various probability algorithms can be used for calculation. For example, the maximum likelihood estimation (MLE) algorithm can be used to calculate the matching probabilities between the phoneme features and at least one preset phoneme. Then, the target phoneme can be determined among at least one preset phoneme according to the matching probabilities, and the target phonemes of each phoneme feature in each speech data can be associated to obtain the target search path.
[0170] For example, as Figure 8 shown, Figure 8 002 in [[ ]] can be a preset phoneme search space. Then, each column in the speech phoneme search space can be a phoneme set. For example, Figure 8 003 in [[ ]] can be a phoneme set. Then, the phoneme set includes at least one preset phoneme. For example, Figure 8 004 in [[ ]] can be a preset phoneme.
[0171] In one embodiment, as Figure 9 shown, by performing a path search on each speech feature frame in the speech data, at least one phoneme search path is obtained. Among them, after obtaining at least one phoneme search path, the cumulative probability of each speech feature frame on each phoneme search path can be calculated, and the target phoneme set of the speech data can be determined according to the cumulative probability.
[0172] For example, as Figure 9 shown, when calculating the cumulative probability of a speech feature frame on the phoneme search path 005, the matching probabilities between each target phoneme and the phoneme feature on the phoneme search path 005 can be cumulatively calculated to obtain the cumulative probability.
[0173] Specifically, the step of "calculating the cumulative probability of each speech feature frame on each phoneme search path" may include:
[0174] Obtain the matching probability of each target phoneme on the phoneme search path
[0175] Cumulatively calculate the matching probabilities of each target phoneme to obtain the cumulative probability.
[0176] Among them, there are various ways to cumulatively calculate the matching probabilities of each target phoneme to obtain the cumulative probability. For example, the matching probabilities of each target phoneme can be added to obtain the cumulative probability. Another example is that the matching probabilities of each target phoneme can be weighted and summed to obtain the cumulative probability.
[0177] In one embodiment, after obtaining the cumulative probability on each phoneme search path, the target phoneme set of the speech data can be determined according to the cumulative probability. Specifically, the step of "determining the target phoneme set of the speech data according to the cumulative probability" may include:
[0178] Compare the cumulative probabilities on each phoneme search path to obtain a comparison result;
[0179] Determine the target phoneme search path from at least one phoneme search path according to the comparison result;
[0180] Concatenate the target phonemes on the target phoneme search path to obtain a set of target phonemes.
[0181] For example, as shown in the figure, by comparing the cumulative probabilities, it is found that the cumulative probability of phoneme search path 005 is the largest. Therefore, phoneme search path 005 can be determined as the target phoneme search path. Then, the target phonemes on the target phoneme search path can be concatenated to obtain a set of target phonemes. For example, as shown in the figure, the target phonemes in the set of target phonemes are respectively "a:", "-", "t", "-", "i", "-", "-", "k", "-", "-", "l", where "-" can refer to a blank phoneme. Therefore, when concatenating the target phonemes, the blank phonemes can be deleted to obtain the set of target phonemes. For example, after concatenating the target phonemes in the figure, the set of target phonemes obtained is "a:tikl".
[0182] By performing path search on each speech feature frame in the preset phoneme search space, the obtained set of target phonemes can fully represent the phonemes that may be included in the speech data, improving the accuracy of phoneme alignment for speech feature frames.
[0183] 103. Perform word unit alignment on at least one speech feature frame to obtain a set of target words in the target language for the speech data, where the set of target words includes the word units corresponding to each speech feature frame.
[0184] In one embodiment, the embodiments of the present application can also perform word unit alignment on at least one speech feature frame to obtain a set of target words in the target language for the speech data.
[0185] Among them, a word unit can be the smallest unit that constitutes a word in the target language. For example, taking English as an example, the word units of the word "hello" can include "he" and "llo". Another example is that the word units of the word "loving" can include "lov" and "ing".
[0186] In one embodiment, the word units of the target language can be obtained through a subword algorithm. The subword algorithm can include algorithms that can convert words into word units. For example, the subword algorithm can include the Byte Pair Encoding (BPE) algorithm, the word-piece algorithm, and so on.
[0187] For example, the word-piece algorithm can be included in the word unit alignment layer in the preset speech recognition model proposed in the embodiments of the present application, so that this unit alignment layer can perform word unit alignment on speech identification frames through the word-piece algorithm, thereby improving the accuracy of word unit alignment.
[0188] In one embodiment, when aligning at least one speech feature frame with word units, a method similar to that for aligning speech feature frames with phonemes can be adopted. The difference is that word unit alignment aligns speech feature frames at the word granularity, while phoneme alignment aligns speech feature frames at the phoneme granularity. Specifically, the step of "respectively aligning the at least one speech feature frame with word units to obtain a set of target words in the target language" may include:
[0189] Perform path search for each speech feature frame in the preset dictionary search space to obtain at least one word unit search path;
[0190] Calculate the cumulative probability of the speech feature frame on each word unit search path;
[0191] Determine the set of target words of the speech feature frame according to the cumulative probability.
[0192] Among them, the preset dictionary search space may include a space composed of word knowledge in the target language. In the preset dictionary search space, the phonemes corresponding to each word unit in the target language are defined, the characteristics of the phonemes corresponding to each word unit, and the relationships between each word unit, etc.
[0193] In one embodiment, the preset dictionary search space may also include multiple word unit sets, and the step of performing path search for each speech feature frame in the preset dictionary search space may refer to the steps of aligning speech feature frames with phonemes. Therefore, the step of "performing path search for each speech feature frame in the preset dictionary search space to obtain at least one word unit search path" may include:
[0194] Enhance the features of the speech feature frame at the word unit granularity to obtain the word unit features of the speech feature frame;
[0195] Filter out the set of target word units from the multiple word unit sets according to the word unit features;
[0196] Perform word search in the set of target word units according to the word unit features, and generate at least one word unit search path according to the search results.
[0197] In one embodiment, in order to improve efficiency, word units with similar word unit features can be grouped together to form a word unit set. And multiple word unit sets grouped together constitute the preset dictionary search space.
[0198] In one embodiment, in order to obtain the word unit search path more precisely, the features of the speech feature frame can be enhanced at the word unit granularity, so as to obtain the word unit features of the speech feature frame. For example, the speech feature frame can be subjected to secondary feature extraction to obtain the word unit features.
[0199] Next, the target word unit set can be screened out from multiple word unit sets according to the word unit features. Among them, the reference steps for screening out the target word unit from multiple word unit sets can be the steps of "screening out the target phoneme set from multiple phoneme sets", which will not be repeated here.
[0200] In one embodiment, after screening out the target word unit, word search can be performed in the target word unit set according to the word unit features, and at least one word unit search path can be generated according to the search results. Specifically, the steps of "performing word search in the target word unit set according to the word unit features and generating at least one word unit search path according to the search results" may include:
[0201] Calculate the matching probabilities between the word unit features and at least one preset word unit respectively;
[0202] Determine the target word unit from at least one preset word unit according to the matching probabilities;
[0203] Associate the target word units of each word unit feature to obtain a word unit search path.
[0204] By performing phoneme alignment and word unit alignment on the speech feature frames, the target phoneme set and target word set of the speech data can be obtained. Then, the initial speech recognition text can be adjusted by using the target word set and target phoneme set, thereby improving the accuracy of the speech recognition text.
[0205] 104. Perform text mapping on at least one speech feature frame to obtain the initial speech recognition text of the speech data in the target language.
[0206] In one embodiment, the embodiments of the present application can also perform text mapping on at least one speech feature frame to obtain the initial speech recognition text of the speech data in the target speech. Specifically, the steps of "performing text mapping on at least one speech feature frame respectively to obtain the initial speech recognition text of the speech data in the target language" may include:
[0207] Extract attention features of the speech feature frames in multiple attention dimensions to obtain the attention features of the speech feature frames in each attention dimension;
[0208] Decode the attention features in each attention dimension to obtain the initial speech recognition text of the speech data in the target language.
[0209] In one embodiment, in order to improve the accuracy of speech recognition, attention features of the speech feature frames can be extracted in multiple attention dimensions to obtain the attention features of the speech feature frames in each attention dimension.
[0210] Among them, the multi-head attention mechanism can be used to extract attention features from speech feature frames, and each head of the attention mechanism can correspond to an attention dimension.
[0211] In one embodiment, after obtaining the attention features, the attention features on each attention dimension can be decoded to obtain the initial speech recognition text of the speech data in the target language.
[0212] For example, the feature distribution of the attention features in the preset text mapping space can be calculated, and then the text information corresponding to the attention features can be determined according to the feature distribution.
[0213] In one embodiment, it should be noted that there is no restriction on the execution order among the steps of "performing phoneme alignment on at least one speech feature frame respectively to obtain the target phoneme set of the speech data in the target language", "performing word unit alignment on at least one speech feature frame respectively to obtain the target word set of the speech data in the target language", and "performing text mapping on at least one speech feature frame respectively to obtain the initial speech recognition text of the speech data in the target language". For example, the above steps can be executed sequentially, or the above steps can be executed in parallel.
[0214] 105. Adjust the initial speech recognition text according to the target phoneme set and the target word set to obtain and output the speech recognition text of the speech data.
[0215] In one embodiment, after obtaining the target phoneme set, the target word set, and the initial speech recognition text, the initial speech recognition text can be adjusted according to the target phoneme set and the target word set to obtain and output the speech recognition text of the speech data. Specifically, the step of "adjusting the initial speech recognition text according to the target phoneme set and the target word set to obtain and output the speech recognition text of the speech data" may include:
[0216] Identify the phoneme information of the initial speech recognition text;
[0217] Use the target phoneme set to adjust the phoneme information to obtain the speech recognition text after phoneme adjustment;
[0218] Use the target word set to adjust the speech recognition text after phoneme adjustment to obtain and output the speech recognition text of the speech data.
[0219] For example, the matching degree between the target phoneme set and the phoneme information of the initial speech recognition text can be calculated. When the matching degree between the two reaches a preset threshold, the initial speech recognition text is used as the speech recognition text after phoneme adjustment. When the matching degree between the two does not reach the preset threshold, the mismatched phonemes in the initial speech recognition text can be replaced with the phonemes in the target phoneme set, so as to obtain the speech recognition text after phoneme adjustment.
[0220] After obtaining the speech recognition text after phoneme adjustment, the speech recognition text after phoneme adjustment can be adjusted by using the target word set, so as to obtain and output the speech recognition text of the voice data. For example, the text similarity between the target word set and the speech recognition text after phoneme adjustment can be calculated by using cosine distance, edit distance, Siamese network, word vector algorithm, etc. When the text similarity between the two reaches a preset threshold, the speech recognition file after phoneme adjustment can be output.
[0221] An embodiment of the present application proposes a speech recognition method, which includes obtaining at least one speech feature frame of voice data in a target language; respectively performing phoneme alignment on the at least one speech feature frame to obtain a target phoneme set of the voice data in the target language; respectively performing word unit alignment on the at least one speech feature frame to obtain a target word set of the voice data in the target language, where the target word set includes word units corresponding to each speech feature frame; respectively performing text mapping on the at least one speech feature frame to obtain an initial speech recognition text of the voice data in the target language; and adjusting the initial speech recognition text according to the target phoneme set and the target word set to obtain and output the speech recognition text of the voice data. In the embodiment of the present application, the speech feature frame can be aligned in two dimensions of phoneme and word unit, and the initial speech recognition text is adjusted by using the target phoneme set and the target word set, so that the matching degree between the speech recognition text and the voice data can be improved, and the accuracy of speech recognition can be improved.
[0222] In addition, an embodiment of the present application also correspondingly proposes a preset speech recognition model, and the preset speech recognition model is trained by using masked phoneme features. Through the masked phoneme features, the preset speech recognition model can use more surrounding information to predict the masked part, thereby improving the robustness of the preset speech recognition model. In addition, the preset speech recognition model includes a phoneme alignment layer and a word unit alignment layer, which can enable the preset speech recognition model to learn more granular alignment messages, thereby improving the performance of the speech recognition model.
[0223] According to the method described in the above embodiments, the following will be further described in detail by way of examples.
[0224] An embodiment of the present application will take the integration of the speech recognition method on a server as an example to introduce the method of the embodiment of the present application.
[0225] In one embodiment, as Figure 9 shown, a speech recognition method has the following specific process:
[0226] 201. The computer device obtains at least one speech feature frame of the speech data in the target language.
[0227] 202. The computer device respectively performs phoneme alignment on at least one speech feature frame to obtain a target phoneme set of the speech data in the target language.
[0228] 203. The computer device respectively performs word unit alignment on at least one speech feature frame to obtain a target word set of the speech data in the target language, where the target word set includes the word unit corresponding to each speech feature frame.
[0229] 204. The computer device respectively performs text mapping on at least one speech feature frame to obtain an initial speech recognition text of the speech data in the target language.
[0230] 205. The computer device adjusts the initial speech recognition text according to the target phoneme set and the target word set, and obtains and outputs the speech recognition text of the speech data.
[0231] In the embodiment of the present application, the computer device can obtain at least one speech feature frame of the speech data in the target language; the computer device can respectively perform phoneme alignment on at least one speech feature frame to obtain a target phoneme set of the speech data in the target language; the computer device can respectively perform word unit alignment on at least one speech feature frame to obtain a target word set of the speech data in the target language, where the target word set includes the word unit corresponding to each speech feature frame; the computer device can respectively perform text mapping on at least one speech feature frame to obtain an initial speech recognition text of the speech data in the target language; the computer device can adjust the initial speech recognition text according to the target phoneme set and the target word set, and obtains and outputs the speech recognition text of the speech data. In the embodiment of the present application, the computer device can perform alignment on the speech feature frame in two dimensions of phoneme and word unit, and use the target phoneme set and the target word set to adjust the initial speech recognition text, so as to improve the matching degree between the speech recognition text and the speech data, and thus improve the accuracy of speech recognition.
[0232] To better implement the speech recognition method provided in the embodiment of the present application, in one embodiment, a speech recognition-based device is further provided, and this speech recognition-based device can be integrated into the computer device. The meanings of the nouns are the same as those in the above speech recognition method, and the specific implementation details can refer to the description in the method embodiment.
[0233] In one embodiment, a voice recognition device is provided. The voice recognition device may be specifically integrated in a computer device, such as Figure 11 As shown, the voice recognition device includes: an acquisition unit 301, a phoneme alignment unit 302, a word unit alignment unit 303, a text mapping unit 304, and an adjustment unit 305, specifically as follows:
[0234] The acquisition unit 301 is configured to acquire at least one voice feature frame of voice data in a target language;
[0235] The phoneme alignment unit 302 is configured to perform phoneme alignment on the at least one voice feature frame respectively to obtain a target phoneme set of the voice data in the target language;
[0236] The word unit alignment unit 303 is configured to perform word unit alignment on the at least one voice feature frame respectively to obtain a target word set of the voice data in the target language, where the target word set includes word units corresponding to each voice feature frame;
[0237] The text mapping unit 304 is configured to perform text mapping on the at least one voice feature frame respectively to obtain an initial voice recognition text of the voice data in the target language;
[0238] The adjustment unit 305 is configured to adjust the initial voice recognition text according to the target phoneme set and the target word set, and obtain and output the voice recognition text of the voice data.
[0239] In one embodiment, the phoneme alignment unit includes:
[0240] A path search subunit, configured to perform path search on each voice feature frame in a preset phoneme search space to obtain at least one phoneme search path;
[0241] A calculation subunit, configured to calculate the cumulative probability of the voice feature frame on each phoneme search path;
[0242] A determination subunit, configured to determine the target phoneme set of the voice data according to the cumulative probability.
[0243] In one embodiment, the path search subunit includes:
[0244] A feature enhancement module, configured to perform feature enhancement on the voice feature frame at the phoneme granularity to obtain the phoneme feature of the voice feature frame;
[0245] A screening module, configured to screen out a target phoneme set from the multiple phoneme sets according to the phoneme feature;
[0246] A phoneme search module, configured to perform phoneme search in the target phoneme set according to the phoneme features, and generate at least one phoneme search path based on the search results.
[0247] In one embodiment, the phoneme search module includes:
[0248] A calculation sub-module, configured to calculate the matching probabilities between the phoneme features and the multiple phoneme nodes respectively;
[0249] A determination sub-module, configured to determine at least one target phoneme node among the multiple phoneme nodes according to the matching probabilities;
[0250] An association sub-module, configured to associate the target phoneme nodes of each phoneme feature to obtain at least one target search path.
[0251] In one embodiment, the word unit alignment unit includes:
[0252] A path search sub-unit, configured to perform path search on each speech feature frame in a preset dictionary search space to obtain at least one word unit search path;
[0253] A calculation sub-unit, configured to calculate the cumulative probabilities of the speech feature frames on each word unit search path;
[0254] A determination sub-unit, configured to determine the target word set of the speech feature frames according to the cumulative probabilities.
[0255] In one embodiment, the text mapping unit includes:
[0256] An attention feature extraction sub-unit, configured to perform attention feature extraction on the speech feature frames in multiple attention dimensions to obtain the attention features of the speech feature frames in each attention dimension;
[0257] A decoding sub-unit, configured to decode the attention features in each attention dimension to obtain the initial speech recognition text of the speech data in the target language.
[0258] In one embodiment, the adjustment unit includes:
[0259] An identification sub-unit, configured to identify the phoneme information of the initial speech recognition text;
[0260] A first adjustment sub-unit, configured to adjust the phoneme information by using the target phoneme set to obtain a phoneme-adjusted speech recognition text;
[0261] A second adjustment sub-unit, configured to adjust the phoneme-adjusted speech recognition text by using the target word set to obtain and output the speech recognition text of the speech data.
[0262] In specific implementation, each of the above units can be implemented as an independent entity, or can be arbitrarily combined and implemented as the same or several entities. For the specific implementation of each of the above units, reference can be made to the foregoing method embodiments, which will not be elaborated herein.
[0263] The above voice recognition-based device can improve the convenience of people taking transportation means.
[0264] The embodiment of the present application further provides a computer device, which may include a terminal or a server. For example, the computer device can be used as a voice recognition-based terminal, and the terminal can be a mobile phone, a tablet computer, etc.; or the computer device can be a server, such as a voice recognition-based server, etc. As Figure 12 shown, it shows a schematic structural diagram of the terminal involved in the embodiment of the present application. Specifically:
[0265] The computer device may include components such as a processor 401 with one or more processing cores, a memory 402 with one or more computer-readable storage media, a power supply 403, and an input unit 404. Those skilled in the art can understand that Figure 12 the structural diagram of the computer device shown in does not constitute a limitation on the computer device, and it may include more or fewer components than shown, or combine certain components, or arrange different components. Among them:
[0266] The processor 401 is the control center of the computer device, connecting various parts of the entire computer device through various interfaces and lines. By running or executing software programs and / or modules stored in the memory 402, and calling data stored in the memory 402, it executes various functions of the computer device and processes data, thereby performing overall detection of the computer device. Optionally, the processor 401 may include one or more processing cores; preferably, the processor 401 may integrate an application processor and a modulation and demodulation processor. Among them, the application processor mainly processes the operating system, user interfaces, and application programs, etc., and the modulation and demodulation processor mainly processes wireless communication. It can be understood that the above modulation and demodulation processor may not be integrated into the processor 401.
[0267] The memory 402 can be used to store software programs and modules. The processor 401 executes various functional applications and data processing by running the software programs and modules stored in the memory 402. The memory 402 mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the computer device. In addition, the memory 402 can include high-speed random access memory, and can also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices. Correspondingly, the memory 402 can also include a memory controller to provide the processor 401 with access to the memory 402.
[0268] The computer device further includes a power supply 403 for powering each component. Preferably, the power supply 403 can be logically connected to the processor 401 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 403 can also include any components such as one or more DC or AC power supplies, a recharge system, a power failure detection circuit, a power converter or inverter, and a power status indicator.
[0269] The computer device may further include an input unit 404, which can be used to receive input digital or character information, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function controls.
[0270] Although not shown, the computer device may further include a display unit, etc., which will not be elaborated here. Specifically, in this embodiment, the processor 401 in the computer device will load the executable files corresponding to the processes of one or more application programs into the memory 402 according to the following instructions, and the processor 401 will run the application programs stored in the memory 402 to realize various functions as follows:
[0271] Obtain at least one speech feature frame of the speech data in the target language;
[0272] Perform phoneme alignment on the at least one speech feature frame respectively to obtain a target phoneme set of the speech data in the target language;
[0273] Perform word unit alignment on the at least one speech feature frame respectively to obtain a target word set of the speech data in the target language, where the target word set includes the word units corresponding to each speech feature frame;
[0274] Perform text mapping on the at least one speech feature frame respectively to obtain the initial speech recognition text of the speech data in the target language;
[0275] Adjust the initial speech recognition text according to the target phoneme set and the target word set, and obtain and output the speech recognition text of the speech data.
[0276] For the specific implementation of each of the above operations, reference may be made to the foregoing embodiments, which will not be elaborated herein.
[0277] According to one aspect of the present application, there is provided a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in various alternative implementations in the foregoing embodiments.
[0278] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the foregoing embodiments can be completed by a computer program, or the relevant hardware can be controlled by a computer program. The computer program can be stored in a computer-readable storage medium and loaded and executed by a processor.
[0279] Therefore, an embodiment of the present application further provides a storage medium, in which a computer program is stored, and the computer program can be loaded by a processor to execute the steps in any speech recognition method provided by the embodiment of the present application. For example, the computer program can execute the following steps:
[0280] Obtain at least one speech feature frame of speech data in the target language;
[0281] Perform phoneme alignment on the at least one speech feature frame respectively to obtain a target phoneme set of the speech data in the target language;
[0282] Perform word unit alignment on the at least one speech feature frame respectively to obtain a target word set of the speech data in the target language, where the target word set includes word units corresponding to each speech feature frame;
[0283] Perform text mapping on the at least one speech feature frame respectively to obtain the initial speech recognition text of the speech data in the target language;
[0284] Adjust the initial speech recognition text according to the target phoneme set and the target word set, and obtain and output the speech recognition text of the speech data.
[0285] Since the computer program stored in the storage medium can execute the steps in any of the speech recognition methods provided by the embodiments of the present application, the beneficial effects achievable by any of the speech recognition methods provided by the embodiments of the present application can be realized. For details, refer to the previous embodiments and will not be elaborated here.
[0286] The above has introduced in detail a speech recognition method, device, computer device, and storage medium provided by the embodiments of the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those skilled in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation on the present application.
Claims
1. A speech recognition method, characterized in that, Including: Obtaining at least one speech feature frame of speech data in a target language; Using a preset speech recognition model to perform phoneme alignment on the at least one speech feature frame to obtain a target phoneme set of the speech data in the target language; Using the preset speech recognition model to perform word unit alignment on the at least one speech feature frame to obtain a target word set of the speech data in the target language, where the target word set includes word units corresponding to each speech feature frame; Using the preset speech recognition model to perform text mapping on the at least one speech feature frame to obtain an initial speech recognition text of the speech data in the target language; Using the preset speech recognition model to adjust the initial speech recognition text according to the target phoneme set and the target word set, and obtaining and outputting the speech recognition text of the speech data; Among them, obtaining the preset speech recognition model includes: Obtaining a plurality of phoneme identification frames and a speech recognition model to be trained; Screening out target phoneme information from the phoneme information of the phoneme identification frames; Performing information conversion processing on the target phoneme information to obtain converted phoneme information; Filling the converted phoneme information into the phoneme identification frames to obtain masked phoneme identification frames; Using the speech recognition model to be trained to perform feature extraction on the masked phoneme identification frames to obtain feature information of the masked phoneme identification frames; Using the speech recognition model to be trained to perform phoneme alignment and word unit alignment on the feature information respectively to obtain target phonemes and target word units of the masked phoneme identification frames; Using the speech recognition model to be trained to perform text mapping on the feature information to obtain the speech recognition text of the masked phoneme identification frames; Performing a joint operation on the target phonemes, the target word units and the speech recognition text to obtain joint loss information; Adjusting model parameters of the speech recognition model to be trained according to the joint loss information to obtain the preset speech recognition model.
2. The method according to claim 1, characterized in that, The performing phoneme alignment on the at least one speech feature frame to obtain a target phoneme set of the speech data in the target language includes: Performing path search on each speech feature frame in a preset phoneme search space to obtain at least one phoneme search path; Calculating the cumulative probability of the speech feature frame on each phoneme search path; Determining the target phoneme set of the speech data according to the cumulative probability.
3. The method according to claim 2, wherein The preset phoneme search space includes a plurality of phoneme sets; the performing path search on each speech feature frame in the preset phoneme search space to obtain at least one phoneme search path includes: Performing feature enhancement on the speech feature frame at the phoneme granularity to obtain the phoneme feature of the speech feature frame; Screening out a target phoneme set from the plurality of phoneme sets according to the phoneme feature; Performing phoneme search in the target phoneme set according to the phoneme feature, and generating at least one phoneme search path according to the search result.
4. The method according to claim 3, wherein The phoneme set includes at least one preset phoneme; the method of performing phoneme search in the target phoneme set according to the phoneme feature and generating at least one phoneme search path includes: Calculating the matching probability between the phoneme feature and the at least one preset phoneme respectively; Determining a target phoneme from the at least one preset phoneme according to the matching probability; Associating the target phonemes of each phoneme feature to obtain a target search path.
5. The method according to claim 1, wherein The method of performing word unit alignment on the at least one speech feature frame to obtain a target word set of the speech data in the target language includes: Performing path search on each speech feature frame in a preset dictionary search space to obtain at least one word unit search path; Calculating the cumulative probability of the speech feature frame on each word unit search path; Determining the target word set of the speech feature frame according to the cumulative probability.
6. The method according to claim 1, characterized in that The method of performing text mapping on the at least one speech feature frame to obtain an initial speech recognition text of the speech data in the target language includes: Performing attention feature extraction on the speech feature frame in multiple attention dimensions to obtain the attention features of the speech feature frame in each attention dimension; Decoding the attention features in each attention dimension to obtain an initial speech recognition text of the speech data in the target language.
7. The method according to claim 1, characterized in that The method of adjusting the initial speech recognition text according to the target phoneme set and the target word set, and obtaining and outputting the speech recognition text of the speech data includes: Identifying the phoneme information of the initial speech recognition text; Adjusting the phoneme information by using the target phoneme set to obtain a phoneme-adjusted speech recognition text; Adjusting the phoneme-adjusted speech recognition text by using the target word set to obtain and output the speech recognition text of the speech data.
8. The method according to claim 1, characterized in that, The method of performing joint operation on the target phoneme, the target word unit and the speech recognition text to obtain joint loss information includes: Calculating the alignment loss information between the target phoneme and the target word unit; Calculating the text loss information between the speech recognition text and a preset identification text; Fusing the alignment loss information and the text loss information to obtain joint loss information.
9. A voice recognition device, characterized in that, Including: An acquisition unit, configured to acquire at least one speech feature frame of speech data in a target language; A phoneme alignment unit, configured to perform phoneme alignment on the at least one speech feature frame by using a preset speech recognition model to obtain a target phoneme set of the speech data in the target language; A word unit alignment unit, configured to perform word unit alignment on the at least one speech feature frame by using the preset speech recognition model to obtain a target word set of the speech data in the target language, where the target word set includes the word unit corresponding to each speech feature frame; A text mapping unit, configured to perform text mapping on the at least one speech feature frame by using the preset speech recognition model to obtain an initial speech recognition text of the speech data in the target language; An adjustment unit, configured to use the preset speech recognition model to adjust the initial speech recognition text according to the target phoneme set and the target word set, and obtain and output the speech recognition text of the speech data; Among them, obtaining the preset speech recognition model includes: Obtaining a plurality of phoneme identification frames and a speech recognition model to be trained; screening out target phoneme information from the phoneme information of the phoneme identification frames; performing information conversion processing on the target phoneme information to obtain converted phoneme information; filling the converted phoneme information into the phoneme identification frames to obtain masked phoneme identification frames; using the speech recognition model to be trained to extract features from the masked phoneme identification frames to obtain the feature information of the masked phoneme identification frames; using the speech recognition model to be trained to perform phoneme alignment and word unit alignment on the feature information respectively to obtain the target phonemes and target word units of the masked phoneme identification frames; using the speech recognition model to be trained to perform text mapping on the feature information to obtain the speech recognition text of the masked phoneme identification frames; performing joint operations on the target phonemes, the target word units, and the speech recognition text to obtain joint loss information; adjusting the model parameters of the speech recognition model to be trained according to the joint loss information to obtain the preset speech recognition model.
10. A computer device, characterized in that, It includes a memory and a processor; the memory stores an application program, and the processor is configured to run the application program in the memory to execute the operations in the speech recognition method according to any one of claims 1 to 8.
11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a plurality of instructions, and the instructions are suitable for being loaded by a processor to execute the steps in the speech recognition method according to any one of claims 1 to 8.
12. A computer program product, characterized in that, The computer program product includes computer instructions, and the computer instructions are stored in a computer-readable storage medium; a processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps in the speech recognition method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Correcting a text recognized by speech recognition through comparison of phonetic sequences in the recognized text with a phonetic transcription of a manually input correction word
CN1555553A
System, method, and program for speech recognition
JP2004302175A