Artificial intelligence-based speech recognition method, apparatus, device, and storage medium
By predicting and fusing the original speech data and the denoised speech data, and selecting the optimal speech data for recognition, the problem of uneven recognition rate of the denoising module under different noise environments is solved, and efficient speech recognition is achieved in both high-noise and low-noise environments.
Patent Information
- Application Number
- CN202210375934.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-11
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2042-04-11
AI Technical Summary
In existing technologies, noise reduction modules cannot achieve a uniform positive effect on speech recognition rate in both high-noise and low-noise environments, and the signal-to-noise ratio judgment is too simplistic and inaccurate, resulting in poor speech recognition performance.
The original speech data is denoised, and the speech recognition effect prediction model is used to predict the speech recognition effect. The speech data to be recognized is determined from the original speech data, denoised speech data and fused speech data according to the target posterior probability, and the optimal speech data is selected for recognition.
It improves the recognition rate and robustness of speech recognition, balances recognition accuracy in both high-noise and low-noise environments, has a wide range of applications, and reduces computational overhead.
Smart Images

Figure CN114822504B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of speech recognition, and in particular to a speech recognition method and device based on artificial intelligence, an equipment and a storage medium. BACKGROUND
[0002] Speech recognition is a multi-disciplinary field closely related to acoustics, phonetics, linguistics, digital signal processing theory, information theory, computer science and many other disciplines. In order to improve the robustness of speech recognition in a noisy environment, the use of a noise reduction module in the front end has been widely applied. However, the introduction of the noise reduction module may cause the recognition rate of speech in a low-noise environment to decrease, thereby having a negative effect. In order to solve this problem, the prior art uses a signal-to-noise ratio (SNR) to make a judgment. If it is a high-noise environment, a noise-reduced speech is used for speech recognition, and if it is a low-noise environment, the speech is directly recognized. This method of using the signal-to-noise ratio to determine whether a noise reduction model needs to be used is too single, hasty, insufficient and imprecise. SUMMARY
[0003] In order to solve the technical problem that the recognition rate of the noise reduction module in the prior art cannot achieve a unified positive effect in a low-noise and high-noise environment, the present application provides a speech recognition method and device based on artificial intelligence, an equipment and a storage medium, which mainly aims to improve the speech recognition rate in a high-noise and low-noise coexisting environment.
[0004] To achieve the above-mentioned purpose, the present application provides a speech recognition method, which comprises the following steps:
[0005] performing noise reduction processing on the obtained original speech data to obtain corresponding noise-reduced speech data;
[0006] inputting the original speech data and the noise-reduced speech data into a trained speech recognition effect prediction model to perform speech recognition effect prediction and obtain a target posterior probability;
[0007] determining the to-be-recognized speech data from the original speech data, the noise-reduced speech data and fused speech data according to the target posterior probability, wherein the fused speech data is obtained by fusing the original speech data and the noise-reduced speech data using the target posterior probability;
[0008] performing speech recognition on the to-be-recognized speech data, and obtaining the target recognition text as the speech recognition result corresponding to the original speech data.
[0009] In addition, to achieve the above-mentioned purpose, the present application further provides a speech recognition device, which comprises a speech noise reduction module, a speech recognition effect prediction module, a speech selection module and a speech recognition module.
[0010] The voice de-noising module is configured to perform de-noising processing on the obtained original voice data to obtain corresponding de-noised voice data, and input the original voice data and the de-noised voice data into the trained voice recognition effect prediction model.
[0011] The voice recognition effect prediction module is configured to perform voice recognition effect prediction based on the trained voice recognition effect prediction model according to the original voice data and the de-noised voice data to obtain a target posterior probability.
[0012] The voice selection module is configured to determine to-be-recognized voice data from the original voice data, the de-noised voice data and fused voice data according to the target posterior probability, wherein the fused voice data is obtained by fusing the original voice data and the de-noised voice data using the target posterior probability.
[0013] The voice recognition module is configured to perform voice recognition on the to-be-recognized voice data, and obtain target recognition text as a voice recognition result corresponding to the original voice data.
[0014] To achieve the above object, the present application further provides a computer device comprising a memory, a processor and computer readable instructions stored in the memory and executable on the processor, wherein the processor executes the steps of the voice recognition method according to any one of the preceding embodiments when executing the computer readable instructions.
[0015] To achieve the above object, the present application further provides a computer readable storage medium, wherein the computer readable storage medium stores computer readable instructions, and the computer readable instructions are executed by a processor to make the processor execute the steps of the voice recognition method according to any one of the preceding embodiments.
[0016] The voice recognition method, device, equipment and storage medium based on artificial intelligence provided by the present application can improve the recognition rate and robustness of voice recognition by predicting the voice recognition effect of original voice data and de-noised voice data, determining to-be-recognized voice data corresponding to the original voice data according to a target posterior probability in the prediction result, and performing voice recognition on the to-be-recognized voice data. Meanwhile, the voice recognition method, device, equipment and storage medium based on artificial intelligence provided by the present application can recognize high-noise voice and low-noise voice, and can ensure the recognition accuracy of voice in high-noise and low-noise environments, and has a wide range of adaptation. The voice recognition system provided by the present application can recognize high-noise voice and low-noise voice. BRIEF DESCRIPTION OF DRAWINGS
[0017] Figure 1 The flowchart of the voice recognition method in an embodiment of the present application;
[0018] Figure 2 The structural block diagram of the voice recognition device in an embodiment of the present application;
[0019] Figure 3An internal structure block diagram of a computer device in an embodiment of the present application.
[0020] The implementation, functional features and advantages of the present application will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION
[0021] To make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative work fall within the protection scope of the present application. It should be understood that the specific embodiments described herein are only used to explain the present application, and are not used to limit the present application.
[0022] Figure 1 A flowchart of a speech recognition method in an embodiment of the present application is shown in FIG. 2. As shown in FIG. 2, the speech recognition method includes the following steps S100-S400. Figure 1
[0023] S100: The obtained original speech data is subjected to noise reduction processing to obtain corresponding noise-reduced speech data.
[0024] Specifically, the speech recognition method is applied to a computer device. The computer device can be, but is not limited to, various servers, personal computers, notebook computers, smart phones, tablet computers and portable wearable devices.
[0025] The original speech data is obtained, and the original speech data is subjected to noise reduction processing to obtain noise-reduced speech data.
[0026] S200: The original speech data and the noise-reduced speech data are input into a trained speech recognition effect prediction model to perform speech recognition effect prediction, to obtain a target posterior probability.
[0027] Specifically, the trained speech recognition effect prediction model is used to perform speech recognition on the original speech data or corresponding noise-reduced speech data, or to perform prediction of the speech recognition effect of the original speech data and corresponding noise-reduced speech data. The posterior probability is used to represent the prediction of the speech recognition effect of the corresponding speech data. The greater the posterior probability, the better the speech recognition effect is likely to be, and the smaller the posterior probability, the worse the speech recognition effect is likely to be.
[0028] The target posterior probability can include a first target posterior probability of the original speech data, and a second target posterior probability of the corresponding noise-reduced speech data can be calculated according to a sum of the first target posterior probability and the second target posterior probability being 1. The target posterior probability can also include the second target posterior probability of the corresponding noise-reduced speech data, and the first target posterior probability of the original speech data can be calculated according to a sum of the first target posterior probability and the second target posterior probability being 1. The target posterior probability can further include the first target posterior probability of the original speech data and the second target posterior probability of the corresponding noise-reduced speech data.
[0029] S300: determining the to-be-recognized speech data from the original speech data, the noise-reduced speech data, and the fused speech data according to the target posterior probability, wherein the fused speech data is obtained by fusing the original speech data and the noise-reduced speech data using the target posterior probability.
[0030] Specifically, the target posterior probability is obtained by performing speech recognition effect prediction on the original speech data before speech recognition.
[0031] Due to the diversity and complexity of speech signals, a speech recognition module can only obtain satisfactory performance under certain limitations, or can only be applied to certain specific occasions. Therefore, the speech recognition model can have different speech recognition effects on the original speech data, the noise-reduced speech data, and the speech data fused from the original speech data and the noise-reduced speech data.
[0032] According to the target posterior probability, it can be determined whether to use the original speech data for final speech recognition, or to use the noise-reduced speech data for final speech recognition, or to calculate the fused speech data according to the target posterior probability and use the fused speech data for final speech recognition.
[0033] S400: performing speech recognition on the to-be-recognized speech data, and obtaining target recognition text as a speech recognition result corresponding to the original speech data.
[0034] Specifically, the speech recognition (Automatic Speech Recognition) technology is to take speech as the research object, and to let the machine automatically recognize and understand the human spoken speech through speech signal processing and pattern recognition. The speech recognition technology is a high technology that lets the machine convert the speech signal into corresponding text or command through the recognition and understanding process. Speech recognition is a very wide cross-discipline, which is closely related to acoustics, phonetics, linguistics, information theory, pattern recognition theory, and neurobiology. In this embodiment, the ASR (Automatic Speech Recognition) technology is used to convert the to-be-recognized speech data into text to obtain the target recognition text.
[0035] If it is determined that the original voice data is the voice data to be recognized, performing voice recognition on the original voice data, and taking the obtained target recognition text as the voice recognition result.
[0036] If it is determined that the noise-reduced voice data is the voice data to be recognized, performing voice recognition on the noise-reduced voice data, and taking the obtained target recognition text as the voice recognition result of the original voice data.
[0037] If it is determined that the fused voice data is the voice data to be recognized, performing voice recognition on the fused voice data, and taking the obtained target recognition text as the voice recognition result of the original voice data.
[0038] The embodiment predicts the voice recognition effect of the original voice data and the noise-reduced voice data, determines the voice data to be recognized corresponding to the original voice data according to the target posterior probability in the prediction result, performs voice recognition on the voice data to be recognized, and improves the recognition rate and robustness of voice recognition. Meanwhile, the recognition of high-noise voice and low-noise voice is taken into account, and the adaptation range is wide.
[0039] In one embodiment, step S200 specifically includes:
[0040] Performing acoustic feature extraction on the original voice data to obtain corresponding first acoustic features, and performing acoustic feature extraction on the noise-reduced voice data to obtain corresponding second acoustic features;
[0041] Performing feature fusion on the first acoustic features and the second acoustic features to obtain first fused features;
[0042] Performing voice recognition effect prediction according to the first fused features to obtain a target posterior probability.
[0043] Specifically, the acoustic features are a sequence of voice features. The acoustic features can be a sequence of MFCC features or a sequence of FBANK features, but are not limited thereto. The first acoustic features and the second acoustic features can be a sequence of features with a dimension of 128, but are not limited thereto, and can be defined according to actual conditions.
[0044] The feature fusion is feature splicing. For example, the first acoustic features and the second acoustic features are a sequence of features with a dimension of 128, and the first fused features obtained after fusion are a sequence of features with a dimension of 256.
[0045] The trained voice recognition effect prediction model performs voice recognition effect prediction according to the first fused features, and the target posterior probability is obtained.
[0046] In one embodiment, before step S200, the method further includes:
[0047] Obtaining estimated noise data of the original voice data,
[0048] The signal-to-noise ratio corresponding to the original speech data is calculated according to the noise-reduced speech data and the estimated noise data, and the signal-to-noise ratio is input into the trained speech recognition effect prediction model;
[0049] The original speech data and the noise-reduced speech data are input into the trained speech recognition effect prediction model for speech recognition effect prediction to obtain a target posterior probability, including:
[0050] The original speech data is subjected to acoustic feature extraction to obtain corresponding first acoustic features, and the noise-reduced speech data is subjected to acoustic feature extraction to obtain corresponding second acoustic features,
[0051] The first acoustic features and the second acoustic features are subjected to first feature fusion to obtain first fusion features,
[0052] The first speech recognition effect prediction is performed according to the first fusion features to obtain an intermediate posterior probability,
[0053] The intermediate posterior probability is subjected to second feature fusion with the signal-to-noise ratio as a second intermediate layer feature to obtain second fusion features,
[0054] The second speech recognition effect prediction is performed according to the second fusion features to obtain the target posterior probability.
[0055] Specifically, the estimated noise data of the original speech data is separated from the original speech data by a speech noise reduction module or estimated from the original speech data in a speech noise reduction process.
[0056] The signal-to-noise ratio of the original speech data is input into the trained speech recognition effect prediction model together with the original speech data and the noise-reduced speech data.
[0057] The acoustic features are a sequence of speech features. The acoustic features can be a sequence of MFCC features or a sequence of FBANK features, but are not limited thereto. The first acoustic features and the second acoustic features can be a sequence of features with a dimension of 128, but are not limited thereto, and can be defined according to actual conditions.
[0058] The trained speech recognition effect prediction model sequentially includes a first feature fusion layer, a 2-layer LSTM model, a first full connection layer, a second feature fusion layer, and a second full connection layer. The output layer of the first full connection layer and the second full connection layer uses a softmax layer, and the activation function after the hidden layer can use a ReLU function.
[0059] The feature fusion, i.e., feature splicing, for example, the first acoustic feature and the second acoustic feature are 128-dimensional feature sequences, and the first fused feature obtained after fusion is a 256-dimensional feature sequence. The first feature fusion layer is configured to perform first feature fusion on the first acoustic feature and the second acoustic feature to obtain the first fused feature. The first fused feature is input into the LSTM model, and the first full connection layer performs first speech recognition effect prediction according to the output of the LSTM model to obtain the intermediate posterior probability. The intermediate posterior probability includes a first intermediate posterior probability of the original speech data and a second intermediate posterior probability of the noise-reduced speech data, and the sum of the first intermediate posterior probability and the second intermediate posterior probability is 1.
[0060] The first full connection layer transmits the intermediate posterior probability to the second feature fusion layer, and the second feature fusion layer performs second feature fusion on the intermediate posterior probability as the first intermediate layer feature and the signal-to-noise ratio of the original speech data as the second intermediate layer feature to obtain the second fused feature. The second connection layer performs speech recognition effect prediction according to the second fused feature, and the target posterior probability is obtained.
[0061] In this embodiment, the original speech data, the noise-reduced speech data, and the signal-to-noise ratio of the original speech data and the noise-reduced speech data are combined to predict the speech recognition effect, and the prediction effect is more accurate, so that the to-be-recognized speech data can be determined more accurately, and the recognition rate or recognition effect of the original speech data is improved.
[0062] In one embodiment, the target posterior probability includes a first target posterior probability and a second target posterior probability, the first target posterior probability represents the recognition effect of the original speech data, the second target posterior probability represents the recognition effect of the noise-reduced speech data, and the sum of the first target posterior probability and the second target posterior probability is 1.
[0063] The step S300 specifically includes:
[0064] If the first target posterior probability is greater than the second target posterior probability, the original speech data is determined as the to-be-recognized speech data, and if the first target posterior probability is less than the second target posterior probability, the noise-reduced speech data is determined as the to-be-recognized speech data of the original speech data.
[0065] Alternatively, the original speech data and the noise-reduced speech data are fused according to the first target posterior probability and the second target posterior probability, and the fused speech data is used as the to-be-recognized speech data corresponding to the original speech data.
[0066] Specifically, if the first target posterior probability of the original speech data is greater than the second target posterior probability of the noise-reduced speech data, it indicates that the predicted recognition effect of the original speech data is better than that of the noise-reduced speech data, and therefore, the original speech data is selected as the to-be-recognized speech data.
[0067] If the first target posterior probability of the original speech data is less than the second target posterior probability of the denoised speech data, it indicates that the prediction recognition effect on the original speech data is not as good as the prediction recognition effect on the denoised speech data, and thus the denoised speech data is selected as the to-be-recognized speech data.
[0068] If the first target posterior probability of the original speech data is equal to the second target posterior probability of the denoised speech data, it indicates that the prediction recognition effect on the original speech data is the same as the prediction recognition effect on the denoised speech data, and thus the original speech data or the denoised speech data can be selected as the to-be-recognized speech data. However, the denoised speech data is more optimal to be selected because the data processing amount of the denoised speech data is smaller in speech recognition.
[0069] In another specific embodiment, regardless of the size of the first target posterior probability of the original speech data and the second target posterior probability of the denoised speech data, the original speech data and the denoised speech data are fused, and the fused speech data obtained after the fusion is selected as the to-be-recognized speech data.
[0070] The fused speech data = the first target posterior probability * the original speech data + the second target posterior probability * the denoised speech data, and the specific formula is shown in formula (1):
[0071]
[0072] Wherein, is the speech data after the fusion, is the denoised speech data and Y is the original speech data. p0 is the first target posterior probability, and 1-p0 is the second target posterior probability.
[0073] The speech recognition method of the present application is applied to a speech recognition system, and the speech recognition system includes a speech denoising module, a speech recognition effect prediction module in which a trained or to-be-trained speech recognition effect prediction model is deployed, a speech selection module, and a speech recognition module. The speech selection module specifically includes an original denoising selection module and / or a speech fusion module for selecting denoised speech and original speech. The present application realizes that a data-driven neural network model better determines whether a denoising module will have a positive effect on a speech recognition module, thereby improving the overall recognition rate of the system in a high-noise and low-noise coexisting environment.
[0074] In one embodiment, before step S200, the method further includes:
[0075] Obtaining different known speech segments and corresponding denoised speech segments;
[0076] generate a data label corresponding to each known speech segment, and label the corresponding training sample according to the data label, wherein each training sample includes a known speech segment and a corresponding noise-reduced speech segment, the data label includes a first posterior probability and a second posterior probability, the first posterior probability represents the recognition effect on the corresponding known speech segment, the second posterior probability represents the recognition effect on the corresponding noise-reduced speech segment of the known speech segment, and the sum of the first posterior probability and the second posterior probability is 1;
[0077] train the pre-trained speech recognition effect prediction model using the labeled training sample until a convergence condition is met, to obtain a trained speech recognition effect prediction model.
[0078] Specifically, the known speech segment is a speech segment whose true speech recognition text is known, and the recognition effect of the known speech segment and its noise-reduced speech segment is also known. The data label is a representation of the known recognition effect of a group of known speech segments and corresponding noise-reduced speech segments.
[0079] A training sample includes a known speech segment and a corresponding noise-reduced speech segment, and the corresponding training sample is labeled using a data label to obtain a labeled training sample. All labeled training samples form a training set.
[0080] The pre-trained speech recognition effect prediction model is trained using the training set, and the training is stopped when the loss function (for example, a cross-entropy loss function, but not limited thereto) is reduced to a threshold value or the number of training reaches a preset number of training. The pre-trained speech recognition effect prediction model is constructed using the model parameters when the convergence condition is met to obtain a trained speech recognition effect prediction model.
[0081] In one embodiment, before step S200, the method further includes:
[0082] obtaining different known speech segments, corresponding noise-reduced speech segments, and signal-to-noise ratios;
[0083] generate a data label corresponding to each known speech segment, and label the corresponding training sample according to the data label, wherein each training sample includes a known speech segment and a corresponding noise-reduced speech segment and a signal-to-noise ratio, the data label includes a first posterior probability and a second posterior probability, the first posterior probability represents the recognition effect on the corresponding known speech segment, the second posterior probability represents the recognition effect on the corresponding noise-reduced speech segment of the known speech segment, and the sum of the first posterior probability and the second posterior probability is 1;
[0084] train the pre-trained speech recognition effect prediction model using the labeled training sample until a convergence condition is met, to obtain a trained speech recognition effect prediction model.
[0085] Specifically, the known speech segment is a speech segment whose real speech recognition text is known, and the known speech segment and its recognition effect of the noise-reduced speech segment speech recognition are also known. The data label is a representation of the known recognition effect of a group of known speech segments and corresponding noise-reduced speech segments.
[0086] The calculation formula of the signal-to-noise ratio is shown in formula (2):
[0087]
[0088] wherein, the noise-reduced speech segment, the noise segment estimated according to the known speech segment.
[0089] A training sample includes a known speech segment, a corresponding noise-reduced speech segment, and a signal-to-noise ratio. The corresponding training sample is labeled using the data label to obtain a labeled training sample. All labeled training samples form a training set.
[0090] The pre-trained speech recognition effect prediction model is trained using the training set, that is, the loss function and the gradient are calculated, the model parameters are updated according to the gradient, and then a new pre-trained speech recognition effect prediction model is constructed using the updated model parameters. The new pre-trained speech recognition effect prediction model is used to
[0091] The training is stopped when the loss function is reduced to a threshold or the number of training reaches a preset number of training. The model parameters when the convergence condition is reached are used to construct the pre-trained speech recognition effect prediction model to obtain a trained speech recognition effect prediction model.
[0092] In one embodiment, the data label corresponding to each known speech segment is generated, including:
[0093] The actual speech text of the known speech segment is obtained;
[0094] The known speech segment is subjected to speech recognition to obtain a first recognition text, and the noise-reduced speech segment is subjected to speech recognition to obtain a second recognition text;
[0095] The similarity between the actual speech text and the first recognition text is calculated to obtain a first similarity, and the similarity between the actual speech text and the second recognition text is calculated to obtain a second similarity;
[0096] The first posterior probability of the known speech segment and the second posterior probability of the noise-reduced speech segment are determined according to the first similarity and the second similarity;
[0097] The first posterior probability and the second posterior probability form the data label.
[0098] Specifically, the actual speech text of the known speech segment is the real text corresponding to the speech in the known speech segment, and the actual speech text can be artificially recognized and provided to the computer device. The speech recognition is performed on the known speech segment and the corresponding noise-reduced speech segment respectively to obtain a first recognized text and a second recognized text.
[0099] The first recognized text and the second recognized text can be the same as or different from the actual speech text as the real text. Therefore, the first similarity between the first recognized text and the actual speech text and the second similarity between the second recognized text and the actual speech text need to be calculated, and the first similarity and the second similarity represent the difference between the first recognized text and the actual speech text and the difference between the second recognized text and the actual speech text. According to the first similarity and the second similarity, the recognition effect of the known speech segment and the noise-reduced speech segment can be determined, that is, the first posterior probability and the second posterior probability are obtained. The higher the similarity, the greater the corresponding posterior probability.
[0100] The similarity can be obtained by calculating the edit distance between the two texts.
[0101] In one embodiment, determining the first posterior probability of the known speech segment and the second posterior probability of the noise-reduced speech segment according to the first similarity and the second similarity comprises:
[0102] If the first similarity is greater than the second similarity, the first posterior probability of the known speech segment is determined as 1, and the second posterior probability of the noise-reduced speech segment is determined as 0.
[0103] If the first similarity is less than or equal to the second similarity, the first posterior probability of the known speech segment is determined as 0, and the second posterior probability of the noise-reduced speech segment is determined as 1.
[0104] Specifically, the posterior probability of the present embodiment has only two values of 1 and 0, which simplifies the training complexity. Even if the fused speech data is used, the essence is to select the speech data with a posterior probability of 1 as the to-be-recognized speech data. The present embodiment reduces the operation overhead.
[0105] In one embodiment, determining the first posterior probability of the known speech segment and the second posterior probability of the noise-reduced speech segment according to the first similarity and the second similarity comprises:
[0106] Calculating the sum of the first similarity and the second similarity to obtain a similarity sum;
[0107] Taking the ratio of the first similarity to the similarity sum as the first posterior probability of the known speech segment;
[0108] Taking the ratio of the second similarity to the similarity sum as the second posterior probability of the noise-reduced speech segment.
[0109] Specifically, the embodiment determines the posterior probability by the ratio of the similarity, can represent that the higher the similarity is, the greater the posterior probability is, and can ensure that the sum of the first posterior probability and the second posterior probability is 1.
[0110] In addition, the embodiment realizes the diversification and refinement of the posterior probability compared to the two values of 1 and 0 of the posterior probability. The data label obtained by the embodiment can more accurately represent the speech recognition effect of the original noise-reduced speech data and the noise-reduced noise-reduced speech data. For model training, the posterior probability prediction result of the pre-trained speech recognition effect prediction model can be more accurate.
[0111] In one embodiment, the similarity between the actual speech text and the first recognized text is calculated to obtain a first similarity, and the similarity between the actual speech text and the second recognized text is calculated to obtain a second similarity, comprising:
[0112] The edit distance between the actual speech text and the first recognized text is calculated to obtain a first edit distance, and the edit distance between the actual speech text and the second recognized text is calculated to obtain a second edit distance.
[0113] The first similarity between the actual speech text and the first recognized text is obtained according to the first edit distance, and the second similarity between the actual speech text and the second recognized text is obtained according to the second edit distance.
[0114] Specifically, the embodiment determines the similarity between texts by edit distance. The greater the edit distance is, the lower the similarity is, and the smaller the edit distance is, the higher the similarity is.
[0115] Edit distance (Edit Distance) is also called Levenshtein distance, which refers to the minimum number of editing operations required to convert one string to another. In the fields of information theory, linguistics and computer science, Levenshtein distance is used to measure the similarity between two sequences. The permitted editing operations include replacing one character with another, inserting a character, and deleting a character. Generally speaking, the smaller the edit distance is, the greater the similarity of the two strings is.
[0116] Edit distance similarity = 1 - edit distance / max (string 1 length, string 2 length).
[0117] Taking the values of the first posterior probability and the second posterior probability as 0 or 1 as an example, the model data label (p 00 ,p 11 ) is generated according to formula (3):
[0118]
[0119]
[0120] wherein, W represents the actual speech text of the known speech segment, W Y represents the first recognition text of speech recognition on the known speech segment (the original noisy speech segment), represents the second recognition result of speech recognition on the noise-reduced speech segment obtained after the denoising module or denoising model, dist(*) represents the edit distance between two texts, i.e., dist(W,W Y ) is the first edit distance between the actual speech text and the first recognition text, is the second edit distance between the actual speech text and the second recognition text.p 00 is the first posterior probability corresponding to the known speech segment, 11 is the second posterior probability corresponding to the noise-reduced speech segment corresponding to the known speech segment.
[0121] Since the similarity is inversely proportional to the edit distance, the model data label (p 00 , p 11 ) can also be generated according to formula (4):
[0122] (p 00 , p 11 ) = [1, 0], the first similarity > the second similarity
[0123] (p 00 , p 11 ) = [0, 1], the first similarity ≤ the second similarity formula (4)
[0124] wherein, p 00 is the first posterior probability corresponding to the known speech segment, p 11 is the second posterior probability corresponding to the noise-reduced speech segment corresponding to the known speech segment.
[0125] The embodiments of the present application can acquire and process related data based on artificial intelligence technology to realize speech recognition. Among them, artificial intelligence (Artificial Intelligence, AI) is to use digital computers or digital computer controlled machines to simulate, extend and expand human intelligence, perceive environment, acquire knowledge and use knowledge to obtain the best results. Theory, method, technology and application system.
[0126] The basic technology of artificial intelligence generally includes technologies such as sensors, special artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction system, mechatronics, etc. Artificial intelligence software technology mainly includes computer vision technology, robot technology, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning, etc. Several major directions.
[0127] The application uses a neural network model to determine whether the noise reduction module has a positive effect on the speech recognition model, and can fuse the noise-reduced speech and the original speech using the output probability or select the original speech or the noise-reduced speech as the to-be-recognized speech according to the output probability, thereby improving the recognition rate of the overall speech recognition system. The introduction of the selection model does not require joint training or fine-tuning training of the ASR model (speech recognition model) and the noise reduction model, saves development costs, maintains the independence of each module, and facilitates maintenance.
[0128] Figure 2 FIG. 1 is a structural block diagram of a speech recognition device according to an embodiment of the present application. Referring to FIG. 1, the device includes a speech denoising module 100, a speech recognition effect prediction module 200, a speech selection module 300, and a speech recognition module 400. Figure 2
[0129] The speech denoising module 100 is configured to perform noise reduction processing on the obtained original speech data to obtain corresponding noise-reduced speech data, and input the original speech data and the noise-reduced speech data to the trained speech recognition effect prediction model.
[0130] The speech recognition effect prediction module 200 is configured to perform speech recognition effect prediction based on the trained speech recognition effect prediction model according to the original speech data and the noise-reduced speech data to obtain a target posterior probability.
[0131] The speech selection module 300 is configured to determine to-be-recognized speech data from the original speech data, the noise-reduced speech data, and fused speech data according to the target posterior probability, wherein the fused speech data is obtained by fusing the original speech data and the noise-reduced speech data using the target posterior probability.
[0132] The speech recognition module 400 is configured to perform speech recognition on the to-be-recognized speech data and obtain a target recognition text as a speech recognition result corresponding to the original speech data.
[0133] The speech recognition device is generally provided in a server / terminal device.
[0134] In an embodiment, the speech recognition effect prediction module 200 includes:
[0135] The feature extraction module is configured to perform acoustic feature extraction on the original speech data to obtain corresponding first acoustic features and perform acoustic feature extraction on the noise-reduced speech data to obtain corresponding second acoustic features.
[0136] The first feature fusion module is configured to perform feature fusion on the first acoustic features and the second acoustic features to obtain first fused features.
[0137] The first prediction module is configured to perform speech recognition effect prediction according to the first fused feature to obtain a target posterior probability.
[0138] In one embodiment, the apparatus further comprises:
[0139] The noise data obtaining module is configured to obtain estimated noise data of the original speech data,
[0140] The signal-to-noise ratio calculating module is configured to calculate a signal-to-noise ratio corresponding to the original speech data according to the de-noised speech data and the estimated noise data, and input the signal-to-noise ratio into the trained speech recognition effect prediction model.
[0141] The speech recognition effect prediction module 200 comprises:
[0142] The feature extraction module is configured to perform acoustic feature extraction on the original speech data to obtain corresponding first acoustic features, and perform acoustic feature extraction on the de-noised speech data to obtain corresponding second acoustic features,
[0143] The first feature fusion module is configured to perform first feature fusion on the first acoustic features and the second acoustic features to obtain first fused features,
[0144] The first prediction module is configured to perform first speech recognition effect prediction according to the first fused features to obtain an intermediate posterior probability,
[0145] The second feature fusion module is configured to perform second feature fusion on the intermediate posterior probability as a first intermediate layer feature and the signal-to-noise ratio as a second intermediate layer feature to obtain second fused features,
[0146] The second prediction module is configured to perform second speech recognition effect prediction according to the second fused features to obtain a target posterior probability.
[0147] In one embodiment, the target posterior probability comprises a first target posterior probability and a second target posterior probability, the first target posterior probability representing a recognition effect on the original speech data, and the second target posterior probability representing a recognition effect on the de-noised speech data, and the sum of the first target posterior probability and the second target posterior probability is 1.
[0148] The speech selection module 300 specifically comprises:
[0149] The original de-noised selection module is configured to determine the original speech data as the to-be-recognized speech data if the first target posterior probability is greater than the second target posterior probability, and determine the de-noised speech data as the to-be-recognized speech data of the original speech data if the first target posterior probability is less than the second target posterior probability.
[0150] Or,
[0151] a voice fusion module configured to fuse the original voice data and the noise-reduced voice data according to the first target posterior probability and the second target posterior probability, and take the fused voice data as the to-be-recognized voice data corresponding to the original voice data.
[0152] In one embodiment, the apparatus further comprises:
[0153] a sample voice acquisition module configured to acquire different known voice segments and corresponding noise-reduced voice segments;
[0154] a label generation module configured to generate a data label corresponding to each known voice segment, and mark corresponding training samples according to the data label, wherein each training sample comprises a known voice segment and a corresponding noise-reduced voice segment, the data label comprises a first posterior probability and a second posterior probability, the first posterior probability represents an identification effect on the corresponding known voice segment, the second posterior probability represents an identification effect on a noise-reduced voice segment corresponding to the known voice segment, and the sum of the first posterior probability and the second posterior probability is 1.
[0155] a training module configured to train a pre-trained voice identification effect prediction model using the marked training samples until a convergence condition is met, and obtain a trained voice identification effect prediction model.
[0156] In one embodiment, the apparatus further comprises:
[0157] a sample voice acquisition and calculation module configured to acquire different known voice segments, corresponding noise-reduced voice segments, and signal-to-noise ratios;
[0158] a label generation module configured to generate a data label corresponding to each known voice segment, and mark corresponding training samples according to the data label, wherein each training sample comprises a known voice segment, a corresponding noise-reduced voice segment, and a signal-to-noise ratio, the data label comprises a first posterior probability and a second posterior probability, the first posterior probability represents an identification effect on the corresponding known voice segment, the second posterior probability represents an identification effect on a noise-reduced voice segment corresponding to the known voice segment, and the sum of the first posterior probability and the second posterior probability is 1.
[0159] a training module configured to train a pre-trained voice identification effect prediction model using the marked training samples until a convergence condition is met, and obtain a trained voice identification effect prediction model.
[0160] In one embodiment, the label generation module specifically comprises:
[0161] a text acquisition module configured to acquire an actual voice text of the known voice segment;
[0162] The voice recognition module is further configured to perform voice recognition on the known voice segment to obtain a first recognized text, and perform voice recognition on the noise-reduced voice segment to obtain a second recognized text.
[0163] The similarity calculation module is configured to calculate a first similarity between the actual voice text and the first recognized text to obtain a first similarity, and calculate a second similarity between the actual voice text and the second recognized text to obtain a second similarity.
[0164] The posterior probability determination module is configured to determine a first posterior probability of the known voice segment and a second posterior probability of the noise-reduced voice segment according to the first similarity and the second similarity.
[0165] The label combination module is configured to combine the first posterior probability and the second posterior probability to obtain the data label.
[0166] In one embodiment, the posterior probability determination module is specifically configured to determine the first posterior probability of the known voice segment as 1 and the second posterior probability of the noise-reduced voice segment as 0 if the first similarity is greater than the second similarity, or determine the first posterior probability of the known voice segment as 0 and the second posterior probability of the noise-reduced voice segment as 1 if the first similarity is less than or equal to the second similarity.
[0167] In one embodiment, the posterior probability determination module specifically includes:
[0168] The sum module is configured to calculate a sum of the first similarity and the second similarity to obtain a sum of similarities.
[0169] The first proportion calculation module is configured to take a ratio of the first similarity to the sum of similarities as the first posterior probability of the known voice segment.
[0170] The second proportion calculation module is configured to take a ratio of the second similarity to the sum of similarities as the second posterior probability of the noise-reduced voice segment.
[0171] In one embodiment, the similarity calculation module specifically includes:
[0172] The edit distance calculation unit is configured to calculate an edit distance between the actual voice text and the first recognized text to obtain a first edit distance, and calculate an edit distance between the actual voice text and the second recognized text to obtain a second edit distance.
[0173] The similarity calculation unit is configured to obtain the first similarity between the actual voice text and the first recognized text according to the first edit distance, and obtain the second similarity between the actual voice text and the second recognized text according to the second edit distance.
[0174] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0175] The terms "first" and "second" in the above-mentioned modules / units are only used to distinguish different modules / units and are not intended to specify which module / unit has a higher priority or any other limiting meaning. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or modules is not necessarily limited to those steps or modules explicitly listed, but may include other steps or modules not explicitly listed or inherent to these processes, methods, products, or devices. The module divisions appearing in this application are merely logical divisions; in actual applications, different division methods may be used.
[0176] For specific limitations regarding the speech recognition device, please refer to the limitations on the speech recognition method above, which will not be repeated here. Each module in the aforementioned speech recognition device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in the computer device in hardware form, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each module.
[0177] Figure 3 This is a block diagram of the internal structure of a computer device according to an embodiment of this application. Figure 3 As shown, the computer device includes a processor, memory, network interface, input device, and display screen connected via a system bus. The processor provides computing and control capabilities. The memory includes storage media and internal memory. The storage media can be non-volatile or volatile. The storage media stores the operating system and may also store computer-readable instructions, which, when executed by the processor, enable the processor to implement a speech recognition method. The internal memory provides an environment for the operation of the operating system and computer-readable instructions in the storage media. The internal memory may also store computer-readable instructions, which, when executed by the processor, enable the processor to perform a speech recognition method. The network interface of the computer device is used for communication with an external server via a network connection. The display screen of the computer device can be an LCD screen or an e-ink screen. The input device of the computer device can be a touch layer covering the display screen, buttons, a trackball, or a touchpad located on the computer device's casing, or an external keyboard, touchpad, or mouse, etc.
[0178] In an embodiment, a computer device is provided, which includes a memory, a processor, and computer readable instructions (e.g., a computer program) stored in the memory and executable on the processor, the processor implements the steps of the voice recognition method in the above embodiments when executing the computer readable instructions, for example Figure 1 the functions of the modules 100-400 shown in the above embodiments when executing the computer readable instructions, for example Figure 2 the functions of the modules 100-400 shown in the above embodiments when executing the computer readable instructions, for example
[0179] The processor can be a central processing unit (CPU), and can also be other general-purpose processors, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor. The processor is a control center of the computer device, and connects all parts of the computer device through various interfaces and lines.
[0180] The memory can be used to store computer readable instructions and / or modules, and the processor implements various functions of the computer device by running or executing the computer readable instructions and / or modules stored in the memory, and calling data stored in the memory. The memory can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, at least one application program required by a function (such as a sound playing function, an image playing function, etc.), etc.; and the data storage area can store data created according to the use of the mobile phone (such as audio data, video data, etc.), etc.
[0181] The memory can be integrated in the processor, or can be separately arranged from the processor.
[0182] Those skilled in the art can understand that Figure 3 The structure shown in the above embodiments is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device to which the scheme of the present application is applied. The specific computer device can include more or less components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0183] In an embodiment, a computer readable storage medium having computer readable instructions stored thereon is provided, the computer readable instructions, when executed by a processor, implement the steps of the speech recognition method in the above embodiments, for example Figure 1 the steps S100 to S400 and extensions of other extensions and related steps of the method. Alternatively, the computer readable instructions, when executed by a processor, implement the functions of the modules / units of the speech recognition apparatus in the above embodiments, for example Figure 2 the modules 100 to 400. To avoid repetition, no further elaboration is made here.
[0184] It is understood by those skilled in the art that all or part of the processes of the above embodiments can be implemented by computer readable instructions instructing the relevant hardware, and the computer readable instructions can be stored in a computer readable storage medium, and when executed, can include the processes of the above embodiments. Any reference to memory, storage, database or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0185] It should be noted that in this document, the terms "comprising", "including", or any other variant thereof are intended to cover non-exclusive inclusion, so that processes, devices, articles or methods including a series of elements not only include those elements, but also include other elements not explicitly listed, or inherent to such processes, devices, articles or methods. Without more limitations, the element defined by the statement "comprising a" does not exclude the presence of other identical elements in the process, device, article or method including the element.
[0186] The above application embodiment serial numbers are only for description, and do not represent the advantages and disadvantages of the embodiments. Through the above description of the embodiments, those skilled in the art can clearly understand that the above embodiment methods can be realized by means of software and the necessary general hardware platform, and of course, they can also be realized by hardware, but in many cases, the former is a better embodiment. Based on this understanding, the technical solutions of the present application can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disc, optical disc) as described above, and includes a plurality of instructions for making a terminal device (which can be a mobile phone, computer, server, or network device, etc.) execute the methods described in various embodiments of the present application.
[0187] The above is only the preferred embodiment of the present application, and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation using the content of the present application specification and drawings, or direct or indirect application in other related technical fields, is also included in the patent protection scope of the present application.
Claims
1. A speech recognition method, characterized in that, The method includes: The acquired raw speech data is subjected to noise reduction processing to obtain the corresponding noise-reduced speech data; The original speech data and the denoised speech data are input into the trained speech recognition effect prediction model to predict the speech recognition effect and obtain the target posterior probability. The speech data to be identified is determined from the original speech data, the denoised speech data, and the fused speech data according to the target posterior probability, wherein the fused speech data is obtained by fusing the original speech data and the denoised speech data using the target posterior probability; The speech data to be recognized is subjected to speech recognition, and the resulting target recognized text is used as the speech recognition result corresponding to the original speech data. The step of inputting the original speech data and the denoised speech data into a trained speech recognition performance prediction model to predict the speech recognition performance and obtain the target posterior probability includes: The original speech data is subjected to acoustic feature extraction to obtain the corresponding first acoustic feature, and the denoised speech data is subjected to acoustic feature extraction to obtain the corresponding second acoustic feature; The first acoustic feature and the second acoustic feature are fused to obtain the first fused feature; Based on the first fused feature, a speech recognition performance prediction is performed to obtain the target posterior probability; The target posterior probability includes a first target posterior probability and a second target posterior probability. The first target posterior probability represents the recognition effect on the original speech data, and the second target posterior probability represents the recognition effect on the denoised speech data. The sum of the first target posterior probability and the second target posterior probability is 1. The step of determining the speech data to be recognized from the original speech data, the denoised speech data, and the fused speech data based on the target posterior probability includes: If the first target posterior probability is greater than the second target posterior probability, then the original speech data is determined to be speech data to be recognized; if the first target posterior probability is less than the second target posterior probability, then the denoised speech data is determined to be speech data to be recognized from the original speech data. Alternatively, the original speech data and the denoised speech data can be fused based on the first target posterior probability and the second target posterior probability, and the fused speech data can be used as the speech data to be recognized corresponding to the original speech data.
2. The method according to claim 1, characterized in that, Before inputting the original speech data and the denoised speech data into the trained speech recognition performance prediction model to predict the speech recognition performance and obtain the target posterior probability, the method further includes: Obtain the estimated noise data from the original speech data. The signal-to-noise ratio (SNR) corresponding to the original speech data is calculated based on the denoised speech data and the estimated noise data, and the SNR is input into the trained speech recognition performance prediction model. The step of inputting the original speech data and the denoised speech data into a trained speech recognition performance prediction model to predict the speech recognition performance and obtain the target posterior probability includes: Acoustic feature extraction is performed on the original speech data to obtain the corresponding first acoustic feature, and acoustic feature extraction is performed on the denoised speech data to obtain the corresponding second acoustic feature. The first acoustic feature and the second acoustic feature are fused to obtain the first fused feature. Based on the first fused features, a first speech recognition performance prediction is performed to obtain the intermediate posterior probability. The intermediate posterior probability, used as a first intermediate layer feature, is fused with the signal-to-noise ratio, used as a second intermediate layer feature, to obtain the second fused feature. Based on the second fusion feature, a second speech recognition effect prediction is performed to obtain the target posterior probability.
3. The method according to claim 1, characterized in that, Before inputting the original speech data and the denoised speech data into the trained speech recognition performance prediction model to predict the speech recognition performance and obtain the target posterior probability, the method further includes: Acquire different known speech segments and their corresponding noise-reduced speech segments; Generate data labels for each known speech segment, and label the corresponding training samples according to the data labels. Each training sample includes a known speech segment and a corresponding denoised speech segment. The data labels include a first posterior probability and a second posterior probability. The first posterior probability represents the recognition effect of the corresponding known speech segment, and the second posterior probability represents the recognition effect of the denoised speech segment of the corresponding known speech segment. The sum of the first posterior probability and the second posterior probability is 1. The pre-trained speech recognition performance prediction model is trained using labeled training samples until the convergence condition is met, thus obtaining the trained speech recognition performance prediction model.
4. The method according to claim 2, characterized in that, Before inputting the original speech data and the denoised speech data into the trained speech recognition performance prediction model to predict the speech recognition performance and obtain the target posterior probability, the method further includes: Acquire different known speech segments and their corresponding denoised speech segments and signal-to-noise ratios; Generate data labels for each known speech segment, and label the corresponding training samples according to the data labels. Each training sample includes a known speech segment, a corresponding denoised speech segment, and a signal-to-noise ratio. The data labels include a first posterior probability and a second posterior probability. The first posterior probability represents the recognition effect of the corresponding known speech segment, and the second posterior probability represents the recognition effect of the denoised speech segment of the corresponding known speech segment. The sum of the first posterior probability and the second posterior probability is 1. The pre-trained speech recognition performance prediction model is trained using labeled training samples until the convergence condition is met, thus obtaining the trained speech recognition performance prediction model.
5. The method according to claim 3 or 4, characterized in that, The generation of data tags corresponding to each known speech segment includes: Obtain the actual speech text of the known speech segment; The known speech segment is subjected to speech recognition to obtain a first recognized text, and the denoised speech segment is subjected to speech recognition to obtain a second recognized text; The similarity between the actual speech text and the first recognized text is calculated to obtain a first similarity, and the similarity between the actual speech text and the second recognized text is calculated to obtain a second similarity. The first posterior probability of the known speech segment and the second posterior probability of the denoised speech segment are determined based on the first similarity and the second similarity. The first posterior probability and the second posterior probability are combined to form the data label.
6. The method according to claim 5, characterized in that, Determining the first posterior probability of the known speech segment and the second posterior probability of the denoised speech segment based on the first similarity and the second similarity includes: If the first similarity is greater than the second similarity, then the first posterior probability of the known speech segment is determined to be 1 and the second posterior probability of the denoised speech segment is determined to be 0. If the first similarity is less than or equal to the second similarity, then the first posterior probability of the known speech segment is determined to be 0 and the second posterior probability of the denoised speech segment is determined to be 1.
7. The method according to claim 5, characterized in that, Determining the first posterior probability of the known speech segment and the second posterior probability of the denoised speech segment based on the first similarity and the second similarity includes: Calculate the sum of the first similarity and the second similarity to obtain the sum of similarities; The ratio of the first similarity to the sum of the similarities is used as the first posterior probability of the known speech segment; The ratio of the second similarity to the sum of the similarities is used as the second posterior probability of the denoised speech segment.
8. The method according to claim 5, characterized in that, The step of calculating the similarity between the actual speech text and the first recognized text to obtain a first similarity, and calculating the similarity between the actual speech text and the second recognized text to obtain a second similarity, includes: Calculate the edit distance between the actual speech text and the first recognized text to obtain the first edit distance; calculate the edit distance between the actual speech text and the second recognized text to obtain the second edit distance. A first similarity between the actual speech text and the first recognized text is obtained based on the first edit distance, and a second similarity between the actual speech text and the second recognized text is obtained based on the second edit distance.
9. A speech recognition device, the device being used to implement the speech recognition method as described in any one of claims 1-8, characterized in that, The device includes a speech denoising module, a speech recognition effect prediction module, a speech selection module, and a speech recognition module; The speech denoising module is used to perform noise reduction processing on the acquired original speech data to obtain corresponding denoised speech data, and input the original speech data and the denoised speech data into the trained speech recognition effect prediction model. The speech recognition effect prediction module is used to predict the speech recognition effect based on the trained speech recognition effect prediction model according to the original speech data and the noise-reduced speech data, and obtain the target posterior probability. The speech selection module is used to determine the speech data to be recognized from the original speech data, the noise-reduced speech data and the fused speech data according to the target posterior probability, wherein the fused speech data is obtained by fusing the original speech data and the noise-reduced speech data using the target posterior probability; The speech recognition module is used to perform speech recognition on the speech data to be recognized, and use the obtained target recognition text as the speech recognition result corresponding to the original speech data.
10. A computer device comprising a memory, a processor, and computer-readable instructions stored in the memory and executable on the processor, characterized in that, When the processor executes the computer-readable instructions, it performs the steps of the speech recognition method as described in any one of claims 1-8.
11. A computer-readable storage medium storing computer-readable instructions, characterized in that, When the computer-readable instructions are executed by a processor, the processor performs the steps of the speech recognition method as described in any one of claims 1-8.
Citation Information
Patent Citations
Voice recognition method and device, computer readable storage medium and electronic equipment
CN113889091A