Speech recognition method, device and storage medium
By using a second speech recognition model with low model complexity and similar vocabulary for training, the recognition delay problem caused by adding model parameters or decoding delay time in the prior art is solved, and efficient speech recognition and accuracy are achieved.
Patent Information
- Application Number
- CN202211730468.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-30
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2042-12-30
AI Technical Summary
When existing speech recognition technology increases model parameters or decoding delay time to improve recognition accuracy, it will lead to a longer recognition delay, affecting the effectiveness of electronic devices.
A second speech recognition model is adopted, which has a lower model complexity than the first speech recognition model and is trained using a similar vocabulary list. The similar vocabulary list is a subset of the general vocabulary list, the amount of data is smaller than the general vocabulary list, and includes at least two vocabulary with similar relationships in the general vocabulary list.
The training efficiency and calculation efficiency of the speech recognition model are improved, the recognition accuracy of similar vocabulary is ensured, and the recognition delay is reduced, which solves the recognition delay problem caused by increasing model parameters or decoding delay time.
Smart Images

Figure CN116246612B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to a speech recognition method, device and storage medium, and belongs to the field of speech recognition technology. Background Art
[0002] With the development of artificial intelligence, many electronic devices have added voice control functions, such as smart bracelets, watches, electric fans, televisions, air conditioners, etc. These electronic devices usually use offline real-time voice recognition technology to recognize the voice data output by the user in real time and provide feedback interaction based on the recognition results. The accuracy of feedback and recognition delay will affect the use effect of electronic devices. Among them, the accuracy can be measured by the recognition rate and the word-crossing rate. The recognition rate refers to the ratio of the number of times that a word in the voice data can be correctly recognized as a word in the vocabulary to the number of times the word is actually spoken in the audio. The word-crossing rate refers to the ratio of the number of times the voice data is mistakenly recognized as other words in the vocabulary to the number of times the word is actually spoken in the audio. The vocabulary refers to the list of words that the electronic device can support for recognition. The recognition delay refers to the length of time from the speaker saying a word to the electronic device recognizing the word.
[0003] In order to improve the accuracy of electronic devices in recognizing voice data, it is generally achieved in the following two ways:
[0004] The first method is to adjust the network structure of the speech recognition model. Specifically, the larger the amount of data in the speech recognition model, the higher the accuracy of the output posterior result. Therefore, the accuracy of the speech recognition result can be improved by increasing the network parameters of the speech recognition model. In addition, since the speech recognition model takes into account past and future information when calculating the recognition result of each speech frame, the more information there is, the more accurate the output result will be in theory. Therefore, the accuracy of the speech recognition result can also be improved by increasing the amount of information used by the speech recognition model.
[0005] However, in the first method, if the model parameters are increased, the model needs to be retrained and then re-verified using the test set, which takes a long time. At the same time, the increase in model parameters or the amount of information used by the model will affect the model calculation speed, resulting in longer recognition delays and affecting the use of electronic equipment.
[0006] The second method is to increase the decoding delay time of the speech recognition model. Specifically, the result of the speech recognition model after decoding may not be correct. It is possible that some noise in the audio causes a mutation that affects the recognition result and outputs an erroneous result. Therefore, after the speech recognition model recognizes the result for the first time, it will wait for a period of time to confirm whether the result is output due to some special circumstances or a true result. Based on this, the accuracy of the speech recognition result can be improved by increasing the waiting time.
[0007] However, the second method requires extending the waiting time of the speech recognition model, which will also lead to longer recognition delays and affect the use of electronic devices. Summary of the invention
[0008] The present application provides a speech recognition method, device and storage medium. Since the model complexity of the second speech recognition model is lower than the model complexity of the first speech recognition model, and the second speech recognition model is trained using a similar word list; and the similar word list is a subset of the general word list, that is, the data volume of the similar word list is smaller than the data volume of the general word list, and the similar word list includes at least two words with similar relationships in the general word list, therefore, the training efficiency and computational efficiency of the second speech recognition model are both high, and at the same time, the accuracy of similar word recognition can be guaranteed, and the problem of large recognition delay caused by increasing the network parameters and / or input volume of the first recognition speech model can be solved; the recognition delay can be reduced while ensuring the accuracy of the recognition result. The present application provides the following technical solutions:
[0009] In a first aspect, a speech recognition method is provided, the method comprising:
[0010] Acquire speech data to be recognized, wherein the speech data to be recognized includes multiple speech frames;
[0011] Input the speech data to be recognized into a pre-trained first speech recognition model in a streaming manner to obtain a first vocabulary corresponding to the nth speech frame currently input into the first speech recognition model, where n is a positive integer;
[0012] determining whether the first word matches a word in a similar word list, the similar word list being a subset of the general word list, and the similar word list including at least two words in the general word list having a similar relationship;
[0013] In the case where the first vocabulary matches a vocabulary in a similar vocabulary list, the nth speech frame is input into a pre-trained second speech recognition model to obtain a second vocabulary corresponding to the nth speech frame; the model complexity of the second speech recognition model is lower than the model complexity of the first speech recognition model, and the second speech recognition model is trained using the similar vocabulary list;
[0014] In the case where the second vocabulary is different from the first vocabulary, inputting the speech frame after the n-th speech frame into the second speech recognition model to obtain a third vocabulary corresponding to the n-th speech frame;
[0015] In the case that there are a preset number of identical third words, the identical third words are output to obtain the word recognition result of the nth speech frame.
[0016] Optionally, the method further comprises:
[0017] In the case that the second vocabulary is the same as the first vocabulary, the second vocabulary is output to obtain the vocabulary recognition result of the nth speech frame.
[0018] Optionally, the method further comprises:
[0019] In the case that the first vocabulary does not match the vocabulary in the similar vocabulary list, the first vocabulary is output to obtain the vocabulary recognition result of the nth speech frame.
[0020] Optionally, the training process of the second speech recognition model includes:
[0021] For each word in the general vocabulary, determining whether the word has the similarity relationship with other words in the general vocabulary;
[0022] If the word has the similarity relationship with the other word, adding the word and the other word to the similar word list;
[0023] Inputting a pre-established first network model based on audio samples corresponding to the words in the similar vocabulary to obtain a model result;
[0024] Based on the difference between the model result and the words in the similar vocabulary, the network parameters of the first network model are updated to obtain the second speech recognition model.
[0025] Optionally, the determining whether the word has the similarity relationship with other words in the general vocabulary includes:
[0026] Determine whether there is a containment relationship between the word and the other words; if there is a containment relationship, determine that there is the similarity relationship;
[0027] and / or,
[0028] Determine whether the edit distance between the phoneme sequence of the vocabulary and the phoneme sequence of the other vocabulary is less than a distance threshold; and determine that the similarity relationship exists when the edit distance is less than the distance threshold.
[0029] Optionally, the first network model is a feedforward neural network.
[0030] Optionally, the training process of the first speech recognition model includes:
[0031] Inputting a pre-established second network model based on audio samples corresponding to the vocabulary in the general vocabulary to obtain a model result;
[0032] Based on the difference between the model result and the vocabulary in the general vocabulary, the network parameters of the second network model are updated to obtain the first speech recognition model.
[0033] Optionally, when there are a preset number of identical third words, outputting the same third word to obtain the word recognition result of the n-th speech frame includes:
[0034] In the case that there are two consecutive identical third words, the identical third words are output to obtain the word recognition result of the n-th speech frame.
[0035] In a second aspect, an electronic device is provided, the device comprising a processor and a memory; a program is stored in the memory, and the program is loaded and executed by the processor to implement the speech recognition method provided in the first aspect.
[0036] In a third aspect, a computer-readable storage medium is provided, wherein a program is stored in the storage medium, and when the program is executed by a processor, it is used to implement the speech recognition method provided in the first aspect.
[0037] The beneficial effects of the present application include at least: obtaining a first vocabulary corresponding to the nth frame of speech currently input into the first speech recognition model by streaming multiple frames of speech data to be recognized into a pre-trained first speech recognition model; determining whether the first vocabulary matches a vocabulary in a similar vocabulary list; in the case where the first vocabulary matches a vocabulary in a similar vocabulary list, inputting the nth frame of speech into a pre-trained second speech recognition model to obtain a second vocabulary corresponding to the nth frame of speech; in the case where the second vocabulary is different from the first vocabulary, inputting the speech frame after the nth frame of speech into the second speech recognition model to obtain a third vocabulary corresponding to the nth frame of speech; in the case where there are a preset number of identical third vocabulary words, outputting the same nth frame of speech. The vocabulary recognition result of the nth frame of speech frame is obtained by using three words; since the model complexity of the second speech recognition model is lower than the model complexity of the first speech recognition model, and the second speech recognition model is trained using a similar vocabulary list; and the similar vocabulary list is a subset of the general vocabulary list, that is, the data volume of the similar vocabulary list is smaller than the data volume of the general vocabulary list, and the similar vocabulary list includes at least two words with similar relationship in the general vocabulary list, therefore, the training efficiency and calculation efficiency of the second speech recognition model are both high, and at the same time, the accuracy of similar vocabulary recognition can be guaranteed, and the problem of large recognition delay caused by increasing the network parameters and / or input amount of the first recognition speech model can be solved; the recognition delay can be reduced under the premise of ensuring the accuracy of the recognition result.
[0038] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application and implement it according to the contents of the specification, the following is a detailed description of the preferred embodiments of the present application in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 is a flow chart of a speech recognition method provided by an embodiment of the present application;
[0040] Figure 2 is a flow chart of a model training method provided by an embodiment of the present application;
[0041] Figure 3 It is a block diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0042] The specific implementation methods of the present application are further described in detail below in conjunction with the accompanying drawings and examples. The following examples are used to illustrate the present application but are not intended to limit the scope of the present application.
[0043] Optionally, the present application uses the speech recognition method provided in each embodiment as an example for use in an electronic device, where the electronic device is a terminal or a server. The terminal may be a mobile phone, a computer, a tablet computer, a wearable device, or a household appliance, etc. This embodiment does not limit the type of electronic device.
[0044] Figure 1 : is a flow chart of a speech recognition method provided by an embodiment of the present application, the method comprising at least the following steps:
[0045] Step 101: Acquire speech data to be recognized, where the speech data to be recognized includes multiple speech frames.
[0046] In one example, after the electronic device collects audio data, it performs voice detection on the audio data; if the audio data is detected to be voice data, the voice data to be recognized is obtained, triggering the execution of step 102; if the audio data is detected not to be voice data, step 102 is not executed and the process ends.
[0047] Optionally, the electronic device collects audio data in real time; or, collects audio data when receiving a voice collection instruction. The voice collection instruction is generated when the electronic device receives a trigger operation acting on a voice collection control, or is sent by other devices. This embodiment does not limit the method for obtaining the voice collection instruction.
[0048] Step 102, streaming the speech data to be recognized into a pre-trained first speech recognition model to obtain a first vocabulary corresponding to the nth speech frame currently input into the first speech recognition model; n is a positive integer.
[0049] Optionally, for each speech frame, the first speech recognition model will calculate the recognition result corresponding to the speech frame in combination with information before the speech frame and / or information after the speech frame, and the recognition result includes the first vocabulary.
[0050] In one example, the recognition result of each speech frame includes at least one word ranked from high to low in confidence, wherein the first word is the word with the highest confidence. Accordingly, the information before the speech frame may be the recognition result of at least one speech frame before the speech frame, and the information after the speech frame may be the recognition result of at least one speech frame after the speech frame.
[0051] Optionally, the number of speech frames before or after the speech frame is considered to identify the delay setting, and the number of speech frames does not exceed a preset threshold, which may be 2 frames or 3 frames. This embodiment does not limit the value of the preset threshold.
[0052] Among them, the training process of the first speech recognition model includes: inputting a pre-established second network model based on audio samples corresponding to words in a general vocabulary to obtain a model result; updating the network parameters of the second network model based on the difference between the model result and the words in the general vocabulary to obtain the first speech recognition model.
[0053] Optionally, the second network model may be based on a bidirectional long short-term memory (LSTM) or a feedforward neural network (FNN). This embodiment does not limit the type of the second network model.
[0054] The general vocabulary is used to store the vocabulary supported by the first speech recognition model.
[0055] Step 103: determine whether the first word matches a word in a similar word list, where the similar word list is a subset of the general word list, and the similar word list includes at least two words in the general word list that have a similar relationship.
[0056] The electronic device compares the first word with each word in the similar word list. If there is a word identical to the first word in the similar word list, it determines whether the first word matches the words in the similar word list and executes step 104. If there is no word identical to the first word in the similar word list, it determines whether the first word does not match the words in the similar word list. If the first word does not match the words in the similar word list, the first word is output to obtain the word recognition result of the nth frame of speech frame.
[0057] Step 104, when the first vocabulary matches the vocabulary in the similar vocabulary list, the nth frame of speech frame is input into a pre-trained second speech recognition model to obtain a second vocabulary corresponding to the nth frame of speech frame; the model complexity of the second speech recognition model is lower than the model complexity of the first speech recognition model, and the second speech recognition model is trained using the similar vocabulary list.
[0058] For each speech frame, the second speech recognition model will at least calculate and obtain a recognition result corresponding to the speech frame in combination with information before the speech frame, and the recognition result includes a second vocabulary.
[0059] In one example, the recognition result of each speech frame includes at least one word ranked from high to low in confidence, wherein the second word is the word with the highest confidence. Accordingly, the information before the speech frame may be the recognition result of at least one speech frame before the speech frame.
[0060] Similarly, the number of speech frames before the speech frame is considered to identify the delay setting, and the number of speech frames does not exceed a preset threshold, which may be 2 frames or 3 frames. This embodiment does not limit the value of the preset threshold.
[0061] The training process of the second speech recognition model includes: for each word in the general vocabulary, determining whether there is a similarity relationship between the word and other words in the general vocabulary; if there is a similarity relationship between the word and other words, adding the word and other words to the similar vocabulary; inputting a pre-established first network model based on audio samples corresponding to the words in the similar vocabulary to obtain a model result; updating the network parameters of the first network model based on the difference between the model result and the words in the similar vocabulary to obtain a second speech recognition model.
[0062] In one example, determining whether a vocabulary has a similarity relationship with other vocabulary in a general vocabulary includes: determining whether a containment relationship exists between the vocabulary and other vocabulary; determining that a similarity relationship exists when a containment relationship exists; and / or determining whether an edit distance between a phoneme sequence of the vocabulary and phoneme sequences of other vocabulary is less than a distance threshold; determining that a similarity relationship exists when the edit distance is less than the distance threshold.
[0063] Optionally, the first network model is a feedforward neural network with a simple structure. Since the first network model has a simple structure and only requires a small amount of data (i.e., a similar word list) for training, the training efficiency and computational efficiency of the obtained second speech recognition model are both high.
[0064] In order to more clearly understand the training process of the second speech recognition model provided by the present application, the training process is described below with an example. Figure 2 , the training process includes at least the following steps:
[0065] Step 21, determining whether the words in the general vocabulary have a containment relationship, if so, inserting at least two words with the containment relationship into the similar vocabulary; if not, executing step 22;
[0066] Step 22, determining whether there are at least two words in the general vocabulary whose edit distance of the phoneme sequence is less than the distance threshold, if so, inserting the at least two words whose edit distance is less than the distance threshold into the similar vocabulary; if not, executing steps 23 and 24;
[0067] Step 23, using the general vocabulary to train the first speech recognition model, and the process ends;
[0068] Step 24, use the similar word list to train the second speech recognition model, and the process ends.
[0069] Step 105, when the second vocabulary is different from the first vocabulary, the speech frame after the n-th speech frame is input into the second speech recognition model to obtain a third vocabulary corresponding to the n-th speech frame.
[0070] Since the first word and the second word are different, it means that at least one of the results is wrong. The second speech recognition model has a higher recognition accuracy for words in the similar word list. Therefore, in this embodiment, by inputting the subsequent speech frames into the second speech recognition model to further recognize the nth frame of speech, the accuracy of recognizing the nth frame of speech can be improved. At this time, the waiting time for outputting the third word can output the result about 100ms earlier than the solution that does not use this real-time method.
[0071] Since the speech frame needs to use the recognition result of the previous speech frame when calculating the second vocabulary corresponding to the current frame, after the speech frame after the nth speech frame is input into the second speech recognition model, the second speech recognition model needs to use the recognition result of the nth speech frame when calculating the second vocabulary corresponding to the speech frame of the frame. At this time, the recognition result used this time is used as the third vocabulary.
[0072] Optionally, when the second vocabulary is the same as the first vocabulary, the second vocabulary is output to obtain the vocabulary recognition result of the nth speech frame. It has been verified that if the second vocabulary is directly outputted, the result can be outputted about 300ms earlier than if this solution is not used.
[0073] Step 106: When there are a preset number of identical third words, output the identical third words to obtain a word recognition result of the nth speech frame.
[0074] In one example, when there are two consecutive identical third words, the identical third words are output to obtain the word recognition result of the n-th speech frame.
[0075] In other embodiments, the third words may not be required to be consecutively the same, or the preset number may be greater. This embodiment does not limit the method for determining the word recognition result.
[0076] In summary, the speech recognition method provided in this embodiment obtains a first vocabulary corresponding to the nth frame of speech currently input into the first speech recognition model by streaming multiple frames of speech data to be recognized into a pre-trained first speech recognition model; determines whether the first vocabulary matches a vocabulary in a similar vocabulary list; if the first vocabulary matches a vocabulary in a similar vocabulary list, inputs the nth frame of speech into a pre-trained second speech recognition model to obtain a second vocabulary corresponding to the nth frame of speech; if the second vocabulary is different from the first vocabulary, inputs the speech frame after the nth frame of speech into the second speech recognition model to obtain a third vocabulary corresponding to the nth frame of speech; if there are a preset number of identical third vocabulary words, output the same The same third vocabulary is used to obtain the vocabulary recognition result of the nth frame of speech frame; since the model complexity of the second speech recognition model is lower than the model complexity of the first speech recognition model, and the second speech recognition model is trained using a similar vocabulary list; and the similar vocabulary list is a subset of the general vocabulary list, that is, the data volume of the similar vocabulary list is smaller than the data volume of the general vocabulary list, and the similar vocabulary list includes at least two words in the general vocabulary list with similar relationships, therefore, the training efficiency and calculation efficiency of the second speech recognition model are both high, and at the same time, the accuracy of similar vocabulary recognition can be guaranteed, and the problem of large recognition delay caused by increasing the network parameters and / or input amount of the first recognition speech model can be solved; the recognition delay can be reduced under the premise of ensuring the accuracy of the recognition result.
[0077] Figure 3 3 is a block diagram of an electronic device provided by an embodiment of the present application. The device at least includes a processor 301 and a memory 302.
[0078] The processor 301 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 301 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 301 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 301 may be integrated with a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 301 may also include an AI (Artificial Intelligence) processor, which is used to process computing operations related to machine learning.
[0079] The memory 302 may include one or more computer-readable storage media, which may be non-transitory. The memory 302 may also include a high-speed random access memory, and a non-volatile memory, such as one or more disk storage devices, flash memory storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 302 is used to store at least one instruction, which is used to be executed by the processor 301 to implement the speech recognition method provided in the method embodiment of the present application.
[0080] In some embodiments, the electronic device may further optionally include: a peripheral device interface and at least one peripheral device. The processor 301, the memory 302 and the peripheral device interface may be connected via a bus or a signal line. Each peripheral device may be connected to the peripheral device interface via a bus, a signal line or a circuit board. Schematically, the peripheral devices include but are not limited to: a radio frequency circuit, a touch display screen, an audio circuit, and a power supply.
[0081] Of course, the electronic device may also include fewer or more components, which is not limited in this embodiment.
[0082] Optionally, the present application also provides a computer-readable storage medium, in which a program is stored, and the program is loaded and executed by a processor to implement the speech recognition method of the above method embodiment.
[0083] Optionally, the present application also provides a computer product, which includes a computer-readable storage medium, in which a program is stored, and the program is loaded and executed by a processor to implement the speech recognition method of the above method embodiment.
[0084] The technical features of the above-described embodiments may be arbitrarily combined. To make the description concise, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0085] The above-mentioned embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the attached claims.
Claims
1. A speech recognition method, It is characterized in that The method comprises: Acquire speech data to be recognized, wherein the speech data to be recognized includes multiple speech frames; Input the speech data to be recognized into a pre-trained first speech recognition model in a streaming manner to obtain a first vocabulary corresponding to the nth speech frame currently input into the first speech recognition model, where n is a positive integer; determining whether the first word matches a word in a similar word list, the similar word list being a subset of a general word list, and the similar word list including at least two words in the general word list having a similar relationship; In the case where the first vocabulary matches a vocabulary in a similar vocabulary list, the nth speech frame is input into a pre-trained second speech recognition model to obtain a second vocabulary corresponding to the nth speech frame; the model complexity of the second speech recognition model is lower than the model complexity of the first speech recognition model, and the second speech recognition model is trained using the similar vocabulary list; In the case where the second vocabulary is different from the first vocabulary, inputting the speech frame after the n-th speech frame into the second speech recognition model to obtain a third vocabulary corresponding to the n-th speech frame; In the case that there are a preset number of identical third words, the identical third words are output to obtain the word recognition result of the nth speech frame.
2. The method according to claim 1, It is characterized in that The method further comprises: In the case that the second vocabulary is the same as the first vocabulary, the second vocabulary is output to obtain the vocabulary recognition result of the nth speech frame.
3. The method according to claim 1, It is characterized in that The method further comprises: In the case that the first vocabulary does not match the vocabulary in the similar vocabulary list, the first vocabulary is output to obtain the vocabulary recognition result of the nth speech frame.
4. The method according to claim 1, It is characterized in that The training process of the second speech recognition model includes: For each word in the general vocabulary, determining whether the word has the similarity relationship with other words in the general vocabulary; If the word has the similarity relationship with the other word, adding the word and the other word to the similar word list; Inputting a pre-established first network model based on audio samples corresponding to the words in the similar vocabulary to obtain a model result; Based on the difference between the model result and the words in the similar vocabulary, the network parameters of the first network model are updated to obtain the second speech recognition model.
5. The method according to claim 4, It is characterized in that The determining whether the word has the similarity relationship with other words in the general vocabulary comprises: Determine whether there is a containment relationship between the word and the other words; if there is a containment relationship, determine that there is the similarity relationship; and / or, Determine whether the edit distance between the phoneme sequence of the vocabulary and the phoneme sequence of the other vocabulary is less than a distance threshold; and determine that the similarity relationship exists when the edit distance is less than the distance threshold.
6. The method according to claim 4, It is characterized in that The first network model is a feedforward neural network.
7. The method according to claim 1, It is characterized in that The training process of the first speech recognition model includes: Inputting a pre-established second network model based on audio samples corresponding to the vocabulary in the general vocabulary to obtain a model result; Based on the difference between the model result and the vocabulary in the general vocabulary, the network parameters of the second network model are updated to obtain the first speech recognition model.
8. The method according to any one of claims 1 to 7, It is characterized in that The step of outputting the same third vocabulary to obtain the vocabulary recognition result of the n-th speech frame when there are a preset number of the same third vocabulary includes: In the case that there are two consecutive identical third words, the identical third words are output to obtain the word recognition result of the n-th speech frame.
9. An electronic device, It is characterized in that The device comprises a processor and a memory; a program is stored in the memory, and the program is loaded and executed by the processor to implement the speech recognition method according to any one of claims 1 to 8.
10. A computer-readable storage medium, It is characterized in that The storage medium stores a program, and when the program is executed by the processor, it is used to implement the speech recognition method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Voice recognition method and device and storage medium
CN110797026A
Voice recognition method and device, terminal and storage medium
CN111199730A