Speech recognition methods, electronic devices and computer-readable storage media
By setting candidate recognition models for different domains and selecting the model that matches the current recognition domain for speech recognition, the problem of balancing computational speed and accuracy is solved, thus improving recognition efficiency.
Patent Information
- Application Number
- CN202211538100.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-01
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2042-12-01
AI Technical Summary
In existing technologies, it is difficult to balance computational speed and accuracy by using a fixed speech recognition model, which cannot meet the needs of different scenarios.
Candidate recognition models adapted to different fields are pre-set, and the model that matches the current recognition field is selected for recognition during the recognition process, reducing the amount of computation without reducing accuracy.
It achieves accurate recognition of speech information in different scenarios while reducing computational load and improving recognition efficiency.
Smart Images

Figure CN115881105B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of this disclosure relate to a speech recognition method, an electronic device, and a computer-readable storage medium. Background Technology
[0002] With the rapid development of science and technology, voice recognition is increasingly being applied in various scenarios, such as in-car navigation, smart homes, and daily office work, bringing great convenience to people's lives. Voice recognition takes speech as its research object, using speech signal processing and model recognition to enable machines to understand human language and convert it into digital signals that can be input into a computer. Summary of the Invention
[0003] At least one embodiment of this disclosure provides a speech recognition method, comprising: acquiring speech information to be recognized; determining at least one recognition domain for the speech information to be recognized; determining at least one recognition model that matches the at least one recognition domain from at least two candidate recognition models; and recognizing the speech information to be recognized using the at least one recognition model to obtain a speech recognition result.
[0004] For example, in a speech recognition method provided in one embodiment of this disclosure, determining at least one recognition domain for the speech information to be recognized includes: determining intent information corresponding to the speech information to be recognized based on the speech information to be recognized; and determining the at least one recognition domain based at least on the intent information.
[0005] For example, in a speech recognition method provided in one embodiment of this disclosure, determining the intent information corresponding to the speech information to be recognized based on the speech information to be recognized includes: performing feature extraction on the speech information to be recognized to obtain feature data; inputting the feature data into an acoustic model to obtain the output result of the acoustic model, wherein the output result of the acoustic model includes characters corresponding to the feature data; decoding K character sequences corresponding to the speech information to be recognized based on the output result of the acoustic model; and determining the intent information based on the K character sequences; wherein K is a positive integer.
[0006] For example, in a speech recognition method provided in one embodiment of this disclosure, determining at least one recognition domain for the speech information to be recognized includes: acquiring current status information of a speech receiving device that receives the speech information to be recognized and / or current status information of an associated device associated with the speech receiving device; determining current scene information based on the current status information of the speech receiving device and / or the current status information of the associated device; and determining the at least one recognition domain based at least on the current scene information.
[0007] For example, in a speech recognition method provided in one embodiment of this disclosure, determining at least one recognition domain for the speech information to be recognized includes: obtaining at least one historical recognition domain corresponding to at least one historical speech information received before receiving the speech information to be recognized; and determining the at least one recognition domain based at least on the at least one historical recognition domain.
[0008] For example, in a speech recognition method provided in one embodiment of this disclosure, determining at least one recognition domain for the speech information to be recognized includes: acquiring M reference information and determining M candidate domains based on the M reference information respectively; determining the at least one recognition domain based on the M candidate domains; wherein M is an integer greater than 1.
[0009] For example, in a speech recognition method provided in one embodiment of this disclosure, the M reference information includes at least one of the following: intent information corresponding to the speech information to be recognized; current scene information determined based on the current state information of the speech receiving device that receives the speech information to be recognized and / or the current state information of the associated device associated with the speech receiving device; and historical recognition domain corresponding to historical speech information received before receiving the speech information to be recognized.
[0010] For example, in a speech recognition method provided in an embodiment of this disclosure, determining at least one recognition domain includes: determining a recognition domain; and determining the at least one recognition domain based on the M candidate domains, including: if all M candidate domains are first domains, then the first domain is taken as the recognition domain; if the M candidate domains include N distinct candidate domains, then one candidate domain is determined from the N candidate domains as the recognition domain; wherein N is a positive integer greater than 1 and less than or equal to M.
[0011] For example, in a speech recognition method provided in one embodiment of this disclosure, determining a candidate domain from the N candidate domains as the recognition domain includes: based on the M candidate domains, counting the frequency of each of the N candidate domains; and selecting the candidate domain with the highest frequency among the N candidate domains as the recognition domain.
[0012] For example, in a speech recognition method provided in an embodiment of this disclosure, determining a candidate domain from the N candidate domains as the recognition domain includes: determining the reference information with the highest priority among the M reference information based on the priority ranking of the M reference information; and taking the candidate domain corresponding to the reference information with the highest priority among the N candidate domains as the recognition domain.
[0013] For example, in a speech recognition method provided in an embodiment of this disclosure, the M reference information includes P intent information and Q other information, and the M candidate domains include P first candidate domains corresponding to the P intent information and Q other candidate domains corresponding to the Q other information, where P is an integer greater than 1 and less than M, and Q is a positive integer less than M; determining a candidate domain from the N candidate domains as the recognition domain includes: if at least one candidate domain among the P first candidate domains and at least one candidate domain among the Q other candidate domains are both second domains, then the second domain is taken as the recognition domain.
[0014] For example, in a speech recognition method provided in one embodiment of this disclosure, determining at least one recognition domain includes: determining multiple recognition domains; and determining the at least one recognition domain based on the M candidate domains. This includes: if the M candidate domains include N distinct candidate domains, then the N candidate domains are used as the multiple recognition domains, where N is a positive integer greater than 1 and less than or equal to M; determining at least one recognition model matching the at least one recognition domain includes: determining multiple recognition models that respectively match the multiple recognition domains; and using the at least one recognition model to recognize the speech information to be recognized to obtain a speech recognition result. This includes: using the multiple recognition models respectively to recognize the speech information to be recognized to obtain multiple candidate recognition results; and determining the speech recognition result based on the multiple candidate recognition results.
[0015] For example, in a speech recognition method provided in one embodiment of this disclosure, the speech information to be recognized is recognized using the plurality of recognition models to obtain a plurality of candidate recognition results, including: recognizing the speech information to be recognized using the plurality of recognition models to obtain a plurality of candidate recognition results and a plurality of scores corresponding to the plurality of candidate recognition results; and determining the speech recognition result based on the plurality of candidate recognition results, including: selecting the candidate recognition result with the highest score from the plurality of candidate recognition results as the speech recognition result.
[0016] For example, one embodiment of the speech recognition method provided in this disclosure further includes: training at least two candidate recognition models for at least two preset recognition domains respectively; generating domain-model correspondence information based on the correspondence between the at least two preset recognition domains and the at least two candidate recognition models; wherein, determining at least one recognition model matching the at least one recognition domain includes: using the domain-model correspondence information to determine at least one recognition model matching the at least one recognition domain.
[0017] For example, in a speech recognition method provided in an embodiment of this disclosure, the recognition model includes a language model; using the at least one recognition model to recognize the speech information to be recognized to obtain a speech recognition result includes: inputting the output result of the acoustic model into the language model to obtain the output result of the language model, and using the output result of the language model as the speech recognition result.
[0018] For example, in a speech recognition method provided in one embodiment of this disclosure, the recognition model includes a language model; the speech information to be recognized is recognized using the at least one recognition model to obtain a speech recognition result. This includes: inputting the K character sequences into the language model to obtain the output result of the language model, and using the output result of the language model as the speech recognition result.
[0019] For example, in a speech recognition method provided in one embodiment of this disclosure, based on the output of the acoustic model, K character sequences corresponding to the speech information to be recognized are decoded, including: performing multiple decoding operations until a termination condition is triggered, to obtain multiple character sequences decoded by the last decoding operation and the sorting of the multiple character sequences, wherein the termination condition indicates the end of the currently decoded statement, and for each decoding operation after the first decoding operation, decoding continues based on the multiple incomplete character sequences obtained by the previous decoding operation; wherein, in the sorting, the multiple character sequences are arranged from high to low quality, and the K character sequences are the first K character sequences in the sorting.
[0020] At least one embodiment of this disclosure provides an electronic device, including a voice receiving device and a voice recognition device. The voice receiving device is configured to receive voice information to be recognized; the voice recognition device is configured to execute the voice recognition method provided in any embodiment of this disclosure.
[0021] At least one embodiment of this disclosure provides an electronic device, including a processor; a memory storing one or more computer program modules; wherein the one or more computer program modules are configured to be executed by the processor to implement the speech recognition method provided in any embodiment of this disclosure.
[0022] At least one embodiment of this disclosure provides a computer-readable storage medium storing non-transitory computer-readable instructions, which, when executed by a computer, can implement the speech recognition method provided in any embodiment of this disclosure. Attached Figure Description
[0023] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the accompanying drawings of the embodiments will be briefly described below. Obviously, the drawings described below only relate to some embodiments of this disclosure and are not intended to limit this disclosure.
[0024] Figure 1 A flowchart of a speech recognition method provided in at least one embodiment of this disclosure is shown;
[0025] Figure 2 A flowchart of a feature extraction operation provided in at least one embodiment of this disclosure is shown;
[0026] Figure 3 A flowchart of a speech recognition process provided by at least one embodiment of the present disclosure is shown;
[0027] Figure 4 A flowchart of another speech recognition process provided by at least one embodiment of the present disclosure is shown;
[0028] Figure 5 A flowchart of another speech recognition process provided by at least one embodiment of the present disclosure is shown;
[0029] Figure 6 A schematic block diagram of a speech recognition device provided in at least one embodiment of the present disclosure is shown;
[0030] Figure 7 A schematic block diagram of an electronic device provided for some embodiments of this disclosure;
[0031] Figure 8 A schematic block diagram of another electronic device provided in at least one embodiment of the present disclosure is shown;
[0032] Figure 9 A schematic block diagram of another electronic device provided in at least one embodiment of the present disclosure is shown; and
[0033] Figure 10 A schematic diagram of a computer-readable storage medium provided in at least one embodiment of the present disclosure is shown. Detailed Implementation
[0034] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this disclosure. All other embodiments obtained by those skilled in the art based on the described embodiments of this disclosure without creative effort are within the scope of protection of this disclosure.
[0035] Unless otherwise defined, the technical or scientific terms used in this disclosure shall have the ordinary meaning understood by one of ordinary skill in the art to which this disclosure pertains. The terms “first,” “second,” and similar terms used in this disclosure do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Similarly, the terms “an,” “a,” or “the,” and similar terms do not indicate a quantity limitation, but rather indicate the presence of at least one. The terms “including,” “comprising,” or “containing,” and similar terms mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects. The terms “connected,” “linked,” or similar terms are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. The terms “upper,” “lower,” “left,” and “right,” etc., are used only to indicate relative positional relationships, and these relative positional relationships may change accordingly when the absolute position of the described objects changes.
[0036] The inventors discovered that in related technologies, a single recognition model is used to perform speech recognition operations, and this fixed model is applied to various scenarios. Using a small-scale recognition model results in faster computation speed but lower accuracy. Using a large-scale recognition model, while achieving higher accuracy, involves greater computation and increases latency. Therefore, using a fixed recognition model makes it difficult to meet the needs of various scenarios and cannot balance computational speed and accuracy.
[0037] At least one embodiment of this disclosure provides a speech recognition method, a speech recognition device, an electronic device, and a computer-readable storage medium. The speech recognition method includes: acquiring speech information to be recognized; determining at least one recognition domain for the speech information to be recognized; determining at least one recognition model that matches the at least one recognition domain from at least two candidate recognition models; and recognizing the speech information to be recognized using the at least one recognition model to obtain a speech recognition result.
[0038] The speech recognition method of this embodiment pre-sets at least two candidate recognition models adapted to different fields. During the recognition process, the recognition model that matches the current recognition field is selected for recognition operation. It can be applied to a variety of scenarios and achieves the purpose of reducing the amount of computation without reducing accuracy.
[0039] Figure 1 A flowchart of a speech recognition method provided in at least one embodiment of the present disclosure is shown.
[0040] like Figure 1 As shown, the method may include steps S110 to S140.
[0041] Step S110: Obtain the speech information to be recognized.
[0042] Step S120: For the speech information to be recognized, determine at least one recognition area.
[0043] Step S130: From at least two candidate recognition models, determine at least one recognition model that matches the at least one recognition domain.
[0044] Step S140: Using the at least one recognition model, the speech information to be recognized is recognized to obtain a speech recognition result.
[0045] For example, in step S110, the voice information to be identified is audio data, which can be acquired through a voice receiving device such as a microphone, or by receiving audio data sent by other devices. The audio data can include audio in any file format and encoding, such as PCM (Pulse Code Modulation) audio data or WAV data (wave audio files). The audio sampling rate is, for example, 16000 Hz, with a bit width of 16 bits, and mono.
[0046] For example, in step S120, one or more related recognition domains are determined for the currently received voice information to be recognized. For example, in voice interaction mode, the recognition domain may include the topic type of the interaction, which can be used to indicate what type of topic is being interacted with. Topic types include, for example, weather, music, navigation, etc.
[0047] For example, before executing step S110, multiple candidate recognition models can be pre-acquired for use in the speech recognition process. For example, for at least two preset recognition domains, at least two corresponding candidate recognition models are trained respectively; domain-model correspondence information is generated based on the correspondence between the at least two preset recognition domains and the at least two candidate recognition models.
[0048] For example, as shown in Table 1, r recognition domains (r being an integer greater than 1) can be pre-defined as the aforementioned preset recognition domains. These r preset recognition domains may include, for example, navigation, weather query, music, train tickets, vehicle control, etc. For each preset recognition domain, a corresponding recognition model can be trained as a candidate recognition model. The candidate recognition model can be a neural network model, such as a convolutional neural network (CNN), a recurrent neural network (RNN), or an encoder-decoder. For each preset recognition domain, the corresponding candidate recognition model can be trained using sample data related to that domain and the corresponding label data. For example, for the navigation domain, recognition model D1 can be trained using navigation samples and labels; for the weather query domain, recognition model D2 can be trained using weather query samples and labels, and so on. Then, Table 1 can be generated based on the r preset recognition domains and the corresponding r candidate recognition models. Table 1 can serve as the aforementioned domain-model correspondence information, storing Table 1 and the model codes of the r candidate recognition models.
[0049] Table 1
[0050] Preset recognition area Candidate recognition model navigation Recognition Model D1 Weather forecast Recognition Model D2 music Recognition Model D3 Train tickets Recognition Model D4 … … Car control Recognition model Dr
[0051] For example, if at least one recognition domain obtained in step S120 is one of the domains included in Table 1, then in step S130, the domain-model correspondence information can be used to determine at least one recognition model that matches the at least one recognition domain. For example, for each recognition domain, a candidate recognition model corresponding to the recognition domain can be found in Table 1 as the recognition model that matches the recognition domain.
[0052] For example, in step S140, the model code of the at least one recognition model can be called to perform the recognition operation and obtain the speech recognition result.
[0053] According to at least one embodiment of the speech recognition method disclosed herein, at least two candidate recognition models adapted to different fields are pre-set. During the recognition process, the recognition model that matches the current recognition field is selected for recognition operation. This method can be applied to a variety of scenarios and achieves the purpose of reducing the amount of computation without reducing accuracy.
[0054] For example, in some examples, step S120 may include: determining intent information corresponding to the speech information to be recognized based on the speech information to be recognized; and determining at least one recognition domain based on the intent information.
[0055] For example, intent information can be the information mentioned in the voice message to be recognized, used to characterize the user's purpose in this voice interaction. For example, if the voice message to be recognized is "Please check the train tickets to City X for tomorrow," then the intent corresponding to this voice message is to check train tickets. Based on this intent, the corresponding recognition domain can be determined to be the train ticket domain.
[0056] For example, the intent information can be determined through the following steps: performing feature extraction on the speech information to be recognized to obtain feature data; inputting the feature data into an acoustic model to obtain the output of the acoustic model, wherein the output of the acoustic model includes characters corresponding to the feature data; decoding the K character sequences corresponding to the speech information to be recognized based on the output of the acoustic model; and determining the intent information based on the K character sequences, wherein K is a positive integer.
[0057] For example, in a feature extraction operation, the input data is audio data (i.e., the speech information to be recognized), and the output is feature data, such as FBANK (Filter-Bank) feature data. FBANK is obtained by summing the squares of the power spectra of the Mel filter and then taking the logarithm. The function of this feature extraction operation is to calculate feature data from the audio data, which is then used as input to the subsequent acoustic model.
[0058] The above example uses FBANK, but it can also be applied to other types of feature data, such as MFCC and PNCC.
[0059] Figure 2 A flowchart of a feature extraction operation provided in at least one embodiment of the present disclosure is shown.
[0060] like Figure 2 As shown, the audio file is first pre-emphasized, which enhances the high-frequency signal. Then, the pre-emphasized data is framed and windowed. During framing, the audio signal can be divided into frames, for example, 10ms. For instance, if the speech information to be recognized is "Check my train tickets to X city tomorrow," each syllable might last for hundreds of milliseconds. During framing, each syllable might be divided into a dozen or more frames. To prevent information loss between frames, a 25ms signal is used to calculate features each time, meaning a 10ms shift is made each time. The actual feature calculation uses a 25ms signal. Then, a Fourier transform is performed to obtain the frequency domain signal, i.e., the spectrum, from the time domain signal. After taking the power spectrum and squared amplitude, the frequency domain at a certain time is accumulated to obtain the speech spectrum. Then, the frequency is mapped to the Mel frequency scale using a Mel filter bank. Finally, the logarithm is taken to obtain the FBANK feature. Up to 80 Mel filter banks can be used, meaning each audio frame corresponds to 80 outputs.
[0061] For example, multiple frames can be used together as a recognition input in a recognition process, such as using 67 frames as a recognition input.
[0062] For example, the acoustic model operation is used to map feature data to characters. The input data for this step is the feature data obtained from the feature extraction operation mentioned above (e.g., FBANK feature data). The output is the character corresponding to each frame and the probability of the character, forming a temporal label matrix. For example, the feature extraction operation mentioned above obtained 160 frames of FBANK feature data. This feature data can be input into the acoustic model. The acoustic model outputs, for example, 5000 characters (including blanks, i.e., no characters) and the probability of each character for each frame, forming a temporal label matrix of [160, 5000]. For example, continuing with the above example, the speech information to be recognized is "Help me check the train tickets to X city tomorrow". For example, "help" is divided into frames 1 to 20. For each frame, 5000 results can be obtained. For a certain frame, the 5000 results obtained are, for example,: the probability of "help" is 0.8; the probability of "ba" is 0.5; the probability of "ban" is 0.3; ...; and the probability of blank is 0.001.
[0063] For example, the acoustic model can be a neural network model, such as a transformer-based network (a type of attention neural network). The acoustic model can be pre-trained using sample data and corresponding label data. In this embodiment, the acoustic model can be modeled based on FBANK feature data, meaning the sample data can be FBANK feature data. In other embodiments, the acoustic model can also be modeled based on other data, such as a phoneme-based acoustic model (e.g., the Kaldi model).
[0064] For example, the decoding operation is used to find the K best decoding paths (i.e., the K decoding paths with the highest probability) based on the temporal tag matrix obtained above. The input of this operation is the temporal tag matrix. For example, if there are 160 frames of FBANK feature data and the acoustic model outputs 5000 characters for each frame, then the input matrix of the decoding operation is [160, 5000], and the output of the decoding operation is the K best decoding paths (top K decoding paths).
[0065] For example, based on the output of the acoustic model, the K character sequences corresponding to the speech information to be recognized are decoded, including: performing multiple decoding operations until the termination condition is triggered, to obtain the multiple character sequences decoded by the last decoding operation and the sorting of the multiple character sequences, wherein the termination condition indicates the end of the currently decoded sentence, and for each decoding operation after the first decoding operation, decoding continues based on the multiple incomplete character sequences obtained by the previous decoding operation; wherein, in the sorting, the multiple character sequences are arranged from high to low quality, and the K character sequences are the first K character sequences in the sorting.
[0066] For example, CTC (Connectionist Temporal Classification) decoding can be used. CTC decoding is a type of streaming decoding that infers and decodes specific audio frames and provides real-time feedback of the decoding results to the user. Therefore, as audio input continues, multiple consecutive CTC decoding operations are required. For instance, if a user says, "Please check my train tickets to City X for tomorrow," the streaming recognition process might look like this:
[0067] First CTC decoding: Help me;
[0068] Second CTC decoding: Please check it for me;
[0069] Third CTC decoding: Please check for me tomorrow;
[0070] 4th CTC decoding: Please check for me tomorrow;
[0071] 5th CTC decoding: Please help me check if I'm going to City X tomorrow;
[0072] 6th CTC Decoding: Please help me check who's going to City X tomorrow;
[0073] 7th CTC Decoding: Help me check the train tickets to City X for tomorrow.
[0074] For example, multiple decoding paths can be obtained each time decoding is performed. The above 7 CTC decodings only show one of the decoding paths. For example, in the first CTC decoding, based on the previous several frames, the characters with relatively high probabilities are "帮 (bāng)", "把 (bǎ)", "班 (bān)", etc. Based on the following several frames, the characters with relatively high probabilities are "我 (wǒ)", "五 (wǔ)", "握 (wò)", etc. The probabilities of characters such as "我 (wǒ)", "五 (wǔ)", "握 (wò)" appearing after characters such as "帮 (bāng)", "把 (bǎ)", "班 (bān)" can be calculated, and several paths with relatively high probabilities are obtained. For example, 10 paths with the highest probabilities are obtained: "帮我 (bāng wǒ)", "把我 (bǎ wǒ)", "把握 (bǎ wò)", …, "班五 (bān wǔ)". Therefore, 10 decoding paths can be obtained in the first CTC decoding. In the second CTC decoding, as the user's voice continues to be input, further backward decoding can be performed based on the above 10 decoding paths. For example, 10 optimal decoding paths are also obtained. One of the decoding paths is, for example, "帮我查一下 (bāng wǒ chá yī xià)". And so on, until the end of the current sentence, multiple decoding paths (for example, top10 decoding paths) of the last CTC decoding are obtained, that is, multiple possible sentences are recognized. One of the decoding paths is, for example, the above "帮我查一下明天去X市的火车票 (bāng wǒ chá yī xià míng tiān qù X shì de huǒ chē piào)". The K best decoding paths obtained from the last CTC decoding can be used as the above K character sequences.
[0075] For example, the CTC decoding process can use the greedy algorithm (greedy search), the beam search algorithm (beam search), or the prefix beam search algorithm (prefix beam search).
[0076] For example, for the greedy algorithm (greedy search), the character with the highest probability is retained at each step. The algorithm is simple, but the accuracy is reduced. For example, Table 2 below shows the probabilities at 3 moments, that is, 3 frames T1, T2, and T3. The labels are, for example, 3: the blank label, the A label, and the B label. As shown in Table 2, the probability that the T1 frame is blank is 0.5, the probability that it is the A label is 0.2, and the probability that it is the B label is 0.3; the probability that the T2 frame is blank is 0.4, the probability that it is the A label is 0.3, and the probability that it is the B label is 0.3; the probability that the T3 frame is blank is 0.6, the probability that it is the A label is 0.3, and the probability that it is the B label is 0.1.
[0077] Table 2
[0078] T1 T2 T3 blank 0.5 0.4 0.6 A 0.2 0.3 0.3 B 0.3 0.3 0.1
[0079] For example, according to Table 2, calculate the probability of each tag. For the blank tag, we need to calculate the probability that all three frames are blank. The probability of the entire path (the product of all tag probabilities) is 0.5 * 0.4 * 0.6 = 0.12. For the A tag, there are three possible path combinations: "A--", "--A", and "-A-", where "-" represents the blank tag. The probability of the A tag is the sum of the probabilities of these three paths. The probability of "A--" is 0.2 * 0.4 * 0.6 = 0.048, the probability of "--A" is 0.5 * 0.4 * 0.3 = 0.06, and the probability of "-A-" is 0.5 * 0.3 * 0.6 = 0.09. Therefore, the sum of the probabilities of these three paths is 0.09 + 0.048 + 0.06 = 0.198, which is the probability of the A tag. Similarly, calculate the probability of the label being B, and then compare the probabilities of the blank label, the A label, and the B label. For example, if the probability of the label being A is the highest, then the decoding result of greedy search is A.
[0080] For example, greedy search selects the path with the highest probability at each time step, but CTC decoding aims to select the route with the highest probability throughout the entire process; these two approaches are sometimes inconsistent. Beam search can handle this inconsistency by maintaining K optimal routes at each step, rather than simply using the character with the highest probability at the current time step, where K is a hyperparameter. Beam search decoding is more computationally intensive, but yields more accurate results.
[0081] For example, in the prefix beam search algorithm, the K branches with the highest probability are retained at each step of the search. If the same line is found at a time node that has been processed before, they are merged. The prefix beam search has higher decoding accuracy.
[0082] For example, after decoding K character sequences, the intent information can be determined based on these K character sequences. For instance, the best character sequence (i.e., the top-1 decoding path) can be used to determine the intent information. The top-1 decoding path is then input into a pre-trained intent recognition model, which outputs the intent information. This intent recognition model is used in the speech recognition process, so no additional model is needed. For example, if the top-1 decoding path is "Check train tickets to X city tomorrow," the intent recognition model will output the intent information "Query train tickets," and the corresponding recognition domain is "train tickets."
[0083] For example, in some embodiments, step S120 may include: obtaining the current status information of the voice receiving device that receives the voice information to be identified and / or the current status information of the associated device associated with the voice receiving device; determining the current scene information based on the current status information of the voice receiving device and / or the current status information of the associated device; and determining the at least one recognition area based at least on the current scene information.
[0084] For example, a voice receiving device includes a microphone and other voice receiving components, such as a car's central control screen or a computer. Associated devices can be those controlled by the voice receiving device, allowing users to control them via voice commands. For instance, in a smart home system, a smart speaker can control lights, a robot vacuum cleaner, and a television; these devices can then be considered associated devices of the smart speaker. Similarly, in a car system, a user can control the air conditioning, seats, and other devices via the central control screen; these devices can then be considered associated devices of the central control screen.
[0085] For example, the current status information of a device may include the task currently being performed by the device, its power on / off status, etc. In some embodiments, the current scene information can be determined based on the current status information of the device received by voice. For example, if the app running on the central control screen is navigation, then the current scene information is "navigation," and the corresponding recognition domain is the navigation domain. As another example, if the interface displayed on the central control screen is a train ticket purchase interface, then the current scene information is "train ticket query," and the corresponding recognition domain is the train ticket domain. As yet another example, if the smart speaker is currently playing music, then the current scene information can be "music," and the corresponding recognition domain is the music domain. In other embodiments, the current scene information can be determined based on the current status information of associated devices. For example, if a smart home device associated with the smart speaker (e.g., a robot vacuum cleaner) is currently running, then the current scene information may include a "smart home control" scene, and the corresponding recognition domain is the smart home domain.
[0086] For example, in some embodiments, step S120 may include: obtaining at least one historical recognition domain corresponding to at least one historical voice message received before receiving the voice information to be recognized; and determining the at least one recognition domain based at least on the at least one historical recognition domain.
[0087] For example, the recognition domain can be determined based on context. Historical voice information can be the voice information from the previous round of interaction. The domain of the previous interaction can be determined based on historical voice information. For instance, if the user's previous voice message was about inquiring about train tickets, then the domain of the previous interaction would be "train tickets." Since some dialogues are continuous, such as when a user books train tickets and needs to go through multiple rounds of voice interaction to inquire about departure point, destination, departure time, train number, and type of train, multiple consecutive voice messages may correspond to the same domain. The current dialogue may be a continuation of the previous dialogue; therefore, the recognition domain of the previous interaction can also be considered as a possible recognition domain for the current interaction.
[0088] For example, the above describes three methods for determining the recognition domain. In some embodiments, one of these three methods can be selected to determine the recognition domain based on the actual situation or selection conditions. For instance, determining the recognition domain based on the intent information corresponding to the speech information to be recognized has high accuracy. Therefore, this method can be used to determine the recognition domain in scenarios where high accuracy is required. As another example, determining the recognition domain based on historical speech information is suitable for scenarios with continuous dialogue. Therefore, if the interval between the current speech information to be recognized and the previous historical speech information is less than a certain threshold, it is considered that there is a correlation between the current speech information and the previous speech information. Therefore, the method of determining the recognition domain based on historical speech information can be selected.
[0089] For example, in some embodiments, two or more methods can be combined to jointly determine the identification area. In some embodiments below, the above three methods are used as examples for illustration. However, this disclosure is not limited to these. In addition to the three methods described above, other methods can also be combined in practical applications.
[0090] For example, step S120 may include: acquiring M reference information and determining M candidate regions based on the M reference information; determining the at least one recognition region based on the M candidate regions; wherein M is an integer greater than 1. For example, the M reference information includes at least one of the following: intent information corresponding to the speech information to be recognized; current scene information determined based on the current state information of the speech receiving device that receives the speech information to be recognized and / or the current state information of the associated device associated with the speech receiving device; and historical recognition regions corresponding to historical speech information received before receiving the speech information to be recognized.
[0091] Figure 3 A flowchart of a speech recognition process provided by at least one embodiment of the present disclosure is shown.
[0092] like Figure 3As shown, for example, the M reference information includes three types: intent information, current scene information, and historical recognition domain. Intent information is obtained, for example, through feature extraction, acoustic model recognition, and CTC decoding of the audio. At least one candidate domain can be obtained based on each type of reference information; for example, a first candidate domain is obtained based on the intent information, a second candidate domain is obtained based on the current scene information, and a third candidate domain is obtained based on the historical recognition domain. The process of determining candidate domains based on these three types of reference information can be found in the relevant descriptions above and will not be repeated here. The number of each type of reference information can be one or more, and a corresponding candidate domain can be determined for each type of reference information. For example, for intent information, the top 3 decoding paths can be selected, three intent information pieces can be determined for each of these three paths, and three first candidate domains can be determined for each of these three intent information pieces.
[0093] For example, determining at least one recognition domain in step S120 includes: determining a recognition domain. That is, only one recognition domain is determined in step S120, a corresponding recognition model is determined based on the one recognition domain in step S130, and a recognition operation is performed using the one recognition model in step S140 to obtain a speech recognition result.
[0094] For example, when only one identification domain is determined, step S120 may include: if all M candidate domains are first domains, then the first domain is taken as the identification domain; if the M candidate domains include N different candidate domains, then one candidate domain is determined from the N candidate domains as the identification domain; where N is a positive integer greater than 1 and less than or equal to M.
[0095] For example, if the first, second, and third candidate domains are all of the same type, such as the navigation domain, then that domain can be determined as the final identification domain. If there are multiple different domains among the first, second, and third candidate domains, then one of those different domains can be selected as the final identification domain.
[0096] For example, determining a candidate domain from the N different candidate domains as the identification domain can include: based on the M candidate domains, counting the frequency of each of the N candidate domains; and taking the candidate domain with the highest frequency among the N candidate domains as the identification domain.
[0097] For example, if the number of candidate regions for the first, second, and third candidates is one, and if the first and second candidate regions are navigation, and the third candidate region is weather query, then the frequency of the navigation region is 2, and the frequency of the weather query region is 1. The navigation region has the highest frequency, therefore it can be used as the final identification region. As another example, if the number of candidate regions for the second and third candidates is one, and the number of candidate regions for the first is three, and if all three first candidate regions are navigation, the second candidate region is navigation, and the third candidate region is weather query, then the frequency of the navigation region is 4, and the frequency of the weather query region is 1. The navigation region has the highest frequency, therefore it can be used as the final identification region.
[0098] For example, in some embodiments, determining a candidate domain as the identification domain from N different candidate domains may include: determining the reference information with the highest priority among the M reference information based on the priority ranking of the M reference information; and taking the candidate domain corresponding to the reference information with the highest priority among the N candidate domains as the identification domain.
[0099] For example, when the first, second, and third candidate domains are all different, the recognition domain can be determined based on priority. The priority order of intent information, current scene information, and historical recognition domains can be predetermined, for example, the priority from high to low is: intent information, historical recognition domains, and current scene information.
[0100] For example, if the number of first, second, and third candidate domains is one, and the first candidate domain is the navigation domain, the second candidate domain is the music domain, and the third candidate domain is the weather query domain, then the navigation domain can be used as the final identification domain because the intent information has the highest priority.
[0101] For example, if multiple first candidate domains are obtained based on multiple intent information, and the intent information has the highest priority, then if all the multiple first candidate domains are the same domain, that domain can be used as the identification domain. If the multiple first candidate domains include two or more different domains, then one of them can be selected as the identification domain based on other conditions (such as frequency).
[0102] For example, the M reference information includes P intent information and Q other information, and the M candidate domains include P first candidate domains corresponding to the P intent information and Q other candidate domains corresponding to the Q other information, where P is an integer greater than 1 and less than M, and Q is a positive integer less than M. That is, there are multiple first candidate domains. In some embodiments, determining a candidate domain from N different candidate domains as the identification domain may include: if at least one candidate domain among the P first candidate domains and at least one candidate domain among the Q other candidate domains are both second domains, then the second domain is taken as the identification domain.
[0103] For example, based on the top 3 intent information, three first candidate domains are obtained, such as navigation, vehicle control, and music. A second candidate domain is obtained based on current scene information, such as navigation. A third candidate domain is obtained based on historical recognition information, such as weather query. Since one of the first and second candidate domains is navigation, the navigation domain can be used as the recognition domain.
[0104] For example, the above describes three selection methods for choosing one recognition domain from N different candidate domains. The first method is selection based on frequency, the second is selection based on priority, and the third is selection based on whether there is a domain that is the same as other candidate domains among multiple first candidate domains. In practical applications, one of these three selection methods can be chosen based on the actual situation or conditions. If one method cannot determine the final recognition domain, the other two selection methods can be combined.
[0105] For example, if the number of first, second, and third candidate domains is all one, the first selection method can be used first, that is, the candidate domain with the highest frequency is selected as the recognition domain. If the three candidate domains do not overlap, that is, the frequency is all 1, then the second selection method can be used, that is, the candidate domain with the highest priority (e.g., the first candidate domain) is selected as the recognition domain.
[0106] For example, in the case where the number of first candidate fields is multiple, for the above three selection methods, it can be first determined whether the multiple first candidate fields are all the same. If the multiple first candidate fields are all the same, the multiple first candidate fields are merged into one field, and then one field is selected from the merged field, the second candidate field, and the third candidate field as the recognition field. For example, the field with the highest frequency can be preferentially selected, or the field with the highest priority can be selected. If there are differences among the multiple first candidate fields, the third method can be preferentially used, that is, it is determined whether there is a field among the multiple first candidate fields that is the same as the second candidate field or the third candidate field. If there is a same field, that field is used as the recognition field. If there is no field among the multiple first candidate fields that is the same as the second candidate field or the third candidate field, the frequency information and the priority information can be combined to select one field as the recognition field. Based on this method, the accuracy of the recognition field can be made higher.
[0107] For example, as Figure 3 shown, after obtaining the recognition field, the recognition field can be used to select a recognition model, and then the recognition model is used for recognition operations to obtain a speech recognition result.
[0108] For example, the recognition model may include a language model.
[0109] For example, the language model may use a weighted finite state transducer WFST (Weighted Finite State Transducer). WFST takes into account the relationship between sequences, that is, the sequence transition weight. The Chinese characters before and after in a sentence are strongly correlated. For example, the two characters "你" and "好" may appear as "你好" and "好你" in a sentence, and the former is obviously more probable than the latter. WSFT is this decoding method, which adds the transition probability of the previous and subsequent characters to the probability of the whole sentence.
[0110] For example, in some embodiments, when using the recognition model for recognition operations, the output result of the above acoustic model can be input into the language model to obtain the output result of the language model, and the output result of the language model is used as the speech recognition result. For example, according to the acoustic model, the characters and their probabilities of each frame are obtained, forming a time series label matrix. The time series label matrix can be input into the language model, and the language model can decode to obtain multiple sentences with higher probabilities. One sentence with the highest probability (i.e., top1) can be used as the speech recognition result.
[0111] For example, in some embodiments, when performing a recognition operation using the recognition model, the K character sequences obtained from the above decoding operation can be input into the language model to obtain the output result of the language model, and the output result of the language model can be used as the speech recognition result. For example, if 10 optimal decoding paths are obtained based on the CTC decoding operation, these 10 optimal decoding paths can be input into the language model, and the language model can be used to re-score these 10 optimal decoding paths. The decoding path with the highest score determined by the language model can be used as the speech recognition result.
[0112] For example, determining at least one recognition domain in step S120 includes determining multiple recognition domains. That is, multiple recognition domains are determined in step S120. In this case, if the M candidate domains include N distinct candidate domains, then the N candidate domains can be used as the multiple recognition domains. In step S130, multiple recognition models that match the multiple recognition domains can be determined. In step S140, the speech information to be recognized is recognized using the multiple recognition models to obtain multiple candidate recognition results; based on the multiple candidate recognition results, the speech recognition result is determined.
[0113] Figure 4 A flowchart of another speech recognition process provided by at least one embodiment of the present disclosure is shown.
[0114] like Figure 4 As shown, a first recognition domain, a second recognition domain, and a third recognition domain are obtained based on intent information, current scene information, and historical recognition domains, respectively. When the number of each of these three recognition domains is one and they are all different, a first recognition model, a second recognition model, and a third recognition model can be matched based on these domains, respectively. Recognition operations are then performed using these three models to obtain a first candidate recognition result, a second candidate recognition result, and a third candidate recognition result, respectively. Finally, one of these three results is selected as the final speech recognition result.
[0115] For example, each recognition domain can have one or more. For instance, the number of first recognition domains can be three. If all three first recognition domains are different, then three corresponding recognition models can be obtained for the three first recognition domains. If two of the three first recognition domains are the same and these two recognition domains are different from the other recognition domain, that is, if there are two different recognition domains among the three first recognition domains, then two recognition models can be determined for these two different recognition domains.
[0116] Figure 5 A flowchart of another speech recognition process provided by at least one embodiment of the present disclosure is shown.
[0117] like Figure 5 As shown, with Figure 4 The difference is that if the first recognition domain and the second recognition domain are the same, then a recognition model can be determined based on the first recognition domain and the second recognition domain.
[0118] For example, for each recognition model, while obtaining candidate recognition results using the model, a score corresponding to each candidate recognition result can also be output. Based on the scores of multiple candidate recognition results, the candidate recognition result with the highest score can be selected as the final speech recognition result. This approach can make speech recognition results more accurate and shorten the recognition operation time.
[0119] According to the speech recognition method of this disclosure, a fixed recognition model is replaced with multiple selectable recognition models. One or more corresponding recognition models are selected for recognition operation based on conditions such as the current scene, intent information and user interaction context. This can reduce the amount of computation without reducing the recognition accuracy.
[0120] The speech recognition method according to the embodiments of this disclosure provides multiple reference conditions for determining the recognition domain and multiple selection methods for selecting the recognition domain from multiple candidate domains. It can be used flexibly according to actual conditions to meet the needs of various scenarios.
[0121] According to the speech recognition method of this disclosure, multiple recognition models can be selected based on multiple recognition domains, and the result with the highest score obtained from multiple recognition models can be determined as the final speech recognition result, which can make the speech recognition result more accurate.
[0122] Figure 6 A schematic block diagram of a speech recognition device 200 provided in at least one embodiment of the present disclosure is shown.
[0123] like Figure 6 As shown, the speech recognition device 200 includes a speech acquisition module 210, a domain determination module 220, a model determination module 230, and a model recognition module 240.
[0124] The voice acquisition module 210 is configured to acquire voice information to be recognized. For example, the voice acquisition module 210 can perform... Figure 1 Step S110 is described.
[0125] The domain determination module 220 is configured to determine at least one recognition domain for the speech information to be recognized. The domain determination module 220 may, for example, perform... Figure 1Step S120 is described.
[0126] The model determination module 230 is configured to determine, from at least two candidate recognition models, at least one recognition model that matches the at least one recognition domain. The model determination module 230 may, for example, perform... Figure 1 Step S130 is described.
[0127] The model recognition module 240 is configured to use the at least one recognition model to recognize the speech information to be recognized, so as to obtain a speech recognition result. For example, the model recognition module 240 may perform... Figure 1 Step S140 is described.
[0128] For example, the voice acquisition module 210, the domain determination module 220, the model determination module 230, and the model recognition module 240 can be hardware, software, firmware, or any feasible combination thereof. For example, the voice acquisition module 210, the domain determination module 220, the model determination module 230, and the model recognition module 240 can be dedicated or general-purpose circuits, chips, or devices, or they can be a combination of a processor and memory. The embodiments of this disclosure do not limit the specific implementation of the above-mentioned units.
[0129] It should be noted that in the embodiments of this disclosure, each unit of the speech recognition device 200 corresponds to each step of the aforementioned speech recognition method. For the specific functions of the speech recognition device 200, please refer to the relevant description of the speech recognition method, which will not be repeated here. Figure 6 The components and structure of the voice recognition device 200 shown are merely exemplary and not limiting. The voice recognition device 200 may also include other components and structures as needed.
[0130] Figure 7 This is a schematic block diagram of an electronic device provided for some embodiments of this disclosure. For example... Figure 7 As shown, the electronic device 300 includes a voice receiving device 310 and a voice recognition device 320. The voice receiving device 310 is configured to receive voice information to be recognized. The voice recognition device 320 is configured to perform the voice recognition method of any of the above embodiments.
[0131] The electronic device 300 may be, for example, a computer, tablet computer, mobile phone, television, or other such device, and the voice receiving device 310 may be, for example, a microphone or other device capable of receiving audio data. The voice receiving device 310 can transmit the received audio data to the voice recognition device 320, which may be, for example, a CPU or other processor, and can execute the voice recognition method of any of the above embodiments after receiving the audio data.
[0132] At least one embodiment of this disclosure also provides an electronic device including a processor and a memory, the memory storing one or more computer program modules. The one or more computer program modules are configured to be executed by the processor to implement the speech recognition method described above.
[0133] Figure 8 This is a schematic block diagram of another electronic device provided for some embodiments of this disclosure. For example... Figure 8 As shown, the electronic device 400 includes a processor 410 and a memory 420. The memory 420 stores non-transitory computer-readable instructions (e.g., one or more computer program modules). The processor 410 is used to execute the non-transitory computer-readable instructions, which, when executed by the processor 410, perform one or more steps in the speech recognition method described above. The memory 420 and the processor 410 can be interconnected via a bus system and / or other forms of connection mechanisms (not shown).
[0134] For example, processor 410 may be a central processing unit (CPU), a graphics processing unit (GPU), or other form of processing unit with data processing and / or program execution capabilities. For example, the central processing unit (CPU) may be an x86 or ARM architecture. Processor 410 may be a general-purpose processor or a special-purpose processor, capable of controlling other components in electronic device 400 to perform desired functions.
[0135] For example, memory 420 may include any combination of one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, erasable programmable read-only memory (EPROM), portable compact disc read-only memory (CD-ROM), USB memory, flash memory, etc. One or more computer program modules may be stored on the computer-readable storage medium, and processor 410 may run one or more computer program modules to implement various functions of electronic device 400. Various application programs and various data, as well as various data used and / or generated by the application programs, may also be stored in the computer-readable storage medium.
[0136] It should be noted that, in the embodiments of this disclosure, the specific functions and technical effects of the electronic device 400 can be referred to the description of the speech recognition method above, and will not be repeated here.
[0137] Figure 9This is a schematic block diagram of another electronic device provided in some embodiments of the present disclosure. The electronic device 500 is, for example, suitable for implementing the speech recognition method provided in the embodiments of the present disclosure. The electronic device 500 may be a terminal device, etc. It should be noted that... Figure 9 The illustrated electronic device 500 is merely an example and does not impose any limitation on the functionality and scope of use of the embodiments disclosed herein.
[0138] like Figure 9 As shown, electronic device 500 may include a processing device (e.g., a central processing unit, a graphics processor, etc.) 510, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 520 or a program loaded from storage device 580 into random access memory (RAM) 530. RAM 530 also stores various programs and data required for the operation of electronic device 500. The processing device 510, ROM 520, and RAM 530 are interconnected via bus 540. Input / output (I / O) interface 550 is also connected to bus 540.
[0139] Typically, the following devices can be connected to I / O interface 550: input devices 560 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 570 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 580 including, for example, magnetic tapes, hard disks, etc.; and communication devices 590. Communication device 590 allows electronic device 500 to communicate wirelessly or wiredly with other electronic devices to exchange data. Although Figure 9 An electronic device 500 with various devices is shown, but it should be understood that it is not required to implement or have all of the devices shown, and the electronic device 500 may alternatively implement or have more or fewer devices.
[0140] For example, according to embodiments of this disclosure, the above-described speech recognition method can be implemented as a computer software program. For instance, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program including program code for executing the above-described speech recognition method. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 590, or installed from a storage device 580, or installed from a ROM 520. When the computer program is executed by the processing device 510, the functions defined in the speech recognition method provided by embodiments of this disclosure can be implemented.
[0141] At least one embodiment of this disclosure also provides a computer-readable storage medium storing non-transitory computer-readable instructions that, when executed by a computer, can implement the above-described speech recognition method.
[0142] Figure 10 This is a schematic diagram of a storage medium provided for some embodiments of this disclosure. For example... Figure 10 As shown, storage medium 600 stores non-transitory computer-readable instructions 610. For example, when the non-transitory computer-readable instructions 610 are executed by a computer, one or more steps in the speech recognition method described above are performed.
[0143] For example, the storage medium 600 can be used in the aforementioned electronic device 400. For example, the storage medium 600 can be... Figure 8 The memory 420 in the illustrated electronic device 400. For example, a description of the storage medium 600 can be found here. Figure 8 The corresponding description of the memory 420 in the illustrated electronic device 400 will not be repeated here.
[0144] The following points need to be explained:
[0145] (1) The accompanying drawings of the embodiments of this disclosure only involve the structures involved in the embodiments of this disclosure. Other structures can be referred to the general design.
[0146] (2) Where there is no conflict, the embodiments of this disclosure and the features in the embodiments can be combined with each other to obtain new embodiments.
[0147] The above description is merely a specific embodiment of this disclosure, but the scope of protection of this disclosure is not limited thereto. The scope of protection of this disclosure should be determined by the scope of protection of the claims.
Claims
1. A speech recognition method, comprising: Obtain the speech information to be recognized; Based on the intent information, current scene information, and at least one piece of historical voice information received before receiving the voice information to be identified, multiple recognition domains corresponding to the intent information, the current scene information, and the historical voice information of the voice information to be identified are determined. In response to the fact that the identified multiple recognition domains are different, multiple recognition models that match the multiple recognition domains are determined from multiple candidate recognition models; The speech information to be identified is identified using the multiple recognition models respectively, so as to obtain multiple candidate recognition results and multiple scores corresponding to the multiple candidate recognition results; The candidate recognition result with the highest score is selected from the multiple candidate recognition results as the speech recognition result of the speech information to be recognized.
2. The speech recognition method according to claim 1 further includes: Feature extraction is performed on the speech information to be identified to obtain feature data; The feature data is input into the acoustic model to obtain the output of the acoustic model, wherein the output of the acoustic model includes characters corresponding to the feature data; Based on the output of the acoustic model, K character sequences corresponding to the speech information to be recognized are decoded. The intent information is determined based on the K character sequences; Where K is a positive integer.
3. The speech recognition method according to claim 1, wherein, The current scene information is determined based on the current status information of the voice receiving device that receives the voice information to be identified and / or the current status information of the associated devices associated with the voice receiving device.
4. The speech recognition method according to claim 1 further includes: Obtain at least one historical recognition domain corresponding to the at least one historical voice information; The recognition domain of the voice information to be recognized is determined based on at least one historical recognition domain.
5. The speech recognition method according to claim 1, further comprising: For each of the multiple preset recognition domains, a corresponding multiple candidate recognition models are trained. Based on the correspondence between the multiple preset recognition domains and the multiple candidate recognition models, domain-model correspondence information is generated; Among them, determining multiple recognition models that match the multiple recognition domains from multiple candidate recognition models includes: using the domain-model correspondence information to determine multiple recognition models that match the multiple recognition domains.
6. The speech recognition method according to claim 2, wherein, The recognition model includes a language model, and the method further includes: The output of the acoustic model is input into the language model to obtain the output of the language model, and the output of the language model is used as the first candidate recognition result among the multiple candidate recognition results.
7. The speech recognition method according to claim 2, wherein, The recognition model includes a language model, and the method further includes: The K character sequences are input into the language model to obtain the output of the language model, and the output of the language model is used as the first candidate recognition result among the multiple candidate recognition results.
8. The speech recognition method according to claim 7, wherein, Based on the output of the acoustic model, K character sequences corresponding to the speech information to be recognized are decoded, including: Multiple decoding operations are performed until the termination condition is triggered to obtain multiple character sequences decoded by the last decoding operation and the order of the multiple character sequences. The termination condition indicates the end of the currently decoded statement. For each decoding operation after the first decoding operation, the decoding continues based on the multiple incomplete character sequences obtained by the previous decoding operation. In the sorting process, the multiple character sequences are arranged from high to low quality, and the K character sequences are the first K character sequences in the sorting process.
9. An electronic device, comprising: A voice receiving device configured to receive voice information to be recognized; A speech recognition device configured to perform the speech recognition method as described in any one of claims 1-8.
10. An electronic device, comprising: processor; Memory, which stores one or more computer program modules; The one or more computer program modules are configured to be executed by the processor to implement the speech recognition method according to any one of claims 1-8.
11. A computer-readable storage medium storing non-transitory computer-readable instructions that, when executed by a computer, implement the speech recognition method according to any one of claims 1-8.
Citation Information
Patent Citations
Voice recognition method and system
CN102074231A
Domain recognition method and device for voice information, storage medium and electronic equipment
CN110705308A
Voice processing method and device
CN112259081A
Voice recognition method and device, electronic equipment, medium and program product
CN113051895A