Speech recognition method and device, and computer-readable storage medium

Through the error correction processing of converting speech information into text in the voice interaction system, the semantic understanding error problem caused by speech information errors is solved, and higher speech recognition and semantic understanding accuracy is achieved.

CN114333795BActive Publication Date: 2025-05-13IFLYTEK CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111592910.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-23
Publication Date
2025-05-13
Estimated Expiration
2041-12-23

AI Technical Summary

Technical Problem

In the existing voice interaction technology, some key information in the voice information input by the user is incorrect, resulting in semantic comprehension errors and the user's expected response cannot be given, resulting in interaction failure.

Method used

By obtaining voice information and converting it into text, we judge whether the text semantics meet the preset standards. If it does not, the entity text sequence in the text will be corrected, replaced with a pronunciation coding sequence and added entity type tags. The error correction model is used to generate the error-corrected text, thereby improving the accuracy of speech recognition and semantic understanding.

Benefits of technology

Through error correction processing, the accuracy of speech recognition and semantic understanding is improved, so that the voice interaction system can understand user intentions more accurately and provide correct responses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114333795B_ABST
    Figure CN114333795B_ABST
Patent Text Reader

Abstract

The present application discloses a speech recognition method and device, and a computer-readable storage medium, and belongs to the field of speech interaction technology. The speech recognition method first obtains a first text according to speech information, and obtains a first semantic; wherein the first text includes a first entity text sequence, and the first semantic includes a first entity semantic corresponding to the first entity text sequence, and the first entity semantic has a corresponding entity type label; then judges whether the first semantic meets the preset standard; if so, the first semantic is used as the speech recognition result; otherwise, the first entity text sequence in the first text is replaced with the corresponding pronunciation coding sequence, and an entity type label is added to the pronunciation coding sequence to obtain an error correction text; a second entity text sequence is obtained according to the pronunciation coding sequence, and the entity type label is matched with the second entity text sequence to obtain a second text; and the speech recognition result is obtained using the second text. The present application improves the accuracy of speech recognition and semantic understanding.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of voice interaction technology, and in particular to a voice recognition method and device, and a computer-readable storage medium. Background Art

[0002] With the development of voice interaction technology, related voice interaction applications are becoming more and more widespread. In voice interaction, it is necessary to obtain semantic information based on the voice information input by the user, so as to respond to the user based on the semantic information. However, if some key information in the voice information input by the user is wrong, wrong semantic information will be obtained. After receiving the wrong semantic information, the backend cannot give the expected response to the user, resulting in interaction failure. Summary of the invention

[0003] The main technical problem solved by the present application is to provide a speech recognition method and device, and a computer-readable storage medium, which can improve the accuracy of speech recognition and semantic understanding.

[0004] In order to solve the above technical problems, a technical solution adopted by the present application is: to provide a speech recognition method, comprising:

[0005] Acquire a first text according to the voice information, and acquire a first semantic according to the first text; wherein the first text includes a first entity text sequence, the first semantic includes a first entity semantic corresponding to the first entity text sequence, and the first entity semantic has a corresponding entity type label;

[0006] Determining whether the first semantics meets a preset standard;

[0007] If yes, taking the first semantics as the speech recognition result;

[0008] Otherwise, replace the first entity text sequence in the first text with the corresponding pronunciation coding sequence, and add the entity type label to the pronunciation coding sequence to obtain a correction text; obtain a second entity text sequence based on the pronunciation coding sequence, and match the entity type label with the second entity text sequence to obtain a second text; and use the second text to obtain the speech recognition result.

[0009] The step of obtaining a second entity text sequence according to the pronunciation code sequence and matching the entity type label with the second entity text sequence to obtain a second text comprises:

[0010] Train the error correction model;

[0011] The error correction text is input into the error correction model, and the output of the error correction model is used as the second text.

[0012] The step of training the error correction model includes:

[0013] Providing a first training text, wherein the first training text includes a pronunciation coding sample sequence, and the pronunciation coding sample sequence has a text annotation sample sequence and a type annotation sample sequence matching the pronunciation coding sample sequence;

[0014] Inputting the first training text into the error correction model to obtain a first prediction result;

[0015] Based on the first training text and the first prediction result, the value of the parameter in the error correction model is adjusted so that the first prediction result is close to a first expected text, wherein the first expected text includes the text annotation sample sequence and the type annotation sample sequence.

[0016] Wherein, before the step of training the error correction model, the method further includes:

[0017] Train the pre-trained language model;

[0018] The parameters of the encoding layer in the error correction model are initialized using the trained parameters of the pre-trained language model.

[0019] The step of training the pre-trained language model includes:

[0020] Providing an initial text, wherein the initial text includes a plurality of characters and the initial text expresses correct semantics;

[0021] Obtaining a third probability that each character in the initial text is replaced by any other character in a preset set, and replacing at least one character in the initial text with another character in the preset set with the third probability to obtain a second training text;

[0022] Inputting the second training text into the pre-trained language model to obtain a second prediction result;

[0023] The value of the parameter in the pre-trained language model is adjusted based on the initial text and the second prediction result so that the second prediction result is close to the initial text.

[0024] The step of obtaining a third probability that each character in the initial text is replaced by any other character in a preset set includes:

[0025] For each character in the initial text, obtain the pronunciation similarity and meaning similarity between the current character and any other character in the preset set;

[0026] Obtaining a first sum value of all the pronunciation similarities associated with the current text and a second sum value of all the meaning similarities associated with the current text;

[0027] Obtaining a first ratio of a pronunciation similarity between the current character and another character to the first sum value, and a second ratio of a meaning similarity between the current character and another character to the second sum value;

[0028] Obtain a first product of the first ratio and a first probability p1, and a second product of the second ratio and a second probability p2, and use the sum of the first product and the second product as the third probability P ab ; wherein the sum of the first probability p1 and the second probability p2 is less than 1.

[0029] The first entity semantics belongs to a collection entity or a feature entity, the first entity text sequence includes a first collection entity text sequence and a first feature entity text sequence, the semantic understanding result of the first collection entity text sequence is the first entity semantics belonging to the collection entity, and the semantic understanding result of the first feature entity text sequence is the first entity semantics belonging to the feature entity; the step of using the second text to obtain the speech recognition result includes:

[0030] Acquire a second semantics according to the second text, and determine whether the difference between the second text and the first text is only related to the first characteristic entity text sequence;

[0031] If yes, taking the second semantics as the speech recognition result;

[0032] Otherwise, in response to the second semantics meeting the preset standard, the second semantics is used as the speech recognition result; in response to the second semantics not meeting the preset standard, the first semantics is used as the speech recognition result.

[0033] The first semantics also includes first intention semantics, and the first intention semantics is obtained by understanding the intention semantics of the first text. The step of obtaining the second semantics according to the second text includes:

[0034] Performing intent semantic understanding on the second text to obtain second intent semantics, and obtaining second entity semantics according to the second entity text sequence and the entity type label that match each other;

[0035] The second intent semantics and the second entity semantics are combined to obtain the second semantics.

[0036] The first semantics includes a combination of the first intention semantics and the first entity semantics, and the first entity semantics belongs to a set entity or a feature entity; and the step of determining whether the first semantics meets the preset standard includes:

[0037] Determining whether the combination included in the first semantics is in a preset reasonable combination list;

[0038] If not, determining that the first semantics does not meet the preset standard;

[0039] If so, further determine whether they are simultaneously satisfied that the first entity semantics all belong to the collection entity and that the first entity semantics are in the preset collection entity list; if they are simultaneously satisfied, determine that the first semantics meets the preset standard; if they are not simultaneously satisfied, determine that the first semantics does not meet the preset standard.

[0040] Before the step of obtaining the speech recognition result by using the second text, the method further includes:

[0041] Determining whether the characters and the sequence of characters in the second text match those in the first text;

[0042] If yes, the step of using the second text to obtain the speech recognition result is not performed, and the first semantics is used as the speech recognition result;

[0043] Otherwise, execute the step of obtaining a speech recognition result by using the second text.

[0044] In order to solve the above technical problems, another technical solution adopted by the present application is to provide a speech recognition device, comprising:

[0045] A first semantics acquisition module, used to acquire a first text according to voice information, and acquire a first semantics according to the first text; wherein the first text includes a first entity text sequence, the first semantics includes a first entity semantics corresponding to the first entity text sequence, and the first entity semantics has a corresponding entity type label;

[0046] A first judging module, used to judge whether the first semantics meets a preset standard;

[0047] A response module is used to use the first semantics as the speech recognition result when the first semantics meets the preset standard; and to replace the first entity text sequence in the first text with a corresponding pronunciation coding sequence when the first semantics does not meet the preset standard, and add the entity type label to the pronunciation coding sequence to obtain a correction text; obtain a second entity text sequence according to the pronunciation coding sequence, and match the entity type label with the second entity text sequence to obtain a second text; and obtain the speech recognition result using the second text.

[0048] To solve the above technical problems, another technical solution adopted in the present application is: to provide a speech recognition device, including a memory and a processor, wherein the memory stores program instructions, and the processor can execute the program instructions to implement the speech recognition method described in the above technical solution.

[0049] In order to solve the above technical problems, another technical solution adopted in the present application is: providing a computer-readable storage medium, on which program instructions are stored, and the program instructions can be executed by a processor to implement the speech recognition method described in the above technical solution.

[0050] The beneficial effects of the present application are as follows: the speech recognition method provided by the present application first obtains the first text according to the speech information, and obtains the first semantics according to the first text; wherein the first text includes the first entity text sequence, the first semantics includes the first entity semantics corresponding to the first entity text sequence, and the first entity semantics has the corresponding entity type label; then it is determined whether the first semantics meets the preset standard; if so, the first semantics is used as the speech recognition result; otherwise, the first entity text sequence in the first text is replaced with the corresponding pronunciation coding sequence, and the entity type label is added to the pronunciation coding sequence to obtain the error correction text; the second entity text sequence is obtained according to the pronunciation coding sequence, and the entity type label is matched with the second entity text sequence to obtain the second text; the speech recognition result is obtained using the second text. It can be seen that the present application performs error correction processing on the first text corresponding to the first semantics that does not meet the preset standard to obtain the second text, and focuses on the error correction of the first entity text sequence related to the semantics rather than the entire text sequence, and then obtains the speech recognition result according to the second text after error correction, so that the accuracy of the speech recognition result is higher. Moreover, when correcting the first text, not only the first entity text sequence is corrected, but also its entity type label is corrected, so that the predicted second entity text sequence can be constrained by the entity type label, further improving the accuracy of speech recognition and semantic understanding. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work. Among them:

[0052] Figure 1 This is a flow chart of an implementation method of the speech recognition method of the present application;

[0053] Figure 2 A flowchart of an implementation method for training an error correction model;

[0054] Figure 3 A flowchart of another implementation method for training an error correction model;

[0055] Figure 4 for Figure 3 A schematic flow chart of an implementation method of step S31;

[0056] Figure 5 for Figure 4 A schematic flow chart of an implementation method of step S42;

[0057] Figure 6 for Figure 1 A schematic diagram of a flow chart of an implementation method of step S14;

[0058] Figure 7 for Figure 6 A schematic flow chart of an implementation method of step S61;

[0059] Figure 8 for Figure 1 A schematic flow chart of another implementation of step S14;

[0060] Fig. 9 for Figure 1 A schematic flow chart of an implementation method of step S12;

[0061] Fig.10 This is a flow chart of another embodiment of the speech recognition method of the present application;

[0062] Fig.11 This is a structural diagram of an implementation method of a speech recognition device of the present application;

[0063] Fig.12 for Fig.11 A schematic diagram of the structure of an implementation scheme of a response module;

[0064] Fig.13 for Fig.11 A structural diagram of an implementation method of the first judgment module;

[0065] Fig.14 This is a structural diagram of an implementation method of a speech recognition device of the present application;

[0066] Fig.15 This is a schematic diagram of the structure of an embodiment of a computer-readable storage medium of the present application. DETAILED DESCRIPTION

[0067] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0068] See also Figure 1 , Figure 1 This is a flow chart of an implementation of a speech recognition method of the present application. The speech recognition method includes the following steps.

[0069] Step S11, obtaining a first text according to voice information, and obtaining a first semantic according to the first text; wherein the first text includes a first entity text sequence, the first semantic includes a first entity semantic corresponding to the first entity text sequence, and the first entity semantic has a corresponding entity type label.

[0070] The speech recognition method provided in this embodiment can be applied to a speech recognition device. After the device collects the speech information sent by the user, it first converts the speech information into text information, i.e., the first text, through a speech conversion model, and then converts the first text into semantic information, i.e., the first semantics, through a semantic understanding model.

[0071] Among them, the speech conversion model is a mature speech conversion technology in the prior art. The first text obtained by conversion includes a first entity text sequence and other text sequences. Other text sequences refer to meaningless text sequences such as spoken language. The first entity text sequence is obtained by removing these other text sequences, including key information for semantic understanding.

[0072] Among them, the semantic understanding model is a mature semantic understanding technology in the prior art. The first semantics obtained include first intention semantics and first entity semantics. The first intention semantics is obtained by understanding the intention semantics according to the context, and the first entity semantics is the semantic understanding result corresponding to the first entity text sequence, and the first entity semantics has a corresponding entity type label, which represents the entity type of the first entity text sequence.

[0073] For example, if a user gives a voice message and wants to express "play Andy Lau's Love Potion", the corresponding correct first semantics can be expressed as: intent = play, slot = song:Love Potion|artist:Andy Lau. Among them, "Andy Lau" and "Love Potion" are both first entity semantics, and their corresponding entity types are "artist" and "song" respectively.

[0074] Step S12, determining whether the first semantics meets a preset standard. If yes, executing the following step S13, otherwise, executing the following step S14.

[0075] The voice information given by the user may not be very accurate, or the expression may not be clear enough. If a response is made directly according to the first semantics, it may not be the response expected by the user. After obtaining the first semantics, this embodiment first determines whether the first semantics meets the preset standard, and then makes different responses according to different judgment results, thereby increasing the probability of giving the response expected by the user. The specific judgment method will be described below.

[0076] Step S13: taking the first semantics as a speech recognition result.

[0077] If the first semantics meets the preset standard, it means that the first semantics obtained in step S11 is likely to meet the user's expectations, and it is directly used as the speech recognition result and a response is made accordingly.

[0078] Step S14: replace the first entity text sequence in the first text with the corresponding pronunciation coding sequence, and add an entity type label to the pronunciation coding sequence to obtain a correction text; obtain a second entity text sequence based on the pronunciation coding sequence, and match the entity type label with the second entity text sequence to obtain a second text; and use the second text to obtain a speech recognition result.

[0079] If the first semantics does not meet the preset standard, it means that the first semantics obtained in step S11 is most likely not in line with the user's expectations, and it needs to be corrected before responding. This embodiment corrects errors on the basis of the first text to obtain a second text, and then further uses the second text to obtain a speech recognition result. Specifically, this embodiment replaces the first entity text sequence in the first text with a corresponding pronunciation coding sequence, and adds an entity type label to the pronunciation coding sequence to obtain a correction text; based on the correction text, a second entity text sequence is obtained according to the pronunciation coding sequence, and the entity type label is matched with the second entity text sequence to obtain a second text.

[0080] For example, the first text obtained according to the voice information given by the user is "Play Liu Dehua's Forget Love Water", in which the first entity text sequence includes "Liu Dehua" and "Forget Love Water", which are replaced with the corresponding pronunciation coding sequences "liudehua" and "wangqingshui", and the entity type tag "artist" is added to the pronunciation coding sequence "liudehua", and the entity type tag "song" is added to the pronunciation coding sequence "wangqingshui", and the error correction text "Play artist(liudehua)'s song(wangqingshui)" is obtained.

[0081] Then, the second entity text sequence is obtained according to the pronunciation coding sequence. The entity type label added according to the first semantics may also be wrong. For example, the first entity text sequence "Liu Dehua" corresponds to the wrong entity type label "song". This embodiment further matches the entity type label with the second entity text sequence, for example, the entity type label corresponding to the first entity text sequence "Liu Dehua" is corrected to "artist". Specifically, a trained neural network model can be used to obtain a second entity text sequence and a corresponding entity type label that are closer to the pronunciation coding sequence and closer to the user's intention, thereby improving the accuracy of the speech recognition results. That is, the second text is obtained by the following steps:

[0082] Step 1: Train the error correction model.

[0083] Step 2: input the error correction text into the error correction model, and use the output of the error correction model as the second text.

[0084] The error correction model includes but is not limited to the seq2seq model. After training, the error correction text can be corrected according to the above error correction process to obtain a second text. The specific training process will be described below.

[0085] This embodiment corrects the first text corresponding to the first semantics that does not meet the preset standard to obtain a second text, and focuses on correcting the first entity text sequence related to the semantics rather than the entire text sequence, and then obtains the speech recognition result based on the corrected second text, and the predicted generated second entity text sequence can be constrained by the entity type label, so that the accuracy of speech recognition and semantic understanding is higher.

[0086] In one embodiment, see Figure 2 , Figure 2 The flowchart of an implementation method for training an error correction model is shown below. The error correction model can be trained through the following steps.

[0087] Step S21, providing a first training text, wherein the first training text includes a pronunciation coding sample sequence, and the pronunciation coding sample sequence has a text annotation sample sequence and a type annotation sample sequence matching it.

[0088] After constructing an initial error correction model based on a neural network model such as a seq2seq model, a first training text is provided, wherein the first training text includes a pronunciation coding sample sequence, and the pronunciation coding sample sequence has a text annotation sample sequence and a type annotation sample sequence that match it. The error correction model mainly includes an encoding (Encoder) layer and a decoding (Decoder) layer, wherein the encoding layer is used to encode the input into an input semantic representation vector, and the decoding layer is used to output a corresponding output semantic representation vector according to the input semantic representation vector, thereby realizing prediction of another language sequence based on a language sequence. This embodiment trains the error correction model, and the prediction result expected to be obtained is to replace the pronunciation coding sample sequence with a text annotation sample sequence and a type annotation sample sequence that match it, thereby obtaining a second text after error correction.

[0089] Step S22: input the first training text into the error correction model to obtain a first prediction result.

[0090] After providing the first training text, it is input into the above error correction model to obtain a first prediction result, so as to facilitate the subsequent comparison of the difference between the first prediction result and the expected prediction result, thereby adjusting the model parameters.

[0091] Step S23, adjusting the values ​​of the parameters in the error correction model based on the first training text and the first prediction result, so that the first prediction result is close to the first expected text, and the first expected text includes a text annotation sample sequence and a type annotation sample sequence.

[0092] Further, the values ​​of the parameters in the error correction model are adjusted according to the first training text and the first prediction result, so that the first prediction result of the error correction model is close to the first expected text. It is understandable that the training process may undergo multiple parameter adjustments until it meets expectations.

[0093] This embodiment uses neural network technology to train an error correction model, which facilitates error correction of the first text. The result is more accurate and a more accurate second text can be obtained, thereby improving the accuracy of speech recognition and semantic understanding.

[0094] In one embodiment, see Figure 3 , Figure 3 This is a flowchart of another implementation method for training the error correction model, that is, the error correction model can also be trained through the following steps.

[0095] Step S31, training the pre-trained language model.

[0096] As mentioned above, the present application corrects the first entity text sequence representing key information in the first text, specifically by first replacing it with a pronunciation coding sequence, and then using an error correction model to correct the error, and the error correction process includes replacing text with similar pronunciation or similar meaning with a certain probability. This embodiment first trains the pre-trained language model so that the error correction model has a greater probability of replacing text with similar pronunciation or similar meaning, thereby improving the accuracy of error correction. The specific training process will be described below.

[0097] Step S32, using the parameters of the trained pre-trained language model to initialize the parameters of the encoding layer in the error correction model.

[0098] After the training of the pre-trained language model is completed, its parameters are further used to initialize the parameters of the encoding layer in the error correction model, so that the error correction model inherits the near-phonetic and near-synonymous word replacement function of the pre-trained language model.

[0099] Step S33, training the error correction model.

[0100] The error correction model is then trained. For details, please refer to the above steps S21-S23, which will not be repeated here.

[0101] This implementation uses a pre-trained language model that has completed training to initialize the encoding layer parameters of the error correction model, so that the error correction model has a stronger learning ability in replacing homophonic and synonymous characters, thereby obtaining a more accurate second text after completing the training of the error correction model.

[0102] In one embodiment, see Figure 4 , Figure 4 for Figure 3 The flowchart of step S31 in the embodiment is as follows: the pre-trained language model can be trained through the following steps.

[0103] Step S41, providing an initial text, wherein the initial text includes a plurality of characters and expresses correct semantics.

[0104] This implementation includes but is not limited to using the MLM (Mask Language Model) task to train the pre-trained language model. First, an initial text is obtained, which contains multiple characters and can express correct semantics.

[0105] Step S42, obtaining a third probability that each character in the initial text is replaced by any other character in the preset set, and replacing at least one character in the initial text with other characters in the preset set with the third probability to obtain a second training text.

[0106] The preset set is, for example, a set of commonly used characters, and the characters in the initial text that express the correct semantics are replaced with other characters, so that it has a high probability of expressing the wrong semantics, so as to train the pre-trained language model with it, and it is expected that the pre-trained language model predicts the initial text that expresses the correct semantics. The characters in the initial text are replaced with other characters corresponding to a greater third probability.

[0107] Please refer to Figure 5 , Figure 5 for Figure 4 The flowchart of step S42 of an implementation method can be used to obtain the third probability that each character in the initial text is replaced by any other character in the preset set through the following steps.

[0108] Step S51, for each character in the initial text, obtain the pronunciation similarity and meaning similarity between the current character and any other character in the preset set.

[0109] Assume that there are m characters in the preset set, and character a is used to represent the current character. aj and β aj They are respectively the pronunciation similarity and meaning similarity between character a and other characters in a preset set, where j is a positive integer from 2 to m.

[0110] Step S52, obtaining a first sum value of all pronunciation similarities associated with the current text and a second sum value of all meaning similarities associated with the current text.

[0111] That is, the first sum is The second sum is

[0112] Step S53, obtaining a first ratio of the pronunciation similarity between the current character and another character to the first sum value, and a second ratio of the meaning similarity between the current character and another character to the second sum value.

[0113] Use text b to represent another text, and the pronunciation similarity between text a and text b is α ab , the content similarity is β ab , then the first ratio is The second ratio is

[0114] Step S54, obtaining a first product of the first ratio and the first probability, and a second product of the second ratio and the second probability, and taking the sum of the first product and the second product as the third probability; wherein the sum of the first probability and the second probability is less than 1.

[0115] The first probability p1 is the probability of replacing the current text with a near-phonetic character, and the second probability p2 is the probability of replacing the current text with a synonym. The third probability P of text a being replaced with text b isab It can be expressed by the following formula (1):

[0116]

[0117] This embodiment defines the third probability so that the pre-trained language model can learn homophonic characters or synonymous characters with a higher probability, thereby improving the learning ability of the error correction model in this regard.

[0118] Step S43: input the second training text into the pre-trained language model to obtain a second prediction result.

[0119] This embodiment trains the pre-trained language model, and the expected prediction result is to replace the characters in the second training text with their homophones or synonyms with a higher probability, thereby improving the learning ability of the error correction model in this regard and obtaining a prediction result that expresses the correct semantics. After obtaining the second training text, it is input into the above-mentioned pre-trained language model to obtain a second prediction result. This facilitates the subsequent comparison of the difference between the second prediction result and the expected prediction result, thereby adjusting the model parameters.

[0120] Step S44: adjusting the values ​​of the parameters in the pre-trained language model based on the initial text and the second prediction result so that the second prediction result is close to the initial text.

[0121] Further, the values ​​of the parameters in the pre-trained language model are adjusted based on the initial text and the second prediction result, so that the second prediction result of the pre-trained language model is close to the initial text, that is, it is replaced with a near-pronunciation word or a near-synonym word with a higher probability to express the correct semantics. It is understandable that the training process may undergo multiple parameter adjustments until it meets expectations.

[0122] This embodiment uses neural network technology to train the pre-trained language model so that it can replace the characters in the text input into it with near-phonetic characters or synonyms to obtain prediction results that express correct semantics, thereby improving the learning ability of the error correction model in this regard and improving the accuracy of speech recognition and semantic understanding.

[0123] In one embodiment, see Figure 6 , Figure 6 for Figure 1 The flowchart of step S14 in an implementation manner is as follows: the speech recognition result can be obtained by using the second text.

[0124] Step S61, obtaining the second semantics according to the second text, and determining whether the difference between the second text and the first text is only related to the first characteristic entity text sequence. If yes, executing the following step S62, otherwise, executing the following step S63.

[0125] As mentioned above, the first text includes a first entity text sequence and other text sequences, and the first entity semantics included in the first semantics is a semantic understanding result corresponding to the first entity text sequence, and the first entity semantics has a corresponding entity type label, which characterizes the entity type of the first entity text sequence. The entity type includes a collection type and a feature type, that is, the first entity semantics belongs to a collection entity or a feature entity, and the first entity text sequence includes a first collection entity text sequence and a first feature entity text sequence. The semantic understanding result of the first collection entity text sequence is the first entity semantics belonging to the collection entity, and the semantic understanding result of the first feature entity text sequence is the first entity semantics belonging to the feature entity. For example, "city", "singer", etc. belong to collection entities, and "time", "place", etc. belong to feature entities.

[0126] The second text is obtained by correcting the first entity text sequence and the corresponding entity type label in the first text. After obtaining the second text, this embodiment uses a semantic understanding device to perform semantic understanding on the second text to obtain the second semantics. At the same time, it is determined whether the difference between the second text and the first text is only related to the first feature entity text sequence, that is, it is determined whether the correction process only corrects the first entity text sequence belonging to the feature class entity, so as to obtain different speech recognition results according to different judgment results.

[0127] Step S62: taking the second semantics as the speech recognition result.

[0128] If the difference between the second text and the first text is only related to the first characteristic entity text sequence, it means that only the first entity text sequence belonging to the characteristic entity has been corrected. It is highly likely that this correction is credible and not prone to errors. In this case, this implementation directly recognizes the second semantics and uses it as the speech recognition result.

[0129] Step S63, in response to the second semantics meeting the preset standard, taking the second semantics as the speech recognition result, and in response to the second semantics not meeting the preset standard, taking the first semantics as the speech recognition result.

[0130] If the difference between the second text and the first text is not only related to the first characteristic entity text sequence, it means that the first entity text sequence belonging to the collection entity and the error correction thereof are also performed. In this case, the present embodiment does not directly recognize the error correction, but further determines whether the second semantics obtained according to the second text after the error correction meets the above preset standard. If it meets the standard, the second semantics is considered to be credible and is used as the speech recognition result. If it does not meet the standard, the first semantics obtained according to the first text before the error correction is still used as the speech recognition result. The specific process of determining whether the second semantics meets the above preset standard will be described below.

[0131] This implementation scheme sets a method for approving the final speech recognition result according to the specific content of the error correction and the second semantics, thereby increasing the probability of obtaining more accurate speech recognition and semantic understanding.

[0132] In one embodiment, see Figure 7 , Figure 7 for Figure 6 The flowchart of step S61 in an implementation manner is as follows: the second semantics can be obtained according to the second text.

[0133] Step S71, performing intent semantic understanding on the second text to obtain second intent semantics, and obtaining second entity semantics based on the second entity text sequence and entity type label that match each other.

[0134] The second text is obtained by correcting the first entity text sequence and the corresponding entity type label in the first text. It is likely that it can reflect the accurate speech recognition result. After obtaining the second text, this embodiment first uses the semantic understanding device to understand the intent semantics to obtain the second intent semantics, which is equivalent to correcting the first intent semantics in the first semantics. After obtaining the corrected second text, the corrected second entity semantics can be obtained based on the second entity text sequence and entity type label that match each other, and can be directly combined with the second intent semantics to obtain the second semantics.

[0135] Step S72, combining the second intention semantics and the second entity semantics to obtain the second semantics.

[0136] Directly combining the second intention semantics with the second entity semantics to obtain the second semantics can reflect accurate speech recognition results with a greater probability.

[0137] This embodiment obtains the second intention semantics based on the corrected second text and corrects the first intention semantics. At the same time, the process of obtaining the second text is equivalent to correcting the first entity semantics. Therefore, the second semantics obtained by combining the second intention semantics and the second entity semantics improves the accuracy of speech recognition and semantic understanding.

[0138] In one embodiment, see Figure 8 , Figure 8 for Figure 1 FIG. 5 is a flow chart of another implementation of step S14 in FIG. 5 . This implementation includes the above steps S61-S63, and before step S61, further includes the following steps.

[0139] Step S81, determine whether the characters and character sequence in the second text match the first text. If yes, execute the following step S82. Otherwise, execute the following step S83.

[0140] After obtaining the corrected second text, this embodiment first determines whether the characters and character sequence in the second text match the first text, so as to predict whether the first semantics is accurate before further using the second text to obtain the second semantics.

[0141] Step S82: The step of obtaining a speech recognition result by using the second text is not performed, and the first semantics is used as the speech recognition result.

[0142] If the characters and character order in the second text match the first text, it means that the error correction process and the predicted results are consistent with those before the error correction, that is, the first text is more likely to be accurate. In this case, this implementation directly uses the first semantics as the speech recognition result to improve the efficiency of the speech recognition process.

[0143] Step S83, executing the step of obtaining the speech recognition result by using the second text, that is, executing the above steps S61-S63, and taking the first semantics or the second semantics as the speech recognition result.

[0144] If the characters and character order in the second text do not match the first text, it means that the error correction process has indeed corrected the first text, that is, the first text is most likely inaccurate. In this case, the present embodiment further utilizes the second text to obtain speech recognition results. For details, please refer to the above steps S61-S63, which will not be repeated here.

[0145] This implementation method predicts whether the first semantics is accurate based on the comparison between the second text and the first text, thereby executing different steps to improve the efficiency and accuracy of the speech recognition and semantic understanding process.

[0146] In one embodiment, see Fig. 9 , Fig. 9 for Figure 1 The flowchart of step S12 of an implementation method can be used to determine whether the first semantics meets the preset standard through the following steps.

[0147] Step S91, determine whether the combination included in the first semantics is in a preset reasonable combination list. If not, execute the following step S92. If yes, execute the following step S93.

[0148] As mentioned above, the first semantics includes a combination of the first intention semantics and the first entity semantics, and the first entity semantics belongs to a collection entity or a feature entity. Whether the present application corrects the first text depends on whether the first semantics meets the preset standard. Specifically, it is first determined whether the combination included in the first semantics is in a preset reasonable combination list. The reasonable combination list can be pre-set in the speech recognition device and includes various common reasonable combinations.

[0149] Step S92: determine that the first semantics does not meet the preset standard.

[0150] If the combination included in the first semantics is not in the above-mentioned reasonable combination list, it is directly determined that the first semantics does not meet the preset standard and needs to be corrected to obtain an accurate speech recognition result. The specific correction process can be referred to the above-mentioned implementation methods.

[0151] Step S93, further determine whether it is simultaneously satisfied that the first entity semantics all belong to a collection entity, and the first entity semantics are in a preset collection entity list. If it is simultaneously satisfied, execute the following step S94. If it is not simultaneously satisfied, execute the following step S95.

[0152] If the combination included in the first semantics is in the above-mentioned reasonable combination list, it is necessary to further determine whether two conditions are met at the same time: one is that the first entity semantics all belong to the collection entity, and the other is that the first entity semantics are in the preset collection entity list, so as to determine whether the first semantics needs to be corrected. The collection entity list can be pre-set in the speech recognition device, including common first entity semantics belonging to the collection entity.

[0153] Step S94: determine whether the first semantics meets a preset standard.

[0154] If the above two conditions are met at the same time, this embodiment determines that the first semantics meets the preset standard, and no error correction is required, and the first semantics is directly used as the speech recognition result.

[0155] Step S95: determining that the first semantics does not meet a preset standard.

[0156] If the above two conditions are not met at the same time, this embodiment determines that the first semantics does not meet the preset standard and needs to be corrected to obtain an accurate speech recognition result. The specific correction process can be found in the above embodiments.

[0157] This embodiment determines whether error correction is required by judging whether the first semantics meets the preset standard, thereby using the first semantics or the second semantics as the speech recognition result, which can improve the efficiency and accuracy of the speech recognition and semantic understanding process.

[0158] In some of the above implementations, it is also necessary to determine whether the second semantics meets the preset standard. Specifically, the same process as steps S91-S95 can be used for determination, which will not be repeated here.

[0159] The following is a specific application scenario to illustrate the process of the speech recognition method of this application. Fig.10 , Fig.10 This is a flow chart of another implementation of the speech recognition method of the present application. The speech recognition method includes the following steps.

[0160] Step S101, obtaining a first text according to voice information, and obtaining a first semantics according to the first text.

[0161] For details, please refer to the above step S11. For example, a user gives a voice message and wants to ask "What will the weather be like in Beijing next Wednesday?" The first text obtained is "What will the weather be like in the background on Wednesday in summer?" The corresponding first semantics understands "Wednesday in summer" as the first entity text sequence of the standard class, and adds the entity type tag of "daytime", and understands "background" as the first entity text sequence of the collection class, and adds the entity type tag of "city".

[0162] Step S102, determining whether the first semantics meets a preset standard. If yes, executing the following step S103, otherwise, executing the following step S104.

[0163] For details, please refer to the above steps S91-S95. If the above first semantics does not meet the preset standard, jump to step S104 to perform an error correction process.

[0164] Step S103: taking the first semantics as a speech recognition result.

[0165] Step S104: correct the first text to obtain a second text, and obtain a second semantics based on the second text.

[0166] Specifically, the error correction text "how is the weather at daytime (xiazhousan) city (beijing)" is first obtained based on the first text, and then the error correction model is used to correct the first entity text sequence and the corresponding entity type label in the first text to obtain the second text "how is the weather at datetime (next Wednesday) city (Beijing)", and then the second semantics is obtained based on it. For details, please refer to the description of the above-mentioned relevant implementation methods.

[0167] Step S105, determining whether the second semantics meets the preset standard. If yes, executing the following step S106, otherwise, executing the above step S103.

[0168] For details, please refer to the above steps S91-S95. If the above second semantics meets the preset standard, jump to step S106 to obtain the speech recognition result.

[0169] Step S106: taking the second semantics as the speech recognition result.

[0170] This implementation can improve the accuracy of speech recognition and semantic understanding. Based on the same inventive concept, this application also provides a speech recognition device, please refer to Fig.11 , Fig.11This is a structural diagram of an embodiment of a speech recognition device of the present application, and the speech recognition device includes a first semantic acquisition module 11, a first judgment module 12 and a response module 13. The first semantic acquisition module 11 is used to acquire a first text according to speech information, and acquire a first semantic according to the first text; wherein the first text includes a first entity text sequence, the first semantic includes a first entity semantic corresponding to the first entity text sequence, and the first entity semantic has a corresponding entity type label.

[0171] The first judgment module 12 is used to judge whether the first semantics meets the preset standard. The response module 13 is used to use the first semantics as the speech recognition result when the first semantics meets the preset standard; and to replace the first entity text sequence in the first text with the corresponding pronunciation code sequence when the first semantics does not meet the preset standard, and add an entity type label to the pronunciation code sequence to obtain an error correction text; obtain a second entity text sequence according to the pronunciation code sequence, and match the entity type label with the second entity text sequence to obtain a second text; and obtain the speech recognition result using the second text.

[0172] This embodiment corrects the first text corresponding to the first semantics that does not meet the preset standard to obtain a second text, and focuses on correcting the first entity text sequence related to the semantics rather than the entire text sequence, and then obtains a speech recognition result based on the corrected second text, and the predicted generated second entity text sequence can be constrained by the entity type label, and then further obtains the final speech recognition result based on whether the second semantics meets the preset standard, so that the accuracy of speech recognition and semantic understanding is higher.

[0173] In one embodiment, see Fig.12 , Fig.12 for Fig.11 The structural diagram of an implementation method of the response module 1 in the figure, the response module 13 includes an execution module 131, an error correction module 132 and a second semantic acquisition module 133. The execution module 131 is used to use the first semantics as the speech recognition result when the first semantics meets the preset standard. The error correction module 132 includes a replacement module 1321 and a first neural network module 1322. The replacement module 1321 is used to replace the first entity text sequence in the first text with the corresponding pronunciation coding sequence when the first semantics does not meet the preset standard, and add an entity type label to the pronunciation coding sequence to obtain an error correction text. The first neural network module 1322 is used to obtain a second entity text sequence according to the pronunciation coding sequence after the replacement module 1321 obtains the error correction text, and match the entity type label with the second entity text sequence to obtain a second text. The second semantic acquisition module 133 is used to obtain the speech recognition result using the second text.

[0174] Among them, the first neural network module 1322 includes a first training module 13221 and a first input-output module 13222. The first training module 13221 is used to train the error correction model, and the first input-output module 13222 is used to input the error correction text into the error correction model and use the output of the error correction model as the second text.

[0175] Specifically, the first training module 13221 is used to provide a first training text, which includes a pronunciation coding sample sequence, and the pronunciation coding sample sequence has a text annotation sample sequence and a type annotation sample sequence that match it; input the first training text into the error correction model to obtain a first prediction result; adjust the value of the parameter in the error correction model based on the first training text and the first prediction result so that the first prediction result is close to the first expected text, and the first expected text includes a text annotation sample sequence and a type annotation sample sequence.

[0176] This implementation can improve the accuracy of speech recognition and semantic understanding.

[0177] In one embodiment, please refer to Fig.12 The error correction module 132 also includes a second training module 1323 and an initialization module 1324. The second training module 1323 is used to train the pre-trained language model before the first training module 13221 trains the error correction model. The initialization module 1324 is used to initialize the parameters of the coding layer in the error correction model using the parameters of the trained pre-trained language model.

[0178] Specifically, the second training module 1323 is used to provide an initial text, which contains multiple characters and expresses correct semantics; obtain a third probability that each character in the initial text is replaced by any other character in a preset set, and replace at least one character in the initial text with other characters with the third probability to obtain a second training text; input the second training text into the pre-trained language model to obtain a second prediction result; adjust the value of the parameter in the pre-trained language model based on the initial text and the second prediction result so that the second prediction result is close to the initial text.

[0179] Specifically, the second training module 1323 is used to obtain, for each character in the initial text, the pronunciation similarity and meaning similarity between the current character and any other character in a preset set; obtain a first sum of all pronunciation similarities related to the current character, and a second sum of all meaning similarities related to the current character; obtain a first ratio of the pronunciation similarity between the current character and another character to the first sum, and a second ratio of the meaning similarity between the current character and another character to the second sum; obtain a first product of the first ratio and a first probability, and a second product of the second ratio and a second probability, and take the sum of the first product and the second product as the third probability; wherein the sum of the first probability and the second probability is less than 1.

[0180] This implementation can improve the accuracy of speech recognition and semantic understanding.

[0181] In one embodiment, the second semantic acquisition module 133 includes a first analysis module 1331 and a calling module 1332. The first entity semantics belongs to a collection entity or a feature entity, the first entity text sequence includes a first collection entity text sequence and a first feature entity text sequence, the semantic understanding result of the first collection entity text sequence is the first entity semantics belonging to the collection entity, and the semantic understanding result of the first feature entity text sequence is the first entity semantics belonging to the feature entity.

[0182] The first analysis module 1331 is specifically used to obtain the second semantics according to the second text, and to determine whether the difference between the second text and the first text is only related to the first characteristic entity text sequence. The calling module 1332 is used to call the execution module 131 to use the second semantics as the speech recognition result when the difference is only related to the first characteristic entity text sequence; and when the difference is not only related to the first characteristic entity text sequence, in response to the second semantics meeting the preset standard, call the execution module 131 to use the second semantics as the speech recognition result, and in response to the second semantics not meeting the preset standard, call the execution module 131 to use the first semantics as the speech recognition result.

[0183] Among them, the first semantics also includes the first intention semantics, which is obtained by performing intention semantic understanding on the first text. The first analysis module 1331 is specifically used to perform intention semantic understanding on the second text to obtain the second intention semantics, and to obtain the second entity semantics based on the matching second entity text sequence and entity type label; and to combine the second intention semantics and the second entity semantics to obtain the second semantics.

[0184] This implementation can improve the accuracy of speech recognition and semantic understanding.

[0185] In one embodiment, please refer to Fig.12The response module 13 also includes a second judgment module 134, which is used to judge whether the characters and character sequences in the second text match the first text before the second semantic acquisition module 133 uses the second text to obtain the speech recognition result; and when there is a match, directly use the first semantics as the speech recognition result, and notify the second semantic acquisition module 133 not to execute the step of using the second text to obtain the speech recognition result; when there is no match, notify the second semantic acquisition module 133 to execute the step of using the second text to obtain the speech recognition result.

[0186] This implementation can improve the accuracy of speech recognition and semantic understanding.

[0187] In one embodiment, see Fig.13 , Fig.13 for Fig.11 The structural diagram of an implementation method of the first judgment module 12, the first judgment module 12 includes a second analysis module 121 and a report module 122. Among them, the first semantics includes a combination of the first intention semantics and the first entity semantics, and the first entity semantics belongs to a collection entity or a feature entity.

[0188] Specifically, the second analysis module 121 is used to determine whether the combination included in the first semantics is in the preset reasonable combination list. The reporting module 122 is used to determine that the first semantics does not meet the preset standard when the combination is not in the reasonable combination list; to further determine whether it is simultaneously satisfied that the first entity semantics all belong to the collection entity, and the first entity semantics is in the preset collection entity list when the combination is in the reasonable combination list; and to determine that the first semantics meets the preset standard when both are satisfied, and to determine that the first semantics does not meet the preset standard when both are not satisfied.

[0189] This implementation can improve the accuracy of speech recognition and semantic understanding.

[0190] Based on the same inventive concept, the present application also provides a speech recognition device, see Fig.14 , Fig.14 This is a structural diagram of an embodiment of a speech recognition device of the present application, which includes a memory 141 and a processor 142, wherein the memory 141 stores program instructions, and the processor 142 can execute the program instructions to implement the speech recognition method described in any of the above embodiments. Please refer to the above embodiments for details, which will not be repeated here.

[0191] In addition, this application also provides a computer-readable storage medium, see Fig.15 , Fig.15This is a schematic diagram of the structure of an embodiment of a computer-readable storage medium of the present application. The storage medium 150 stores program instructions 151, which can be executed by a processor to implement the speech recognition method described in any of the above embodiments. Please refer to the above embodiments for details, which will not be repeated here.

[0192] The above description is only an implementation method of the present application, and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly used in other related technical fields, are also included in the patent protection scope of the present application.

Claims

1. A speech recognition method, characterized in that: include: Acquire a first text according to the voice information, and acquire a first semantic according to the first text; wherein the first text includes a first entity text sequence, the first semantic includes a first entity semantic corresponding to the first entity text sequence, and the first entity semantic has a corresponding entity type label; Determining whether the first semantics meets a preset standard; If yes, taking the first semantics as the speech recognition result; Otherwise, replace the first entity text sequence in the first text with the corresponding pronunciation coding sequence, and add the entity type label to the pronunciation coding sequence to obtain a correction text; obtain a second entity text sequence based on the pronunciation coding sequence, and match the entity type label with the second entity text sequence to obtain a second text; and use the second text to obtain the speech recognition result.

2. The speech recognition method according to claim 1, characterized in that: The step of obtaining a second entity text sequence according to the pronunciation coding sequence, and matching the entity type label with the second entity text sequence to obtain a second text comprises: Train the error correction model; The error correction text is input into the error correction model, and the output of the error correction model is used as the second text.

3. The speech recognition method according to claim 2, characterized in that: The step of training the error correction model includes: Providing a first training text, wherein the first training text includes a pronunciation coding sample sequence, and the pronunciation coding sample sequence has a text annotation sample sequence and a type annotation sample sequence matching the pronunciation coding sample sequence; Inputting the first training text into the error correction model to obtain a first prediction result; Based on the first training text and the first prediction result, the value of the parameter in the error correction model is adjusted so that the first prediction result is close to a first expected text, wherein the first expected text includes the text annotation sample sequence and the type annotation sample sequence.

4. The speech recognition method according to claim 2, characterized in that: Before the step of training the error correction model, the method further includes: Train the pre-trained language model; The parameters of the encoding layer in the error correction model are initialized using the trained parameters of the pre-trained language model.

5. The speech recognition method according to claim 4, characterized in that: The step of training the pre-trained language model includes: Providing an initial text, wherein the initial text includes a plurality of characters and the initial text expresses correct semantics; Obtaining a third probability that each character in the initial text is replaced by any other character in a preset set, and replacing at least one character in the initial text with another character in the preset set with the third probability to obtain a second training text; Inputting the second training text into the pre-trained language model to obtain a second prediction result; The value of the parameter in the pre-trained language model is adjusted based on the initial text and the second prediction result so that the second prediction result is close to the initial text.

6. The speech recognition method according to claim 5, characterized in that: The step of obtaining a third probability that each character in the initial text is replaced by any other character in a preset set includes: For each character in the initial text, obtain the pronunciation similarity and meaning similarity between the current character and any other character in the preset set; Obtaining a first sum value of all the pronunciation similarities associated with the current text and a second sum value of all the meaning similarities associated with the current text; Obtaining a first ratio of a pronunciation similarity between the current character and another character to the first sum value, and a second ratio of a meaning similarity between the current character and another character to the second sum value; Obtain a first product of the first ratio and a first probability p1, and a second product of the second ratio and a second probability p2, and use the sum of the first product and the second product as the third probability P ab ; wherein the sum of the first probability p1 and the second probability p2 is less than 1.

7. The speech recognition method according to claim 1, characterized in that: The first entity semantics belongs to a collection entity or a feature entity, the first entity text sequence includes a first collection entity text sequence and a first feature entity text sequence, the semantic understanding result of the first collection entity text sequence is the first entity semantics belonging to the collection entity, and the semantic understanding result of the first feature entity text sequence is the first entity semantics belonging to the feature entity; The step of obtaining the speech recognition result by using the second text comprises: Acquire a second semantics according to the second text, and determine whether the difference between the second text and the first text is only related to the first characteristic entity text sequence; If yes, taking the second semantics as the speech recognition result; Otherwise, in response to the second semantics meeting the preset standard, the second semantics is used as the speech recognition result; in response to the second semantics not meeting the preset standard, the first semantics is used as the speech recognition result.

8. The speech recognition method according to claim 7, characterized in that: The first semantics also includes first intention semantics, and the first intention semantics is obtained by understanding the intention semantics of the first text. The step of obtaining second semantics according to the second text includes: Performing intent semantic understanding on the second text to obtain second intent semantics, and obtaining second entity semantics according to the second entity text sequence and the entity type label that match each other; The second intent semantics and the second entity semantics are combined to obtain the second semantics.

9. The speech recognition method according to claim 1, characterized in that: The first semantics includes a combination of the first intention semantics and the first entity semantics, and the first entity semantics belongs to a set entity or a feature entity; the step of determining whether the first semantics meets the preset standard includes: Determining whether the combination included in the first semantics is in a preset reasonable combination list; If not, determining that the first semantics does not meet the preset standard; If so, further determine whether they are simultaneously satisfied that the first entity semantics all belong to the collection entity and that the first entity semantics are in the preset collection entity list; if they are simultaneously satisfied, determine that the first semantics meets the preset standard; if they are not simultaneously satisfied, determine that the first semantics does not meet the preset standard.

10. The speech recognition method according to claim 1, characterized in that: Before the step of using the second text to obtain the speech recognition result, the method further includes: Determining whether the characters and the sequence of characters in the second text match those in the first text; If yes, the step of using the second text to obtain the speech recognition result is not performed, and the first semantics is used as the speech recognition result; Otherwise, execute the step of obtaining a speech recognition result by using the second text.

11. A speech recognition device, characterized in that: include: A first semantics acquisition module, used to acquire a first text according to voice information, and acquire a first semantics according to the first text; wherein the first text includes a first entity text sequence, the first semantics includes a first entity semantics corresponding to the first entity text sequence, and the first entity semantics has a corresponding entity type label; A first judging module, used to judge whether the first semantics meets a preset standard; A response module is used to use the first semantics as a speech recognition result when the first semantics meets the preset standard; and to replace the first entity text sequence in the first text with a corresponding pronunciation coding sequence when the first semantics does not meet the preset standard, and add the entity type label to the pronunciation coding sequence to obtain a correction text; obtain a second entity text sequence according to the pronunciation coding sequence, and match the entity type label with the second entity text sequence to obtain a second text; and use the second text to obtain the speech recognition result.

12. A speech recognition device, characterized in that: It comprises a memory and a processor, wherein the memory stores program instructions, and the processor can execute the program instructions to implement the speech recognition method according to any one of claims 1 to 10.

13. A computer-readable storage medium, characterized in that: The storage medium stores program instructions, which can be executed by a processor to implement the speech recognition method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Text error correction method, related equipment and readable storage medium

    CN111554295A

  • Voice customer service text error correction method and device

    CN111985213A