Speech recognition error correction method, device, electronic device, and storage medium

By providing real-time error correction function in the speech recognition system, users can correct speech recognition errors through error correction instructions, solving the problem that speech recognition errors cannot be corrected in time in the prior art, and improving the accuracy of speech recognition and the accuracy of semantic understanding.

CN114360549BActive Publication Date: 2025-06-06HISENSE VISUAL TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111642506.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-29
Publication Date
2025-06-06
Estimated Expiration
2041-12-29

AI Technical Summary

Technical Problem

The prior art lacks effective real-time error correction methods in speech recognition, which leads to inability to correct speech recognition errors in time, affecting the accuracy of semantic understanding and the reach rate of media resources.

Method used

Provide a speech recognition error correction method, by obtaining voice information and converting it into text to be verified, receiving error correction instructions issued by users or devices, determining the text to be verified, and establishing an error correction mapping relationship based on the text information of the error text and the target text, and correcting the text to be verified in real time.

Benefits of technology

It realizes user active error correction in real time, improves the accuracy of speech recognition, and enhances the accuracy of semantic understanding and the reach of media.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114360549B_ABST
    Figure CN114360549B_ABST
Patent Text Reader

Abstract

The present application provides a speech recognition error correction method, device, electronic device, and storage medium. The method includes: obtaining speech information, converting the speech information into a text to be verified and displaying it; when receiving a voice instruction for instructing to correct the erroneous text to the target text, obtaining the text information of the erroneous text and the text information of the target text from the voice instruction, the text information at least including the phonetic code information; determining the text to be corrected in the text to be verified according to the text information of the erroneous text and the text information of the target text; obtaining the error correction mapping relationship between the text to be corrected and the target text according to the text information of the erroneous text and the text information of the target text; and correcting the text to be corrected to the target text based on the error correction mapping relationship. The method of the present application can correct errors that are prone to occur in speech recognition, so as to improve the accuracy of subsequent semantic error correction and the reach of media resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to speech recognition technology, and in particular to a speech recognition error correction method, device, electronic device, and storage medium. Background Art

[0002] With the development of artificial intelligence, many electronic devices (such as smart TVs) support voice search functions. For example, when the user says "search movie A", the electronic device will search for movie A.

[0003] However, the voice recognition part of the voice search function still has the problem of easy errors in voice recognition, such as easy errors in the recognition of homophones and when the user's pronunciation is not standard. For example, the movie the user wants to search for contains the word "Hao", but it is recognized as "Hao", so the wrong movie will be displayed.

[0004] In order to improve the accuracy of speech recognition, it is necessary to correct the errors that are prone to occur in speech recognition. However, the existing technology does not have a relatively complete method for real-time error correction of speech recognition. Therefore, how to correct the errors that are prone to occur in speech recognition to improve the accuracy of semantic understanding in speech recognition and the reach of media resources is still a problem worth studying. Summary of the invention

[0005] The present application provides a speech recognition error correction method, device, electronic device, and storage medium for correcting errors that are prone to occur in speech recognition to improve the accuracy of speech recognition.

[0006] On the one hand, the present application provides a speech recognition error correction method, comprising:

[0007] Acquire voice information, convert the voice information into text to be verified and display it;

[0008] When receiving a voice instruction for instructing to correct the erroneous text into a target text, acquiring text information of the erroneous text and text information of the target text from the voice instruction, the text information at least including phonetic code information;

[0009] Determining the text to be corrected in the text to be verified according to the text information of the erroneous text and the text information of the target text;

[0010] Acquire an error correction mapping relationship between the text to be corrected and the target text according to the text information of the erroneous text and the text information of the target text;

[0011] The text to be corrected is corrected into the target text based on the error correction mapping relationship.

[0012] Optionally, the acquiring the error correction mapping relationship between the text to be corrected and the target text according to the text information of the erroneous text and the text information of the target text includes:

[0013] According to the text information of the erroneous text and the text information of the target text, obtaining an error correction mapping relationship between the text to be corrected and the target text from a plurality of stored error correction mapping relationships;

[0014] When there is no error correction mapping relationship between the text to be corrected and the target text in the stored multiple error correction mapping relationships, an error correction mapping relationship between the text to be corrected and the target text is established according to the text information of the erroneous text and the text information of the target text.

[0015] Optionally, determining the text to be corrected in the text to be verified according to the text information of the erroneous text and the text information of the target text includes:

[0016] Generate phonetic and graphic code information for each word in the text to be verified;

[0017] Determine the similarity between the phonetic code information of each word in the text to be verified and the phonetic code information of the target text by using a phonetic code similarity algorithm;

[0018] Determine that the text composed of characters whose similarity reaches a preset similarity in the text to be verified is the text to be corrected;

[0019] The step of establishing an error correction mapping relationship between the text to be corrected and the target text according to the text information of the erroneous text and the text information of the target text includes:

[0020] Establishing a mapping relationship between the glyph information of the erroneous text and the glyph information of the target text;

[0021] According to the phonetic-graphic code information of the target text or the phonetic-graphic code information of the text to be corrected, and the mapping relationship between the glyph information of the erroneous text and the glyph information of the target text, a correction mapping relationship between the text to be corrected and the target text is established.

[0022] Optionally, also include:

[0023] Acquire text information of a plurality of corrected texts and text information of texts to be corrected corresponding to the plurality of corrected texts;

[0024] The multiple error correction mapping relationships are established and stored according to the text information of the multiple error-corrected texts and the text information of the texts to be corrected corresponding to the multiple error-corrected texts.

[0025] Optionally, the acquiring text information of the error text and text information of the target text from the voice instruction includes:

[0026] Converting the voice command into command text;

[0027] The text information of the error text and the text information of the target text indicated by the instruction text are obtained through a sequence labeling algorithm.

[0028] Optionally, also include:

[0029] When the received voice instruction is not for instructing to correct the erroneous text into the target text, the text to be verified is stored.

[0030] On the other hand, the present application provides a speech recognition error correction device, comprising:

[0031] A voice conversion module, used to obtain voice information, convert the voice information into text to be verified and display it;

[0032] An acquisition module, configured to, when receiving a voice instruction for instructing to correct an erroneous text into a target text, acquire text information of the erroneous text and text information of the target text from the voice instruction, wherein the text information at least includes phonetic code information;

[0033] A processing module, used for determining the text to be corrected in the text to be verified according to the text information of the error text and the text information of the target text;

[0034] The acquisition module is also used to acquire an error correction mapping relationship between the text to be corrected and the target text according to the text information of the error text and the text information of the target text;

[0035] An error correction module is used to correct the text to be corrected into the target text based on the error correction mapping relationship.

[0036] Optionally, the acquisition module is specifically used for:

[0037] According to the text information of the erroneous text and the text information of the target text, obtaining an error correction mapping relationship between the text to be corrected and the target text from a plurality of stored error correction mapping relationships;

[0038] When there is no error correction mapping relationship between the text to be corrected and the target text in the stored multiple error correction mapping relationships, an error correction mapping relationship between the text to be corrected and the target text is established according to the text information of the erroneous text and the text information of the target text.

[0039] On the other hand, the present application provides an electronic device, comprising: a processor, and a memory communicatively connected to the processor;

[0040] The memory stores computer-executable instructions;

[0041] The processor executes the computer-executable instructions stored in the memory to implement the speech recognition error correction method as described in the first aspect.

[0042] On the other hand, the present application provides a computer-readable storage medium, in which computer-executable instructions are stored. When the instructions are executed, the computer executes the speech recognition error correction method as described in the first aspect.

[0043] On the other hand, the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the speech recognition error correction method as described in the first aspect.

[0044] The speech recognition error correction method provided by the embodiment of the present application, after acquiring speech information and converting it into a text to be verified, if an error correction instruction initiated by a user or other device is received, the text information of the error text and the target text in the error correction instruction is obtained. The text to be corrected in the text to be verified is determined according to the text information of the error text and the text information of the target text, and the error correction mapping relationship between the text to be corrected and the target text is obtained according to the text information of the error text and the text information of the target text. When correcting the text to be corrected, the text to be corrected is corrected to the target text according to the error correction mapping relationship between the text to be corrected and the target text in the text to be verified. The error correction mapping relationship can be stored before or re-established. Thus, when the speech recognition is wrong, the user actively initiates a correction instruction, so that the electronic device corrects the text to be verified (above) displayed in the first round, so that the correct text (below) is displayed in the second round, and the user's real-time active error correction of the above is realized, so that the place where the speech recognition is wrong can be corrected, and after the correction, if the subsequent speech recognition is wrong, the accuracy of the semantic subsequent error correction will also be greatly improved.

[0045] In addition, the speech recognition error correction method provided in the embodiment of the present application will continuously update the error correction mapping relationship stored in the ES cache, that is, remember a lot of error correction mapping relationships and error correction results, and can directly display the correct text according to the error correction result or perform error correction according to the error correction mapping relationship during the next error correction. In this way, the accuracy of semantic subsequent error correction and the reach of media resources are greatly improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.

[0047] Figure 1 A schematic diagram of an application scenario of the speech recognition error correction method provided in this application.

[0048] Figure 2 A flowchart of a speech recognition error correction method provided for one embodiment of the present application.

[0049] Figure 3 A flowchart of a partial speech recognition error correction method provided for one embodiment of the present application.

[0050] Figure 4 Another schematic diagram of a speech recognition error correction method provided for an embodiment of the present application.

[0051] Figure 5 A flowchart of a partial speech recognition error correction method provided for one embodiment of the present application.

[0052] Figure 6 A schematic diagram of a speech recognition error correction device provided for one embodiment of the present application.

[0053] Figure 7 A schematic diagram of an electronic device provided for one embodiment of the present application.

[0054] The above drawings show clear embodiments of the present disclosure, which will be described in more detail below. These drawings and text descriptions are not intended to limit the scope of the present disclosure in any way, but to illustrate the concepts of the present disclosure to those skilled in the art by referring to specific embodiments. DETAILED DESCRIPTION

[0055] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0056] With the development of artificial intelligence, many electronic devices (such as smart TVs) support voice search functions. For example, when a user says "Search for movie A", the electronic device will search for movie A. However, in the voice recognition part of the voice search function, there is still the problem that voice recognition is prone to errors, such as the recognition of homophones is prone to errors, and the recognition is also prone to errors when the user's pronunciation is not standard. For example, the name of the movie that the user wants to search for contains the word "Hao", but it is recognized as "Hao", so the wrong movie will be displayed. In order to improve the accuracy of voice recognition, it is necessary to correct the errors that are prone to occur in voice recognition. However, in the prior art, after the voice recognition error, the user cannot actively correct the error in real time, and can only let the electronic device display the wrong recognition result. Therefore, the prior art does not have a relatively complete method for real-time error correction of voice recognition. How to correct the errors that are prone to occur in voice recognition to improve the accuracy of semantic understanding in voice recognition and the reach of media resources is still a problem worth studying.

[0057] Based on this, the present application provides a method, device, electronic device, and storage medium for speech recognition error correction. After acquiring speech information and converting it into a text to be verified, if a voice instruction actively initiated by a user or other device to indicate error correction is received, the text to be corrected in the text to be verified is corrected to the target text. Specifically, when correcting the text to be corrected, the text to be corrected is corrected to the target text according to the error correction mapping relationship between the text to be corrected and the target text. The error correction mapping relationship can be stored before or re-established. Thus, when speech recognition is wrong, the user actively initiates a voice instruction indicating error correction, so that the electronic device corrects the text to be verified displayed in the first round (i.e., above), so that the text displayed in the second round (i.e., below) meets the user's expectations, and the user's real-time active error correction is realized, so that the error in speech recognition can be corrected. After correction, if errors occur in subsequent speech recognition, the accuracy of semantic subsequent error correction will also be greatly improved.

[0058] The speech recognition error correction method provided in the present application is applied to electronic devices, such as smart TVs, tablet computers, mobile phones, etc. Figure 1 This is a schematic diagram of the application of the speech recognition error correction method provided by the present application. In the figure, the electronic device converts the acquired speech information into text to be verified (for example, Figure 1(as shown, convert the voice information into "I want to watch xx Hao xx" and display it). The user finds that "Hao" is incorrect, and the correct one should be "郝". At this time, the user can send a voice command for indicating correction to the electronic device (for example, "It's not this Hao, but the one with a '赤' on the left and an '阝' on the right"). When the electronic device receives the voice command for indicating correction, it obtains the text information of the incorrect text ("Hao") and the text information of the target text ("郝") from the voice command, and then determines the text to be corrected in the text to be verified according to the text information of the incorrect text and the text information of the target text (for example Figure 1 the text to be corrected shown is "Hao"). According to the correction mapping relationship between the text to be corrected and the target text, correct the text to be corrected in the text to be verified to the target text (correct "Hao" to "郝").

[0059] Please refer to Figure 2 , one embodiment of the present application provides a voice recognition correction method, including:

[0060] S210, obtain voice information, convert the voice information into a text to be verified and display it.

[0061] The voice information is voice information with a search nature input from the outside (such as the user). For example, the voice information is "I want to watch xx Hao xx", and "xx Hao xx" is the name of a movie. When the electronic device converts the voice information into a text to be verified, if there is an identification error, the displayed text to be verified may be "I want to watch xx Hao xx".

[0062] S220, when receiving a voice command for indicating correcting an incorrect text to a target text, obtain the text information of the incorrect text and the text information of the target text from the voice command, and the text information includes at least phonetic and shape code information.

[0063] When the text to be verified displayed by the electronic device is incorrect, the user can send a voice command again to correct the text to be verified. The voice command is, for example, "It's not this Hao, but the one with a '赤' on the left and an '阝' on the right". After receiving the voice command for indicating correcting the incorrect text ("Hao") to the target text ("郝"), obtain the phonetic and shape code information P11 = {['h"ao']} of the incorrect text from the voice command, and obtain the shape information S11 = {['女”子']} of the incorrect text, the phonetic and shape code information P12 = {['h"ao']} of the target text, and the shape information S12 = {['赤”阝']} of the target text.

[0064] Optionally, when obtaining the text information of the error text and the text information of the target text, first convert the voice command into command text, and then use the sequence labeling algorithm to obtain the text information of the error text and the text information of the target text indicated by the command text. First, use the sequence labeling algorithm to label the error text and the target text in the command text. When determining the text information of the target text, it is also necessary to use the sequence labeling algorithm to label the text related to the target text. After being labeled by the sequence labeling algorithm, the command text is "not this one above [good], but [on the (left) side, a (naked) 'chi', on the (right) side, a 'radical of 'ear'] hao". When determining the text information of the target text, from "[on the (left) side, a (naked) 'chi', on the (right) side, a 'radical of 'ear']", determine the phonetic shape code P21L = {['ch”i']['l”u”o']['l”u”o']}, the glyph information S21L = {['赤']['礻”果']['礻”果']}, and also determine the radical information C23R = {['radical of 'ear']}. Determine that the glyph information of the target text is S12 = {['赤”阝']} according to the phonetic shape code, glyph information, and radical information determined from "[on the (left) side, a (naked) 'chi', on the (right) side, a 'radical of 'ear']".

[0065] Optionally, if the received voice command is not used to indicate correcting the error text in the text to be verified to the target text, store the text to be verified. For example, when the user finds that the text to be verified is correct or does not want to correct the error, and issues other voice commands (such as "I want to listen to music"), then execute this other voice command and directly store the text to be verified.

[0066] Optionally, as Figure 3 shown, when identifying whether the voice command is used to indicate error correction, the voice command can first be identified as command text, and then the command text can be analyzed according to the Abnf semantic rule analysis method. If the command text is used to indicate error correction, then use the sequence labeling algorithm to obtain the text information of the error text and the text information of the target text indicated by the command text, and then obtain the error correction mapping relationship between the text to be corrected and the target text described in step S240.

[0067] S230. Determine the text to be corrected in the text to be verified according to the text information of the error text and the text information of the target text.

[0068] For example, if the text information includes phonetic shape code information, and the phonetic shape code information includes pinyin information and glyph information, then the text to be corrected "good" can be determined from the text to be verified according to the pinyin information P11 = P12 = {['h”ao']} of the error text and the target text, and the glyph information S11 = {['女”子']} of the error text.

[0069] Optionally, the phonetic and graphic code information (including pinyin information and graphic information) of each character in the text to be verified can be generated. For example, if the text to be verified is "I want to see xx good xx", the pinyin information of each character in the text to be verified is PP = {['w”o']['x”i”ang']['k”an']…['h”ao']…}. Then, the similarity between the pinyin information of each character in the text to be verified and the pinyin information P11 = P12 = {['h”ao']} of the target text (or the error text) is determined through the phonetic and graphic code similarity algorithm, and the text composed of the characters in the text to be verified whose similarity reaches the preset similarity is determined as the text to be corrected ("good"). The preset similarity should be as high as possible so that the text to be corrected can be accurately found.

[0070] S240. Obtain the error correction mapping relationship between the text to be corrected and the target text according to the text information of the error text and the text information of the target text.

[0071] The error correction mapping relationship between the text to be corrected and the target text is, for example, {"pronunciation": "[‘h”ao’]", "ErrorCharacters": "[‘女”子’]", "CorrectCharacters": "[‘赤”阝’]"}.

[0072] Optionally, multiple error correction mapping relationships have been stored in the recognition text cache (Elasticsearch, abbreviated as ES) of the electronic device. These multiple error correction mapping relationships are obtained based on many previously corrected texts. That is, obtain the text information of multiple corrected texts and the text information of the corresponding texts to be corrected, and then establish and store these multiple error correction mapping relationships according to the text information of the multiple corrected texts and the text information of the corresponding texts to be corrected. For example, when the electronic device leaves the factory, it may have cached some error correction mapping relationships established for the voice information that is often misrecognized in daily life. The electronic device will also store the error correction mapping relationships established at that time when continuously correcting texts later.

[0073] Therefore, when obtaining the error correction mapping relationship between the text to be corrected and the target text according to the text information of the error text and the text information of the target text, it is possible to first search in the ES cache according to the text information of the error text and the text information of the target text, that is, as Figure 4As shown in the figure, first determine whether there is a matching result in the ES cache. That is, according to the text information of the error text and the text information of the target text, match the error correction mapping relationship from the stored multiple error correction mapping relationships. If a match can be found, the result stored in the ES cache can be obtained, that is, the error correction mapping relationship between the text to be corrected in the text to be verified and the target text can be obtained.

[0074] If the error correction mapping relationship between the text to be corrected in the text to be verified and the target text cannot be obtained from the ES cache, the error correction mapping relationship between the text to be corrected in the text to be verified and the target text needs to be established. As Figure 4 shown in the figure, the text information of the error text and the text information of the target text can be confirmed first through the sequence labeling algorithm described in step S230, and the text to be corrected can be confirmed through the phonetic and shape code similarity algorithm. Then, according to the pinyin information of the target text or the pinyin information of the text to be corrected, and the mapping relationship between the glyph information of the error text and the glyph information of the target text, the error correction mapping relationship between the text to be corrected and the target text is established.

[0075] S250, correct the text to be corrected in the text to be verified to the target text based on the error correction mapping relationship between the text to be corrected and the target text.

[0076] The error correction mapping relationship between the text to be corrected and the target text is, for example, {"pronunciation":"[‘h”ao’]","ErrorCharacters":"[‘女”子’]","CorrectCharacters":"[‘赤”阝’]"}. After correcting the text to be corrected in the text to be verified to the target text based on the error correction mapping relationship between the text to be corrected and the target text, the obtained error correction result is CorrectResult = {"pronunciation":"[‘h”ao’]","ErrorCharacters":"[‘女”子’]","CorrectCharacters":"[‘赤”阝’]","CorrectRetext":"[“我想看xx郝xx”]".}

[0077] As Figure 4 shown in the figure, after correcting the text to be verified, execute the normal semantic understanding logic. At this time, the film and television displayed by the electronic device is "xx郝xx".

[0078] Optionally, as Figure 4As shown, it is possible to re-check whether there is an error correction mapping relationship between the text to be corrected and the target text in the ES cache. If not, the established error correction mapping relationship between the text to be corrected and the target text is stored in the ES cache. Figure 4 As shown, after the error correction mapping relationship between the text to be corrected and the target text is established, it is also possible to detect whether the error correction mapping relationship related to the target text stored in the ES cache is consistent with the newly established error correction mapping relationship. If it is inconsistent, the error correction mapping relationship related to the target text established in the previous round stored in the ES cache is deleted, and then the established error correction mapping relationship between the text to be corrected and the target text is stored in the ES cache.

[0079] Optional, such as Figure 5 As shown, when the error correction result of the text to be verified is generated and the error correction mapping relationship between the newly created text to be corrected and the target text is stored, the newly created error correction mapping relationship can be stored in the form of a phonetic code generation table (or also includes a radical mapping table), and then a mapping code is created for the newly created error correction mapping relationship, and the mapping code can be linked to the newly created error correction mapping relationship. In this way, when the text to be corrected needs to be corrected to the target text later, the corresponding error correction result can be determined according to the mapping code, and the error correction text can be corrected directly according to the unique error correction result.

[0080] After creating the mapping code, the phonetic code information in the newly created error correction mapping relationship and the corresponding target text can be segmented to generate a marked entity based on an open source phonetic code generation tool, and the error correction result is updated after recording the entity phonetic code and the target text, and the updated error correction result is stored in the ES cache. If the phonetic code information mapping needs to be updated in the later stage, the word segmentation interface can be called to update. After creating the phonetic code information mapping (or also including radical mapping) between the text to be corrected and the target text, the phonetic code information of the target text only corresponds to the phonetic code information of the text to be corrected. In this way, the newly created phonetic code information mapping (or also including radical mapping) makes the text to be corrected and the target text correspond one to one. After obtaining the phonetic code information of the target text, the text to be corrected can be directly matched, and the text to be corrected in the text to be verified is directly corrected to the target text.

[0081] In summary, after acquiring voice information and converting it into a text to be verified for display, the speech recognition error correction method provided in the present embodiment, if an error correction instruction actively initiated by a user or other device is received, obtains the text information of the error text and the target text in the error correction instruction. The text to be corrected in the text to be verified is determined based on the text information of the error text and the text information of the target text, and the error correction mapping relationship between the text to be corrected and the target text is obtained based on the text information of the error text and the text information of the target text. When correcting the text to be corrected, the text to be corrected is corrected to the target text based on the error correction mapping relationship between the text to be corrected in the text to be verified and the target text. The error correction mapping relationship may have been stored before or may be newly established. Therefore, when there is an error in voice recognition, the user actively initiates a voice command to indicate correction, so that the electronic device corrects the text to be verified (the above text) displayed in the first round, and then displays the correct text (the following text) in the second round. This enables the user to actively correct the above text in real time, so that the errors in voice recognition can be corrected. After the correction, if errors occur in subsequent voice recognition, the accuracy of subsequent semantic error correction will also be greatly improved.

[0082] In addition, the speech recognition error correction method provided in this embodiment will continuously update and store the error correction mapping relationship in the ES cache, that is, remember a lot of error correction mapping relationships and error correction results, and can directly display the correct text according to the error correction result or perform error correction according to the error correction mapping relationship during the next error correction. In this way, the accuracy of semantic subsequent error correction and the reach of media resources are greatly improved.

[0083] See also Figure 6 One embodiment of the present application provides a speech recognition error correction device 10, comprising:

[0084] The voice conversion module 11 is used to obtain voice information, convert the voice information into text to be verified and display it;

[0085] An acquisition module 12, configured to acquire text information of the erroneous text and text information of the target text from the voice instruction when receiving a voice instruction for instructing to correct the erroneous text into the target text, wherein the text information at least includes phonetic code information;

[0086] A processing module 13, used for determining the text to be corrected in the text to be verified according to the text information of the error text and the text information of the target text;

[0087] The acquisition module 12 is also used to acquire the error correction mapping relationship between the text to be corrected and the target text according to the text information of the error text and the text information of the target text;

[0088] The error correction module 14 is used to correct the text to be corrected into the target text based on the error correction mapping relationship.

[0089] The acquisition module 12 is specifically used to obtain the error correction mapping relationship between the text to be corrected and the target text from multiple stored error correction mapping relationships based on the text information of the erroneous text and the text information of the target text; when there is no error correction mapping relationship between the text to be corrected and the target text in the multiple stored error correction mapping relationships, establish the error correction mapping relationship between the text to be corrected and the target text based on the text information of the erroneous text and the text information of the target text.

[0090] The phonetic-graphic code information includes pinyin information and glyph information. The processing module 13 is specifically used to generate the phonetic-graphic code information of each character in the text to be verified; determine the similarity between the pinyin information of each character in the text to be verified and the pinyin information of the target text through a phonetic-graphic code similarity algorithm; and determine that the text composed of characters in the text to be verified whose similarity reaches a preset similarity is the text to be corrected.

[0091] The acquisition module 12 is specifically used to establish a mapping relationship between the glyph information of the erroneous text and the glyph information of the target text; based on the pinyin information of the target text or the pinyin information of the text to be corrected, and the mapping relationship between the glyph information of the erroneous text and the glyph information of the target text, establish a correction mapping relationship between the text to be corrected and the target text.

[0092] The speech recognition error correction device 10 also includes a storage module 15, which is used to obtain text information of multiple corrected texts and text information of texts to be corrected corresponding to the multiple corrected texts; and establish and store the multiple error correction mapping relationships based on the text information of the multiple corrected texts and the text information of the texts to be corrected corresponding to the multiple corrected texts.

[0093] The acquisition module 12 specifically converts the voice instruction into an instruction text; and acquires text information of the error text and text information of the target text indicated by the instruction text through a sequence labeling algorithm.

[0094] The storage module 15 is also used to store the text to be verified when the received voice instruction is not used to instruct to correct the erroneous text into the target text.

[0095] See also Figure 7 One embodiment of the present application provides an electronic device 20, which includes a processor 21 and a memory 22 connected to the processor. The memory 22 stores computer-executable instructions, and the processor 21 executes the computer-executable instructions stored in the memory to implement the speech recognition error correction method described in any of the above embodiments.

[0096] The present application also provides a computer-readable storage medium, which stores computer execution instructions. When the instructions are executed, the computer execution instructions are executed by a processor to implement the speech recognition error correction method provided in any of the above embodiments.

[0097] The present application also provides a computer program product, including a computer program, which, when executed by a processor, implements the speech recognition error correction method provided in any of the above embodiments.

[0098] It should be noted that the computer-readable storage medium may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic random access memory (FRAM), a flash memory, a magnetic surface memory, an optical disk, or a compact disc read-only memory (CD-ROM), etc. It may also be various electronic devices including one or any combination of the above memories, such as mobile phones, computers, tablet devices, personal digital assistants, etc.

[0099] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.

[0100] The serial numbers of the above-mentioned embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.

[0101] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course, by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, a magnetic disk, or an optical disk), and includes a number of instructions for a terminal device (which can be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in each embodiment of the present application.

[0102] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0103] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0104] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the computer or other programmable device. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0105] The above are only preferred embodiments of the present application, and are not intended to limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.

Claims

1. A speech recognition error correction method, It is characterized in that include: Acquire voice information, convert the voice information into text to be verified and display it; When receiving a voice instruction for instructing to correct the erroneous text into a target text, acquiring text information of the erroneous text and text information of the target text from the voice instruction, the text information at least including phonetic code information; the phonetic code information includes pinyin information and glyph information; Determining the text to be corrected in the text to be verified according to the text information of the erroneous text and the text information of the target text; Acquire an error correction mapping relationship between the text to be corrected and the target text according to the text information of the error text and the text information of the target text; the error correction mapping relationship includes the mapping relationship between the pinyin information of the target text or the pinyin information of the text to be corrected, the glyph information of the error text and the glyph information of the target text; Correcting the text to be corrected into the target text based on the error correction mapping relationship; The step of determining the text to be corrected in the text to be verified according to the text information of the error text and the text information of the target text comprises: Generate phonetic and graphic code information for each word in the text to be verified; Determine the similarity between the pinyin information of each word in the text to be verified and the pinyin information of the target text by using a phonetic code similarity algorithm; Determine that the text composed of characters whose similarity reaches a preset similarity in the text to be verified is the text to be corrected.

2. The method according to claim 1, It is characterized in that The acquiring the error correction mapping relationship between the text to be corrected and the target text according to the text information of the error text and the text information of the target text comprises: According to the text information of the erroneous text and the text information of the target text, obtaining an error correction mapping relationship between the text to be corrected and the target text from a plurality of stored error correction mapping relationships; When there is no error correction mapping relationship between the text to be corrected and the target text in the stored multiple error correction mapping relationships, an error correction mapping relationship between the text to be corrected and the target text is established according to the text information of the erroneous text and the text information of the target text.

3. The method according to claim 2, It is characterized in that The step of establishing an error correction mapping relationship between the text to be corrected and the target text according to the text information of the erroneous text and the text information of the target text includes: Establishing a mapping relationship between the glyph information of the erroneous text and the glyph information of the target text; According to the pinyin information of the target text or the pinyin information of the text to be corrected, and the mapping relationship between the glyph information of the erroneous text and the glyph information of the target text, an error correction mapping relationship between the text to be corrected and the target text is established.

4. The method according to claim 2, It is characterized in that Also includes: Acquire text information of a plurality of error-corrected texts and text information of texts to be corrected corresponding to the plurality of error-corrected texts; The multiple error correction mapping relationships are established and stored according to the text information of the multiple error-corrected texts and the text information of the texts to be corrected corresponding to the multiple error-corrected texts.

5. The method according to claim 1, It is characterized in that The acquiring the text information of the error text and the text information of the target text from the voice instruction comprises: Converting the voice command into command text; The text information of the error text and the text information of the target text indicated by the instruction text are obtained through a sequence labeling algorithm.

6. The method according to claim 1, It is characterized in that Also includes: When the received voice instruction is not for instructing to correct the erroneous text into the target text, the text to be verified is stored.

7. A speech recognition error correction device, It is characterized in that include: A voice conversion module, used to obtain voice information, convert the voice information into text to be verified and display it; An acquisition module, for, when receiving a voice instruction for instructing to correct an erroneous text into a target text, acquiring text information of the erroneous text and text information of the target text from the voice instruction, wherein the text information at least includes phonetic code information; the phonetic code information includes pinyin information and glyph information; A processing module, used for determining the text to be corrected in the text to be verified according to the text information of the error text and the text information of the target text; The acquisition module is also used to acquire an error correction mapping relationship between the text to be corrected and the target text according to the text information of the error text and the text information of the target text; the error correction mapping relationship includes the mapping relationship between the pinyin information of the target text or the pinyin information of the text to be corrected, the glyph information of the error text and the glyph information of the target text; An error correction module, used for correcting the text to be corrected into the target text based on the error correction mapping relationship; The processing module is specifically used to generate the phonetic and graphical code information of each word in the text to be verified; Determine the similarity between the pinyin information of each word in the text to be verified and the pinyin information of the target text by using a phonetic code similarity algorithm; Determine that the text composed of characters whose similarity reaches a preset similarity in the text to be verified is the text to be corrected.

8. The device according to claim 7, It is characterized in that The acquisition module is specifically used for: According to the text information of the erroneous text and the text information of the target text, obtaining an error correction mapping relationship between the text to be corrected and the target text from a plurality of stored error correction mapping relationships; When there is no error correction mapping relationship between the text to be corrected and the target text in the stored multiple error correction mapping relationships, an error correction mapping relationship between the text to be corrected and the target text is established according to the text information of the erroneous text and the text information of the target text.

9. An electronic device, It is characterized in that include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the speech recognition error correction method according to any one of claims 1 to 6.

10. A computer-readable storage medium, It is characterized in that The computer-readable storage medium stores computer-executable instructions, and when the instructions are executed, the computer executes the speech recognition error correction method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Method and user equipment for voice amending

    CN103207769A

  • Speech recognition error correction method and device based on artificial intelligence and storage medium

    CN107220235A

  • Identification correction method and device

    CN108121455A