Error correction method and apparatus for speech recognition result
By receiving and judging the second voice input to automatically correct the voice recognition results, the problem of secondary editing caused by voice recognition errors is solved, thus improving input efficiency and user experience.
Patent Information
- Application Number
- PCT/CN2025/075146
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-17
- Filing Date
- 2025-01-26
- Publication Date
- 2025-12-26
AI Technical Summary
Even with high accuracy, existing speech recognition technologies may still produce errors, requiring users to use additional tools for secondary editing, which impacts input efficiency and user experience.
By receiving a second voice input, performing voice recognition and determining whether it is an error correction command, and revising the first voice recognition result according to the second voice input, automatic error correction is achieved.
It can automatically correct speech recognition errors without the need for tools such as a mouse or keyboard, improving the efficiency of speech recognition error correction and enhancing the user experience.
Smart Images

Figure CN2025075146_26122025_PF_FP_ABST
Abstract
Description
A method and apparatus for correcting speech recognition results
[0001] Cross-references to related applications
[0002] This application claims priority to Chinese Patent Application No. 202410781142.X, filed on June 17, 2024, entitled "A Method and Apparatus for Correcting Speech Recognition Results", the entire contents of which are incorporated herein by reference. Technical Field
[0003] This application relates to the field of speech recognition technology, and in particular to a method and apparatus for correcting speech recognition results. Background Technology
[0004] To make operating terminal devices more convenient and efficient for users, many terminal devices and their installed applications now support voice input. On terminal devices that support voice input, users only need to input voice data to easily issue various commands and convey information, greatly simplifying the operation process and improving the user experience. Summary of the Invention
[0005] In view of this, embodiments of this application provide a method and apparatus for correcting speech recognition results, which improves the efficiency of correcting speech recognition results.
[0006] To achieve the above objectives, the technical solutions provided in this application are as follows:
[0007] In a first aspect, embodiments of this application provide a method for correcting speech recognition results, including:
[0008] Display the first text, which is the speech recognition result obtained by performing speech recognition on at least one speech input;
[0009] Receive second voice input;
[0010] Speech recognition is performed on the second voice input to obtain the second text;
[0011] Based on the first text and the second text, determine whether the second voice input is an error correction command;
[0012] In response to the second voice input being an error correction command, the first text is revised based on the second text to obtain the third text.
[0013] Secondly, embodiments of this application provide an error correction device for speech recognition results, comprising:
[0014] The display module is used to display first text, which is the speech recognition result obtained by performing speech recognition on at least one voice input.
[0015] The user input module is used to receive a second voice input.
[0016] The speech recognition module is used to perform speech recognition on the second speech input to obtain the second text;
[0017] The discrimination module is used to determine whether the second voice input is an error correction command based on the first text and the second text;
[0018] The processing module is configured to, in response to the second voice input being an error correction instruction, revise the first text according to the second text to obtain the third text.
[0019] Thirdly, embodiments of this application provide an electronic device, including: a memory and a processor, wherein the memory is used to store a computer program and the processor is used to cause the electronic device to implement the error correction method for speech recognition results described in any of the above embodiments when executing the computer program.
[0020] Fourthly, embodiments of this application provide a computer-readable storage medium that, when executed by a computing device, causes the computing device to implement the error correction method for speech recognition results described in any of the above embodiments.
[0021] Fifthly, embodiments of this application provide a computer program product that, when run on a computer, enables the computer to implement the error correction method for speech recognition results described in any of the above embodiments. Attached Figure Description
[0022] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0023] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the accompanying drawings that need to be called in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0024] Figure 1 is one of the flowcharts of the speech recognition result error correction method provided in the embodiment of this application;
[0025] Figure 2 is a schematic diagram of the user interface for displaying the first text provided in an embodiment of this application;
[0026] Figure 3 is a schematic diagram of the interface for voice input provided in an embodiment of this application;
[0027] Figure 4 is a second flowchart of the error correction method for speech recognition results provided in an embodiment of this application;
[0028] Figure 5 is a schematic diagram of one scenario of the error correction method for speech recognition results provided in the embodiments of this application;
[0029] Figure 6 is a second scenario illustration of the error correction method for speech recognition results provided in the embodiments of this application;
[0030] Figure 7 is a flowchart of the third step of the speech recognition result error correction method provided in the embodiment of this application;
[0031] Figure 8 is a third scenario illustration of the error correction method for speech recognition results provided in the embodiments of this application;
[0032] Figure 9 is a schematic diagram of the speech recognition result correction device provided in the embodiment of this application;
[0033] Figure 10 is a schematic diagram of the hardware structure of the electronic device provided in an embodiment of this application. Detailed Implementation
[0034] To better understand the above-mentioned objectives, features, and advantages of this application, the solution of this application will be further described below. It should be noted that, unless otherwise specified, the embodiments and features described in these embodiments can be combined with each other.
[0035] Many specific details are set forth in the following description in order to provide a full understanding of this application, but this application may also be implemented in other ways different from those described herein. Obviously, the embodiments in the specification are only some embodiments of this application, and not all embodiments.
[0036] In the embodiments of this application, the terms "exemplary" or "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design solutions. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a specific manner. Furthermore, in the description of the embodiments of this application, unless otherwise stated, "multiple" means two or more.
[0037] Although current speech recognition technology has achieved a fairly high level of accuracy, errors can still occur due to various factors such as the input environment and the accuracy of the user's pronunciation. When errors exist, users often need to use tools such as a mouse, keyboard, or virtual keyboard to edit the speech recognition results and correct them. However, this secondary editing inevitably reduces input efficiency, negatively impacting the user experience.
[0038] This application provides a method for correcting speech recognition results. Referring to FIG1, the method for correcting speech recognition results includes the following steps:
[0039] S101, Display the first text.
[0040] The first text is the speech recognition result obtained by performing speech recognition on at least one speech input.
[0041] Referring to Figure 2, Figure 2 illustrates the application of the speech recognition result correction method provided in this embodiment of the application to an instant messaging scenario. As shown in Figure 2, in an instant messaging scenario, the user inputs instant messaging messages via speech-to-text conversion. When the terminal device receives the user's speech input, it performs speech recognition on the received speech input and displays the speech recognition result in the message editing area 200 of the instant messaging application's user interface. The text content displayed in the message editing area 200, "The weather is so nice today, let's go have some fun," is the first text obtained by performing speech recognition on at least one speech input in this embodiment of the application.
[0042] S102, Receive second voice input.
[0043] Referring to Figure 3, the example in Figure 3 illustrates the application of the speech recognition result correction method provided in this embodiment to an instant messaging scenario. As shown in Figure 3, the user interface of the instant messaging application includes a voice input control 31. The implementation of receiving second voice input may include: firstly, receiving a user's click operation on the voice input control 31, and displaying a voice recording control 32 in response to the user's click operation on the voice input control 31; during the long-press operation on the voice recording control 32, receiving the user's voice input through an audio receiving device such as a microphone on the terminal device.
[0044] S103. Perform speech recognition on the second voice input to obtain the second text.
[0045] In some embodiments, the second voice input line can be speech-recognized using Automatic Speech Recognition (ASR) technology to obtain the second text.
[0046] S104. Based on the first text and the second text, determine whether the second voice input is an error correction instruction.
[0047] In some embodiments, determining whether the second voice input is an error correction instruction based on the first text and the second text includes:
[0048] The first text and the second text are input into a pre-trained judgment model, and the output of the judgment model is used to determine whether the second voice input is an error correction instruction.
[0049] In some embodiments, the discriminant model and the speech recognition model for recognizing the speech input can be independent models. When the discriminant model and the speech recognition model are independent models, the discriminant model can be a model trained on a binary classification machine learning model based on a first training dataset. The first training dataset includes multiple sets of sample data, each set of sample data including a first sample text, a second sample text, and label information representing whether the speech input corresponding to the second text is an error correction instruction. The output of the discriminant model is a binary classification result representing whether the second speech input is an error correction instruction.
[0050] In some embodiments, the discriminant model can be nested within a speech recognition model that performs speech recognition on the speech input. When the discriminant model is nested within the speech recognition model, the speech recognition model can be a model trained on a pre-defined machine learning model based on a second training dataset. The first training dataset includes multiple sets of sample data, each set of sample data including a first sample speech input, a second sample speech input, and label information representing whether the second sample speech input is an error correction instruction. The output of the speech recognition model includes two parts: a speech recognition result of the second speech input and a binary classification result indicating whether the second speech input is an error correction instruction.
[0051] In step S104 above, if it is determined that the second voice input is an error correction command, then the following step S105 is executed:
[0052] S105. In response to the second voice input being an error correction instruction, the first text is revised according to the second text to obtain the third text.
[0053] The speech recognition result correction method provided in this application embodiment, when displaying a first text obtained by speech recognition of at least one speech input and receiving a second speech input, firstly performs speech recognition on the second speech input to obtain a second text, then determines whether the second speech input is a correction instruction based on the first text and the second text, and in response to the second speech input being a correction instruction, revises the first text based on the second text to obtain a third text. Since the speech recognition result correction method provided in this application embodiment can determine whether the current speech input is a correction instruction based on the speech recognition results of previous speech inputs and the speech recognition result of the current speech input, and if it is determined that the current speech input is a correction instruction, revises the speech recognition results of previous speech inputs based on the speech recognition results of the current speech input, the speech recognition result correction method provided in this application embodiment can avoid the need to use tools such as a mouse, keyboard, or virtual keyboard to edit the speech recognition results to eliminate errors, thereby improving the efficiency of speech recognition result correction.
[0054] As an extension and refinement of the above embodiments, this application provides another method for correcting speech recognition results. Referring to FIG4, the method for correcting speech recognition results includes the following steps:
[0055] S401, Display the first text.
[0056] The first text is the speech recognition result obtained by performing speech recognition on at least one speech input.
[0057] S402, Receive second voice input.
[0058] S403. Perform speech recognition on the second voice input to obtain the second text.
[0059] S404. Obtain the voice intent of the second voice input based on the second text.
[0060] In some embodiments, the second text input can be pre-trained with an intent recognition model, and the intent recognition model can be used as the speech intent of the second speech input.
[0061] In some embodiments, the intent recognition model can be a semantic understanding model built on a Large Language Model (LMM), a Natural Language Understanding Model (NLU), or the like.
[0062] S405. Determine whether the voice intent includes a text modification intent.
[0063] That is, determining whether the speech intent of the second speech input includes the semantics of modifying the text.
[0064] In step S405 above, if the voice intent includes a text modification intent, then step S406 is executed as follows:
[0065] S406. In response to the voice intent including a text modification intent, determine whether the first text includes the character to be corrected indicated by the text modification intent.
[0066] In some embodiments, determining whether the first text includes the character to be corrected indicated by the text modification intention includes: converting the character to be corrected indicated by the text modification intention and each character in the first text into pinyin, and determining whether the pinyin obtained by converting each character in the first text includes the pinyin corresponding to the character to be corrected.
[0067] In some embodiments, during the process of determining whether the pinyin obtained from converting each character in the first text includes the pinyin corresponding to the character to be corrected, the differences between front and back nasal sounds, as well as the differences between alveolar and retroflex sounds, are ignored. For example, the difference between "tan" and "tang" is ignored. Another example is the difference between "zong" and "zhong".
[0068] In step S406 above, if it is determined that the first text includes the character to be corrected in the text modification intention, then step S407 is executed as follows:
[0069] S407. In response to the first text including the character to be corrected indicated by the text modification intention, determine that the second voice input is an error correction instruction.
[0070] S408. Determine the character to be corrected in the first text and the target character corresponding to the character to be corrected based on the second text.
[0071] The method of determining the character to be corrected in the first text based on the second text may include: obtaining the text modification intention of the second text; obtaining the character to be corrected indicated by the text modification intention; and determining the correction result corresponding to the character to be corrected in the text modification intention as the target character corresponding to the character to be corrected.
[0072] S409. Replace the character to be corrected with the target character to obtain the third text.
[0073] S410, Replace the first text with the third text.
[0074] Exemplarily, as shown in FIG. 5, the second voice input in FIG. 5 is an error correction instruction. And taking the character to be corrected as "嗨编" and the target character corresponding to the character to be corrected as "海边" as an example. As shown in FIG. 5, the text content of the first text 51 is "The weather is really nice today. Let's go to 嗨编 to play", the text content of the second text 52 is "Change 嗨编 to 海边", and the text content of the third text 53 obtained by replacing the character to be corrected "嗨编" with the corresponding target character "海边" is "The weather is really nice today. Let's go to the seaside to play", and the first text 51 is replaced and displayed as the third text 53.
[0075] It should be noted that in some embodiments, the characters to be corrected determined according to the second text include multiple ones, and each character to be corrected in the multiple characters to be corrected corresponds to a target character respectively. Then the above step S409 includes: replacing each character to be corrected in the multiple characters to be corrected with the corresponding target character respectively to obtain the third text. For example: the text content of the first text is "The weather is 针好 today. Let's go to the seaside to play", the text content of the second text is "Change 针好 and 嗨编 to 真好 and 海边", then the characters to be corrected determined are: "针好" and "嗨编", and the target character corresponding to "针好" is "真好", the target character corresponding to "嗨编" is "海边", and the text content of the third text obtained by replacing each character to be corrected in the multiple characters to be corrected with the corresponding target character is "The weather is really nice today. Let's go to the seaside to play".
[0076] If it is determined in the above step S405 that the voice intention does not include the text modification intention, or it is determined in the above step S406 that the first text does not include the character to be corrected indicated by the text modification intention, then the following step S411 is executed:
[0077] S411. In response to the voice intention not including the text modification intention or the first text not including the character to be corrected indicated by the text modification intention, determine that the second voice input is not an error correction instruction.
[0078] S412. In response to the second voice input not being an error correction instruction, append the second text to the tail of the first text to obtain the fourth text.
[0079] S413. Replace and display the first text with the fourth text.
[0080] For example, referring to Figure 6, Figure 6 illustrates an example where the second voice input is not an error correction command. As shown in Figure 6, the text content of the first text 61 is "The weather is so nice today, let's go to the beach," and the text content of the second text 62 is "Remember to bring your sun hat." By appending the second text 62 to the end of the first text 61, the resulting fourth text 63 has the text content "The weather is so nice today, let's go to the beach, remember to bring your sun hat," and the first text 61 is replaced by the fourth text 63.
[0081] In some embodiments, referring to FIG7, based on the embodiment shown in FIG4, after replacing the first text with the third text or replacing the first text with the fourth text, the speech recognition result correction method provided in this application embodiment further includes:
[0082] S71, Receive text output operation.
[0083] S72. In response to the text output operation, clear the currently displayed text.
[0084] For example, referring to Figure 8, Figure 8 shows an example where the currently displayed text 81 reads "The weather is so nice today, let's go to the beach." When the user's operation on the send control 82 is received, the currently displayed text 81 is sent, and the displayed text 81 is cleared.
[0085] S73, Receive third voice input.
[0086] S74. Recognize the third voice input to obtain the fifth text.
[0087] S75. Display the fifth text.
[0088] That is, if the speech recognition result obtained from the speech input is not included before the current speech input, the speech recognition result of the current speech input will be displayed directly, and the judgment of whether the current speech input is an error correction command will no longer be made.
[0089] Based on the same inventive concept, as an implementation of the above method, this application embodiment also provides a speech recognition result error correction device. This embodiment corresponds to the aforementioned method embodiment. For ease of reading, this embodiment will not repeat the details of the aforementioned method embodiment one by one, but it should be clear that the speech recognition result error correction device in this embodiment can correspondingly implement all the contents of the aforementioned method embodiment.
[0090] This application provides an error correction device for speech recognition results. Figure 9 is a schematic diagram of the structure of the error correction device for speech recognition results. As shown in Figure 9, the error correction device 900 for speech recognition results includes:
[0091] Display module 91 is used to display first text, which is the speech recognition result obtained by performing speech recognition on at least one voice input;
[0092] User input module 92 is used to receive second voice input;
[0093] The speech recognition module 93 is used to perform speech recognition on the second speech input to obtain the second text;
[0094] The discrimination module 94 is used to determine whether the second voice input is an error correction instruction based on the first text and the second text;
[0095] Processing module 95 is configured to, in response to the second voice input being an error correction instruction, revise the first text according to the second text to obtain the third text.
[0096] As an optional implementation of this application, the display module 91 is further configured to replace the first text with the third text after revising the first text according to the second text to obtain the third text.
[0097] As an optional implementation of this application, the processing module 95 is specifically used to determine the character to be corrected in the first text and the target character corresponding to the character to be corrected based on the second text; and to replace the character to be corrected with the target character to obtain the third text.
[0098] As an optional implementation of this application, the processing module 95 is further configured to, in response to determining that the second voice input is not an error correction instruction, append the second text to the end of the first text to obtain the fourth text;
[0099] The display module 91 is also used to replace the first text with the fourth text.
[0100] As an optional implementation of this application, the discrimination module 94 is specifically used to obtain the voice intent of the second voice input based on the second text; determine whether the voice intent includes a text modification intent; in response to the voice intent including a text modification intent, determine whether the first text includes the character to be corrected indicated by the text modification intent; in response to the first text including the character to be corrected indicated by the text modification intent, determine that the second voice input is an error correction instruction.
[0101] As an optional implementation of this application, the user input module 92 is further configured to receive text output operations;
[0102] The processing module 95 is also used to clear the currently displayed text in response to the text output operation.
[0103] As an optional implementation of this application, the user input module 92 is further configured to receive a third voice input after clearing the currently displayed text in response to the text output operation;
[0104] The speech recognition module 93 is also used to recognize the third speech input in order to obtain the fifth text;
[0105] The display module 91 is also used to display the fifth text.
[0106] The speech recognition result correction device provided in this application embodiment can execute the speech recognition result correction method provided in any of the above embodiments. Its implementation principle and technical effect are similar, and will not be described again here.
[0107] Based on the same inventive concept, this application also provides an electronic device. Figure 10 is a schematic diagram of the structure of the electronic device provided in this application embodiment. As shown in Figure 10, the electronic device provided in this embodiment includes: a memory 101 and a processor 102. The memory 101 is used to store a computer program, and the processor 102 is used to execute the error correction method for the speech recognition results provided in the above embodiment when executing the computer program.
[0108] Based on the same inventive concept, this application also provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the computing device implements the error correction method for the speech recognition results provided in the above embodiments.
[0109] Based on the same inventive concept, this application also provides a computer program product that, when run on a computer, enables the computing device to implement the error correction method for speech recognition results provided in the above embodiments.
[0110] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program code.
[0111] The processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor.
[0112] Memory may include non-persistent memory in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0113] Computer-readable media include both permanent and non-permanent, removable and non-removable storage media. Storage media can store information using any method or technology; the information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0114] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some or all of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.
Claims
1. A method for correcting speech recognition results, comprising: Display the first text, which is the speech recognition result obtained by performing speech recognition on at least one speech input; Receive second voice input; Speech recognition is performed on the second voice input to obtain the second text; Based on the first text and the second text, determine whether the second voice input is an error correction command; In response to the second voice input being an error correction command, the first text is revised based on the second text to obtain the third text.
2. The method according to claim 1, further comprising: Replace the first text with the third text.
3. The method of claim 1, wherein revising the first text according to the second text to obtain the third text comprises: Based on the second text, determine the character to be corrected in the first text and the target character corresponding to the character to be corrected; The character to be corrected is replaced with the target character to obtain the third text.
4. The method according to claim 1, further comprising: In response to the second voice input not being an error correction command, the second text is appended to the end of the first text to obtain the fourth text; Replace the first text with the fourth text.
5. The method according to claim 1, wherein determining whether the second voice input is an error correction instruction based on the first text and the second text includes: The voice intent of the second voice input is obtained from the second text; Determine whether the voice intent includes a text modification intent; In response to the voice intent including a text modification intent, determine whether the first text includes the character to be corrected indicated by the text modification intent; In response to the first text including the character to be corrected as indicated by the text modification intention, the second voice input is determined to be an error correction instruction.
6. The method according to claim 2 or 4, further comprising: Receive text output operations; In response to the text output operation, clear the currently displayed text.
7. The method of claim 6, wherein after clearing the currently displayed text in response to the text output operation, the method further comprises: Receive third-party voice input; The third voice input is recognized to obtain the fifth text; The fifth text is displayed.
8. An error correction device for speech recognition results, comprising: The display module is used to display first text, which is the speech recognition result obtained by performing speech recognition on at least one voice input. The user input module is used to receive a second voice input. The speech recognition module is used to perform speech recognition on the second speech input to obtain the second text; The discrimination module is used to determine whether the second voice input is an error correction command based on the first text and the second text; The processing module is configured to revise the first text based on the second text to obtain a third text when it is determined that the second voice input is an error correction instruction.
9. An electronic device, comprising: A memory and a processor, wherein the memory is used to store a computer program and the processor is used to cause the electronic device to implement the error correction method for the speech recognition results according to any one of claims 1-7 when executing the computer program.
10. A computer-readable storage medium storing a computer program that, when executed by a computing device, causes the computing device to implement the error correction method for speech recognition results according to any one of claims 1-7.
11. A computer program product, wherein when the computer program product is run on a computer, the computer implements the error correction method for the speech recognition result according to any one of claims 1-7.
Citation Information
Patent Citations
Speech recognition error correction method and device based on artificial intelligence and storage medium
CN107220235A
Voice recognition error correction method, mobile terminal and computer readable storage medium
CN111243593A
Voice interaction error correction method and device
CN112669833A
Speech recognition error correction method and device, electronic equipment and storage medium
CN114360549A
Speech processing apparatus and method
KR1020120110751A