Text recognition method and device, electronic equipment and readable storage medium
By setting text conditions and scoring conditions to make two judgments on the user feedback text, the problem of the inability to effectively identify the correctness of user feedback in the existing technology is solved, thereby improving the accuracy of speech recognition and user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-29
- Publication Date
- 2026-03-31
AI Technical Summary
Existing speech recognition technology uses overly simplistic methods to determine the validity of user feedback text, failing to effectively identify cases where the user's feedback is completely correct, resulting in a high error rate in the text.
The user feedback text is evaluated twice by setting text conditions and scoring conditions. First, feedback text that meets the text format is filtered. Then, its validity is judged by the alignment results of the acoustic model, including character format, text length, edit distance and scoring conditions.
It improved the accuracy of user feedback text, reduced the error rate of the manuscript, and ensured the accuracy of the manuscript and the user experience.
Smart Images

Figure CN115269777B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of speech recognition technology, and more specifically, to a text recognition method, apparatus, electronic device, and readable storage medium. Background Technology
[0002] In speech recognition applications, one scenario involves converting audio into text and then displaying the transcript to the user for easy reading. However, speech-recognized text often contains errors. To improve user experience, a feedback function can be enabled, allowing users to correct any recognized errors.
[0003] To determine the validity of user feedback text, perfectly correct feedback text can be promptly updated in the document, reducing the error rate. Existing technology uses edit distance to determine the similarity between the feedback text and the original text; however, this method is too simplistic and ineffective, and it cannot handle cases where the user feedback is completely correct. Summary of the Invention
[0004] One objective of this invention is to provide a text recognition method, apparatus, electronic device, and readable storage medium for accurately identifying the validity of user feedback content. Embodiments of this invention can be implemented as follows:
[0005] In a first aspect, the present invention provides a text recognition method, the method comprising: acquiring feedback text and original text; wherein, the original text is text obtained by speech recognition of target audio; the feedback text is text based on the original text after error correction processing; if the feedback text does not meet preset text conditions, the feedback text is determined to be invalid feedback text; the text conditions are conditions characterizing the validity of the text form of the feedback text; if the feedback text meets the text conditions, the scores corresponding to the feedback text and the original text are determined; the scores characterize the alignment results of the feedback text, the original text, and the real text in the target audio, respectively; if the scores corresponding to the feedback text and the original text meet preset text scoring conditions, the feedback text is determined to be valid feedback text; if the scores corresponding to the feedback text and the original text do not meet the text scoring conditions, the feedback text is determined to be invalid feedback text.
[0006] In a second aspect, the present invention provides a text recognition device, comprising: an acquisition module and a recognition module;
[0007] The acquisition module is used to acquire feedback text and original text; wherein, the original text is text obtained by speech recognition of the target audio; the feedback text is text based on the original text after error correction processing; the recognition module is used to: determine the feedback text as invalid feedback text if the feedback text does not meet preset text conditions; the text conditions are conditions that characterize the text form of the feedback text as valid; determine the scores corresponding to the feedback text and the original text respectively if the feedback text meets the text conditions; the scores characterize the alignment results of the feedback text, the original text and the real text in the target audio respectively; determine the feedback text as valid feedback text if the scores corresponding to the feedback text and the original text respectively meet preset text scoring conditions; determine the feedback text as invalid feedback text if the scores corresponding to the feedback text and the original text do not meet the text scoring conditions.
[0008] Thirdly, the present invention provides an electronic device including a processor and a memory, the memory storing a computer program executable by the processor, the processor being able to execute the computer program to implement the method described in the first aspect.
[0009] Fourthly, the present invention provides a readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described in the first aspect.
[0010] The present invention provides a text recognition method, apparatus, electronic device, and readable storage medium. The method includes: after obtaining the original text recognized from speech and the feedback text provided by the user, first determining whether the feedback text meets text conditions. If it does not meet the text conditions, the feedback text is directly determined to be invalid. If it meets the text conditions, the method can continue to determine whether the feedback text and the original text meet scoring conditions. That is, first determining the scores corresponding to the feedback text and the original text respectively, and then, if the scores corresponding to the feedback text and the original text respectively meet preset text scoring conditions, the feedback text is determined to be valid feedback text. In this invention, the feedback text is first initially judged based on text conditions. If the feedback text meets the text form requirements, a second judgment is made on the feedback text, that is, whether the alignment results of the feedback text, the original text, and the real text in the target audio respectively meet the conditions. This invention performs two judgments on the feedback text through text conditions and scoring conditions, thereby accurately determining the validity of the feedback text. Attached Figure Description
[0011] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0012] Figure 1 This is a schematic diagram of a speech recognition scenario provided in an embodiment of the present invention;
[0013] Figure 2 A structural block diagram of an electronic device provided in an embodiment of the present invention;
[0014] Figure 3 A schematic flowchart illustrating the text recognition method provided in an embodiment of the present invention;
[0015] Figure 4 A flowchart illustrating the process of determining whether feedback text meets text conditions, provided in an embodiment of the present invention;
[0016] Figure 5 A schematic diagram illustrating the principle of determining text scores in an embodiment of the present invention;
[0017] Figure 6 A schematic flowchart of step S330 provided in an embodiment of the present invention;
[0018] Figure 7 A schematic flowchart illustrating another text recognition method provided in an embodiment of the present invention;
[0019] Figure 8 A flowchart illustrating the process of determining the confidence level of feedback text, provided in an embodiment of the present invention;
[0020] Figure 9 This is a functional block diagram of the text recognition device provided in the embodiments of this application. Detailed Implementation
[0021] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.
[0022] Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.
[0023] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.
[0024] In the description of this invention, it should be noted that if terms such as "upper," "lower," "inner," or "outer" are used to indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings, or the orientation or positional relationship in which the product of this invention is usually placed, they are only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this invention.
[0025] Furthermore, the terms "first" and "second" are used only to distinguish descriptions and should not be interpreted as indicating or implying relative importance.
[0026] It should be noted that, where there is no conflict, the features in the embodiments of the present invention can be combined with each other.
[0027] Speech recognition technology can provide users with voice content recognition services. This technology can be applied to various scenarios, such as speech-to-text, voice wake-up, and human-computer interaction.
[0028] Please see Figure 1 , Figure 1 This is a schematic diagram of a scenario for the text recognition method provided in an embodiment of the present invention. The scenario includes a terminal device 101 and a server 102, which are connected via a network.
[0029] Terminal device 101 can send a voice recognition request to server 102. The voice recognition request contains the voice to be recognized. After obtaining the audio, server 102 can recognize the original text corresponding to the audio through voice recognition technology. Then, server 102 can send the original text to terminal device 101, and terminal device 101 can display the original text to the user.
[0030] In practice, the original text obtained from speech recognition may contain errors. To improve the user experience, a feedback function can be enabled for users. After viewing the original text, if the user believes that there are errors in the original text compared with the real text in the speech to be recognized, the user can input feedback text on the terminal device 101. The feedback text is the text after the user corrects the errors in the original text. The terminal device 101 sends this feedback text to the server 102.
[0031] To determine the validity of user feedback text, perfectly correct feedback text can be promptly updated in the document, reducing the error rate. Existing technology uses edit distance to determine the similarity between the feedback text and the original text; however, this method is too simplistic and ineffective, and it cannot handle cases where the user feedback is completely correct.
[0032] To address the aforementioned technical problems, this invention provides a text recognition method. It should be noted that the text recognition method in this invention identifies user-provided feedback. This method not only identifies whether the user-provided text is correct and valid, but also, when the feedback text is valid, further determines the confidence level of the feedback text. This provides a data foundation for subsequent updates to the speech recognition transcript, reducing the transcript error rate.
[0033] Please see Figure 2 , Figure 2 This is a structural block diagram of an electronic device provided in an embodiment of the present invention. The electronic device may be... Figure 1 The server 102 in the text can also be a terminal device 101 (in which case the terminal has text recognition capability), and there is no limitation here.
[0034] like Figure 2 As shown, the electronic device 200 includes a memory 201, a processor 202, and a communication interface 203. The memory 201, processor 202, and communication interface 203 are electrically connected to each other directly or indirectly to realize data transmission or interaction. For example, these components can be electrically connected to each other through one or more communication buses or signal lines.
[0035] The memory 201 can be used to store software programs and modules, such as the instructions / modules of the text recognition device 400 provided in this embodiment of the invention. These can be stored in the memory 201 in the form of software or firmware, or embedded in the operating system (OS) of the electronic device 200. The processor 202 executes various functional applications and data processing by executing the software programs and modules stored in the memory 201. The communication interface 203 can be used for signaling or data communication with other node devices.
[0036] The memory 201 may be, but is not limited to, random access memory (RAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), etc.
[0037] Processor 202 can be an integrated circuit chip with signal processing capabilities. Processor 202 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0038] Understandable. Figure 2 The structure shown is for illustrative purposes only; the electronic device 200 may also include more than [other components]. Figure 2 The more or fewer components shown, or having the same Figure 2 The different configurations shown. Figure 2 The components shown can be implemented using hardware, software, or a combination thereof.
[0039] Please see Figure 3 , Figure 3 This is a schematic flowchart illustrating the text recognition method provided in an embodiment of the present invention. The execution entity of this method may be... Figure 2 The electronic device shown may include the following methods:
[0040] S310, retrieve the feedback text and the original text.
[0041] The original text is the text obtained by speech recognition of the target audio; the feedback text is the text that has undergone error correction based on the original text.
[0042] S320, if the feedback text does not meet the preset text conditions, then the feedback text is determined to be invalid feedback text.
[0043] Among them, the text condition is the condition that characterizes the text form of the feedback text as valid.
[0044] S330, If the feedback text meets the text conditions, then determine the scores corresponding to the feedback text and the original text respectively;
[0045] The score represents the alignment results of the feedback text and the original text with the real text in the target audio.
[0046] S340, if the scores corresponding to the feedback text and the original text respectively meet the preset text scoring conditions, then the feedback text is determined to be a valid feedback text.
[0047] S350, if the scores corresponding to the feedback text and the original text do not meet the preset text scoring conditions, then the feedback text is determined to be invalid feedback text.
[0048] According to the text recognition method provided in this embodiment of the invention, after obtaining the original text recognized by speech and the feedback text provided by the user, it first determines whether the feedback text meets the text conditions. If it does not meet the text conditions, the feedback text is directly determined to be invalid. If it meets the text conditions, it can continue to determine whether the feedback text and the original text meet the scoring conditions. That is, it first determines the scores corresponding to the feedback text and the original text respectively. Then, if the scores corresponding to the feedback text and the original text respectively meet the preset text scoring conditions, the feedback text is determined to be valid feedback text. In this invention, the feedback text is first judged based on the text conditions. If the feedback text meets the text form, it can be judged a second time, that is, whether the alignment results of the feedback text and the original text with the real text in the target audio meet the conditions. This invention performs two judgments on the feedback text through the text conditions and the scoring conditions, thereby accurately obtaining the validity of the feedback text.
[0049] The following will be combined with the appendix Figure 4 To be continued Figure 6 The steps S310 to S350 described above will be explained in detail.
[0050] In step S310, the feedback text and the original text are obtained.
[0051] In this embodiment of the invention, the original text is the text obtained by speech recognition of the target audio. The original text may be consistent with the real text in the target audio, or it may differ from the real text. The target audio may be the speech to be recognized requested by the user, it may be a speech recorded by the user in real time, or it may be speech pre-stored locally on the terminal device. The feedback text is the text provided by the user, which is the text that the user uses to correct errors by combining the original text and the real text.
[0052] In optional implementations, the feedback text may be manually entered by the user on the terminal device, or it may be text obtained by speech recognition based on the user's recorded voice, or it may be text obtained by text recognition of an image with feedback text uploaded by the user. No limitation is made here.
[0053] In step S320, if the feedback text does not meet the preset text conditions, the feedback text is determined to be invalid feedback text.
[0054] In this embodiment of the invention, the text condition is a condition that characterizes the validity of the text form of the feedback text. The text form includes the text format, text length, and edit distance between the feedback text and the original text. The edit distance is a quantification of the degree of difference between the two strings, the feedback text and the original text.
[0055] In this embodiment of the invention, the text conditions may include: character format conditions, text length conditions, and edit distance conditions. The character format conditions represent the condition that the character format in the feedback text is valid. The text length conditions represent the condition that the text length of the feedback text is valid. The edit distance conditions represent the condition that the edit distance between the feedback text and the original text is valid.
[0056] Based on the above text conditions, in an optional implementation, the method for determining whether the feedback text meets the text conditions can be found in [reference needed]. Figure 4 , Figure 4 A flowchart illustrating the process of determining whether feedback text meets text conditions, provided in an embodiment of the present invention:
[0057] a1 determines whether the format of the characters in the feedback text is a preset character format. If yes, execute a2; otherwise, execute a5.
[0058] In this embodiment of the invention, the preset character format is any one or a combination of the following: Chinese characters; English characters; numbers; punctuation marks.
[0059] a2. Determine whether the length of the feedback text is within the preset text length range. If yes, proceed to step a3; otherwise, proceed to step a5.
[0060] In this embodiment of the invention, the text length range can be set according to the length of the original text. In an optional implementation, assuming the length of the original text is s, the text length range can be set to (0.5s, 1.5s).
[0061] a3 determines whether the editing distance between the feedback text and the original text is within the preset distance range. If yes, execute a4; otherwise, execute a5.
[0062] In this embodiment of the invention, the editing distance can be set according to actual needs, for example, the editing distance can be 10.
[0063] a4, confirm that the feedback text meets the text conditions.
[0064] a5 indicates that the feedback text does not meet the text conditions.
[0065] If the feedback text does not meet any of the character format condition, text length condition, and edit distance condition, then the feedback text is determined to not meet the text condition; if the feedback text meets the character format condition, text length condition, and edit distance condition in sequence, then the feedback text is determined to meet the text condition.
[0066] It should be noted that the execution order between a1 and a3 mentioned above is merely an example and is not a limitation on the execution order between a1 and a3. In other words, the judgment process of character format condition, text length condition and edit distance condition can be carried out simultaneously, or they can be executed in different orders according to different character format condition, text length condition and edit distance condition. No limitation is made here.
[0067] The above text conditions can be used to initially screen the feedback text, eliminating feedback text that does not meet the text format requirements and avoiding unnecessary processing. When the feedback text meets the text format requirements, step S330 can be executed to perform a second screening of the feedback text to determine whether the feedback text is valid.
[0068] In step S330, if the feedback text meets the text conditions, the scores corresponding to the feedback text and the original text are determined respectively.
[0069] The score represents the alignment results of the feedback text, the original text, and the real text in the target audio. The alignment result is obtained by decoding the target audio with the original text and the feedback text respectively.
[0070] In this embodiment of the invention, to determine the scores corresponding to the feedback text and the original text respectively, this embodiment provides an optional implementation method, please refer to... Figure 5 , Figure 5 This is a schematic diagram illustrating the principle of determining text scores in an embodiment of the present invention, as shown below. Figure 5 As shown:
[0071] Reference text: refers to the original text or feedback text;
[0072] Graph Composition: The graph to be aligned and decoded is obtained by performing HCLG operation on a pre-trained acoustic model and a pronunciation dictionary.
[0073] Acoustic Model: The acoustic model calculates the posterior probability of acoustic features belonging to each phoneme. It is trained using over 100 hours of well-pronounced audio. The training process involves: first, segmenting the audio into frames, then extracting 40-dimensional Mel-frequency cepstral coefficients (MFCC) features. After feature extraction, the audio text is expanded into phonemes according to a dictionary, and the acoustic model is trained using a time-delay neural network (TDNN).
[0074] Decoding: Based on the MFCC features of the audio input, combined with the likelihood and composition output by the acoustic model, the Viterbi algorithm is used for decoding, with the aim of selecting the optimal path.
[0075] Output word alignment results: For example, in the sequence "Today is a good day", output the posterior probability and duration of the phonemes in each word, and then calculate the score of each word.
[0076] Combination Figure 5 The schematic diagram shown illustrates one implementation method of step S330 in this embodiment of the invention. Please refer to [link / reference]. Figure 6 , Figure 6 A schematic flowchart of step S330 provided in an embodiment of the present invention:
[0077] S331, input the original text and feedback text into the pre-trained acoustic model respectively to obtain the decoding maps corresponding to the original text and feedback text respectively.
[0078] S332, decode the decoding maps corresponding to the original text and the feedback text respectively, and decode the audio features corresponding to the target audio to obtain the posterior probability and duration of each phoneme contained in the original text and the feedback text respectively.
[0079] S333, based on the posterior probability and duration of each phoneme contained in the original text and the feedback text respectively, obtain the score of each character contained in the original text and the feedback text respectively, as well as the score of the original text and the feedback text respectively.
[0080] In step S333, after obtaining the scores of each character in the original text and the feedback text, the scores of each character can be summed, and then the sum can be divided by the text length to obtain the scores corresponding to the original text and the feedback text respectively.
[0081] After determining the scores obtained by the feedback text and the original text, steps S340 and S350 can be executed to perform a second judgment on the feedback text and finally determine the validity of the feedback text.
[0082] In step S340, if the scores corresponding to the feedback text and the original text respectively meet the preset text scoring conditions, then the feedback text is determined to be a valid feedback text.
[0083] In step S350, if the scores corresponding to the feedback text and the original text do not meet the text scoring conditions, the feedback text is determined to be invalid feedback text.
[0084] In this embodiment of the invention, the text scoring conditions include: the score of the feedback text is greater than the score of the original text, or the score of the original text is greater than the score of the feedback text, and the difference between the score of the original text and the score of the feedback text is within a preset difference range.
[0085] In an optional implementation, the difference range can be set according to actual needs, for example, the difference range can be (0,5).
[0086] After determining that the feedback text is valid, this embodiment of the invention can further determine the confidence level of the feedback text. Therefore, in Figure 3 Based on this, embodiments of the present invention also provide another text recognition method, please refer to [link to relevant documentation]. Figure 7 , Figure 7 A schematic flowchart of another text recognition method provided in an embodiment of the present invention, the method further comprising:
[0087] S360, when the feedback text is determined to be valid, the target character to be edited in the feedback text and the original text is determined based on the editing sequence corresponding to the feedback text and the original text.
[0088] In this embodiment of the invention, the edit sequence includes the edit operation corresponding to the original text and the target character corresponding to the edit operation. The edit operation includes any one or a combination of the following: insertion operation; deletion operation; replacement operation.
[0089] For example, if the original text is "Today is a good day", by inserting "see", replacing "good" with "bad", and then deleting "qi", we can get the feedback text "See today is a bad day". The target characters are "see" in the feedback text, "good" in the original text, "bad" in the feedback text, and "qi" in the original text.
[0090] S370, based on the editing operation corresponding to the target character, determine the character scoring conditions corresponding to the target character.
[0091] Continuing to refer to the above example, among which, the editing operation corresponding to the target character "look" is "insertion", the editing operations corresponding to "good" and "bad" are "replacement", and the editing operation corresponding to "anger" is "deletion".
[0092] In an embodiment of the present invention, the scoring conditions corresponding to the insertion operation, the deletion operation, and the replacement operation are different from each other.
[0093] For the insertion operation, the corresponding character scoring condition is: the score of the target character is greater than the first threshold.
[0094] For the deletion operation, the corresponding character scoring condition is: the score of the target character is less than the second threshold.
[0095] For the replacement operation, the corresponding character scoring condition is: the score of the target character after replacement is greater than the score of the target character before replacement, or, the score of the target character after replacement is less than the score of the target character before replacement, and the score difference between the score of the target character after replacement and the score of the target character before replacement is less than the third threshold; wherein, the target character after replacement is located in the feedback text, and the target character before replacement is located in the original text.
[0096] For the above replacement condition, wherein, the target character after replacement is located in the feedback text, and the target character before replacement is the character at the same character position in the original text as that in the feedback text. For example, continuing to refer to the above example, "good" in the original text and "bad" in the feedback text, among which, "good" is the target character before replacement, and "bad" is the target character after replacement.
[0097] It should be noted that the first threshold, the second threshold, and the third threshold in the embodiment of the present invention can be set according to actual needs. For example, the above first threshold can be: greater than 50 points; the second threshold can be less than 30 points; the third threshold can be: less than 10 points.
[0098] It can be understood that when the score of the inserted character is greater than the first threshold, it indicates that the inserted character improves the confidence of the feedback text. When the score of the deleted character is less than the second threshold, it indicates that deleting the character will not have a great impact on the effectiveness of the feedback text. When the score of the character after replacement is greater than the score of the character after replacement, it indicates that the inserted character improves the confidence of the feedback text, or, the score of the target character after replacement is less than the score of the target character before replacement, and the score difference between the score of the target character after replacement and the score of the target character before replacement is less than the third threshold, it indicates that replacing the character will not have a great impact on the effectiveness of the feedback text.
[0099] S380, determine the score of each target character. If the score of any target character does not meet the character score condition corresponding to the target character, then determine the feedback text as low confidence feedback text.
[0100] It is understandable that after obtaining the score of each character in the feedback text and the original text through the above step S330, the score of the target character can be obtained after the target character is determined.
[0101] S390, if the score of each target character satisfies the character score condition corresponding to each target character, then the feedback text is determined to be a high-confidence feedback text.
[0102] Based on the above, in optional implementations, the method for determining whether the feedback text is of low or high confidence level can be found in [reference needed]. Figure 8 , Figure 8 A flowchart illustrating the process of determining the confidence level of feedback text, as provided in this embodiment of the invention:
[0103] b1 determines whether the score of the inserted target character is greater than the first threshold. If so, b2 is executed; otherwise, b5 is executed.
[0104] b2. Determine whether the score of the target character to be deleted is less than the second threshold. If yes, proceed to step b3; otherwise, proceed to step b5.
[0105] b3. Determine whether the score of the target character after replacement is greater than the score of the target character before replacement, or whether the score of the target character after replacement is less than the score of the target character before replacement, and the difference between the scores of the target character after replacement and the target character before replacement is less than the third threshold. If so, execute b4; otherwise, execute b5.
[0106] b4 indicates that the feedback text is a high-confidence feedback text.
[0107] b5 indicates that the feedback text is a low-confidence feedback text.
[0108] It should be noted that, Figure 8 The steps b1 to b3 shown do not have a specific execution order. It can be understood that steps b1 to b3 can be executed in the following order: Figure 8 The operations shown can be executed in the order indicated, or simultaneously. Furthermore, if the editing operation does not involve any of the insert, replace, or delete operations, then these operations can be skipped. Figure 8 The execution conditions for each corresponding operation.
[0109] By implementing the above methods, when the feedback text is valid, the confidence level of the feedback text can be determined, and then the high confidence level can be updated to the manuscript in a timely manner, thereby reducing the error rate of the manuscript.
[0110] The text recognition method provided in this application can be executed in a hardware device or as a software module. When the text recognition method is implemented as a software module, this application also provides a text recognition method apparatus. Please refer to [link to relevant documentation]. Figure 9 , Figure 9 This is a functional block diagram of a text recognition device provided in an embodiment of this application. The text recognition device 400 may include:
[0111] The acquisition module 410 is used to acquire feedback text and original text; wherein, the original text is the text obtained by speech recognition of the target audio; and the feedback text is the text based on the original text after error correction.
[0112] The recognition module 420 is used to: determine that the feedback text is invalid if the feedback text does not meet the preset text conditions; the text conditions are the conditions that characterize the validity of the text form of the feedback text; if the feedback text meets the text conditions, determine the scores corresponding to the feedback text and the original text respectively; the scores characterize the alignment results of the feedback text, the original text and the real text in the target audio respectively; if the scores corresponding to the feedback text and the original text respectively meet the preset text scoring conditions, determine that the feedback text is valid; if the scores corresponding to the feedback text and the original text respectively do not meet the text scoring conditions, determine that the feedback text is invalid.
[0113] It is understandable that the acquisition module 410 and the recognition module 420 can execute collaboratively. Figure 3 Each step in the process is used to achieve the corresponding technical effect.
[0114] In an optional implementation, the text conditions include: character format conditions, text length conditions, and edit distance conditions. The recognition module 420 is specifically used to: determine that the feedback text does not meet any one of the character format conditions, text length conditions, and edit distance conditions if the feedback text does not meet any one of the conditions; and determine that the feedback text meets the text conditions if the feedback text meets the character format conditions, text length conditions, and edit distance conditions in sequence.
[0115] In an optional implementation, the recognition module 420 is specifically used to: if the format of the characters in the feedback text is a preset character format, then the feedback text satisfies the character format condition; if the format of the characters in the feedback text is not a preset character format, then the feedback text satisfies the character format condition; wherein, the preset character format is any one or a combination of the following: Chinese characters; English characters; numbers; punctuation marks; if the text length of the feedback text is within a preset text length range, then the feedback text satisfies the text length condition; if it is not within the text length range, then the feedback text does not satisfy the text length condition; if the editing distance between the feedback text and the original text is within a preset distance range, then the feedback text satisfies the editing distance condition; if it is not within the distance range, then the feedback text does not satisfy the editing distance condition.
[0116] In an optional implementation, the text scoring conditions include: the score of the feedback text is greater than the score of the original text, or the score of the original text is greater than the score of the feedback text, and the difference between the score of the original text and the score of the feedback text is within a preset range.
[0117] In an optional implementation, the identification module 420 is further configured to: when the feedback text is determined to be valid feedback text, determine the target character of the edited operation in the feedback text and the original text based on the edit sequence corresponding to the feedback text and the original text; determine the character scoring condition corresponding to the target character based on the edit operation corresponding to the target character; determine the score of each target character; if the score of any target character does not meet the character scoring condition corresponding to the target character, then the feedback text is determined to be low-confidence feedback text; if the score of each target character meets the character scoring condition corresponding to each target character, then the feedback text is determined to be high-confidence feedback text.
[0118] In an optional implementation, the recognition module 420 is further specifically configured to: input the original text and the feedback text into a pre-trained acoustic model to obtain the decoding maps corresponding to the original text and the feedback text respectively; decode the decoding maps corresponding to the original text and the feedback text respectively with the audio features corresponding to the target audio to obtain the posterior probability and duration of each phoneme contained in the original text and the feedback text respectively; and obtain the score of each character contained in the original text and the feedback text respectively, as well as the score of the original text and the feedback text respectively, based on the posterior probability and duration of each phoneme contained in the original text and the feedback text respectively.
[0119] In an optional implementation, the editing operation includes any one or a combination of the following: insertion operation; deletion operation; replacement operation; if the editing operation of the target character is an insertion operation, the character score condition corresponding to the target character is: the score of the target character is greater than a first threshold; if the editing operation of the target character is a deletion operation, the character score condition corresponding to the target character is: the score of the target character is less than a second threshold; if the editing operation of the target character is a replacement operation, the character score condition corresponding to the target character is: the score of the replaced target character is greater than the score of the original target character, or the score of the replaced target character is less than the score of the original target character, and the difference between the score of the replaced target character and the score of the original target character is less than a third threshold; wherein, the replaced target character is located in the feedback text, and the original target character is located in the original text.
[0120] This application also provides a readable storage medium storing a computer program thereon, which, when executed by a processor, implements the text recognition method as described in any of the foregoing embodiments. The computer-readable storage medium may be, but is not limited to, various media capable of storing program code, such as a USB flash drive, portable hard drive, ROM, RAM, PROM, EPROM, EEPROM, magnetic disk, or optical disk.
[0121] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A text recognition method, characterized by, The method comprises: obtaining a feedback text and an original text; wherein the original text is a text obtained by performing speech recognition on a target audio; and the feedback text is a text obtained by performing error correction processing on the original text; if the feedback text does not satisfy a preset text condition, determining that the feedback text is an invalid feedback text; wherein the text condition is a condition representing that the text form of the feedback text is valid; if the feedback text satisfies the text condition, determining scores corresponding to the feedback text and the original text respectively; wherein the scores represent alignment results of the feedback text and the original text with a real text in the target audio respectively; if the scores corresponding to the feedback text and the original text respectively satisfy a preset text score condition, determining that the feedback text is a valid feedback text; wherein the text score condition comprises: the score of the feedback text is greater than the score of the original text, or the score of the original text is greater than the score of the feedback text, and the difference between the score of the original text and the score of the feedback text is within a preset difference range; if the scores corresponding to the feedback text and the original text respectively do not satisfy the text score condition, determining that the feedback text is an invalid feedback text.
2. The text recognition method of claim 1, wherein, The text condition comprises: a character format condition, a text length condition and an edit distance condition, and the method further comprises: if the feedback text does not satisfy any one of the character format condition, the text length condition and the edit distance condition, determining that the feedback text does not satisfy the text condition; if the feedback text satisfies the character format condition, the text length condition and the edit distance condition in turn, determining that the feedback text satisfies the text condition.
3. The text recognition method of claim 2, wherein, The method further comprises: if the format of characters in the feedback text is a preset character format, the feedback text satisfies the character format condition; if the format of characters in the feedback text is not the preset character format, the feedback text does not satisfy the character format condition; wherein the preset character format is any one or a combination thereof: Chinese characters, English characters, numbers and punctuation marks; if the text length of the feedback text is within a preset text length range, the feedback text satisfies the text length condition; if not, the feedback text does not satisfy the text length condition; if the edit distance between the feedback text and the original text is within a preset distance range, the feedback text satisfies the edit distance condition; if not, the feedback text does not satisfy the edit distance condition.
4. The text recognition method of claim 1, wherein, The method further comprises: when determining that the feedback text is a valid feedback text, determining target characters in the feedback text and the original text that are edited based on an edit sequence corresponding to the feedback text and the original text; determining a character score condition corresponding to the target characters based on an edit operation corresponding to the target characters. determining a score of each of the target characters, and determining that the feedback text is a low-confidence feedback text if the score of any of the target characters does not satisfy a character score condition corresponding to the target character; determining that the feedback text is a high-confidence feedback text if the score of each of the target characters satisfies the character score condition corresponding to the target character.
5. The text recognition method of claim 4, wherein, if the feedback text satisfies the text condition, determining a score corresponding to each of the feedback text and the original text, comprising: inputting the original text and the feedback text into a pre-trained acoustic model respectively to obtain a decoding graph corresponding to each of the original text and the feedback text; decoding the decoding graph corresponding to each of the original text and the feedback text respectively and the audio feature corresponding to the target audio to obtain a phoneme posterior probability and a duration of each character contained in the original text and the feedback text respectively; obtaining a score of each character contained in the original text and the feedback text respectively and a score of the original text and the feedback text respectively according to the phoneme posterior probability and the duration of each character contained in the original text and the feedback text respectively.
6. The text recognition method of claim 5, wherein, The editing operation includes any one of the following and combinations thereof: an insertion operation; a deletion operation; a replacement operation; if the editing operation of the target character is the insertion operation, the character score condition corresponding to the target character is that the score of the target character is greater than a first threshold value; if the editing operation of the target character is the deletion operation, the character score condition corresponding to the target character is that the score of the target character is less than a second threshold value; if the editing operation of the target character is the replacement operation, the character score condition corresponding to the target character is: the score of the target character after replacement is greater than the score of the target character before replacement, or the score of the target character after replacement is less than the score of the target character before replacement, and the difference between the score of the target character after replacement and the score of the target character before replacement is less than a third threshold value; wherein the target character after replacement is located in the feedback text, and the target character before replacement is located in the original text.
7. A text recognition apparatus characterized by comprising: comprising: an acquisition module and an identification module; the acquisition module is configured to acquire a feedback text and an original text; wherein the original text is a text obtained by performing speech recognition on a target audio; and the feedback text is a text obtained by performing error correction processing based on the original text; the identification module is configured to: if the feedback text does not satisfy a preset text condition, determining that the feedback text is an invalid feedback text; the text condition is a condition representing that the text form of the feedback text is valid; if the feedback text satisfies the text condition, determining a score corresponding to each of the feedback text and the original text; the score represents an alignment result of the feedback text, the original text and the real text in the target audio respectively. If the scores corresponding to the feedback text and the original text respectively satisfy a preset text score condition, the feedback text is determined as a valid feedback text; if the scores corresponding to the feedback text and the original text respectively do not satisfy the text score condition, the feedback text is determined as an invalid feedback text; the text score condition comprises: the score of the feedback text is greater than the score of the original text, or the score of the original text is greater than the score of the feedback text, and the score difference between the score of the original text and the score of the feedback text is within a preset score difference range.
8. An electronic device, comprising: A computer program product comprising a processor and a memory storing a computer program executable by the processor, the processor being executable to implement the method of any one of claims 1 to 6.
9. A readable storage medium, having stored thereon a computer program, characterized in that, The computer program product is executable by the processor to implement the method of any one of claims 1 to 6.
Citation Information
Patent Citations
Error correction method and device of voice recognition text
CN106847288A