A text correction method and device, computer equipment and storage medium

By performing dual-dimensional error correction on speech recognition text using both text and pinyin, and by utilizing preset sets and confusion sets for error correction processing, the problem of low error correction accuracy in existing technologies has been solved, achieving a higher error correction accuracy.

CN115600594BActive Publication Date: 2026-03-17CHINA CONSTRUCTION BANK +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-11
Publication Date
2026-03-17

AI Technical Summary

Technical Problem

In existing technologies, the conversion between pinyin and text is not one-to-one, resulting in many text recognition errors during speech recognition. Existing text-based error correction methods have low accuracy and cannot effectively improve the accuracy of automatic correction.

Method used

By performing dual error correction on the text to be checked in terms of both text and pinyin dimensions, the pre-error text is determined by using a preset set of segmented text and a set of confused text, and the pre-error pinyin text is determined by combining a preset set of pinyin and a set of confused pinyin. Finally, the corrected text is determined by inverse pinyin conversion and text replacement.

Benefits of technology

It improves the accuracy of text error detection and automatic correction in speech recognition application scenarios, and significantly enhances the error correction effect through dual-dimensional processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115600594B_ABST
    Figure CN115600594B_ABST
Patent Text Reader

Abstract

The present specification relates to the technical field of computer, and particularly relates to a text correction method and device, computer equipment and storage medium. The method comprises: performing word segmentation processing and pinyin conversion processing on a text to be checked, to obtain a plurality of to-be-checked character texts and a plurality of to-be-checked pinyin texts; determining at least one pre-error character text and at least one pre-error pinyin text from the plurality of to-be-checked character texts and the plurality of to-be-checked pinyin texts by using a preset word segmentation character set, a confused character set, a preset pinyin set and a confused pinyin set; performing pinyin reverse conversion processing on each pre-error pinyin text to obtain a pre-converted error character text; and processing the pre-error character text and the pre-converted error character text in the text to be checked by using the confused character set and a preset target character set, to determine a corrected text. By using the embodiments of the present specification, the accuracy of text checking and automatic correction in the application scenario of speech recognition is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of computer technology, and in particular to a text correction method, apparatus, computer device, and storage medium. Background Technology

[0002] Currently, the application scenarios requiring speech recognition are increasing, necessitating the conversion of collected speech into text. However, during the speech-to-text process, the non-one-to-one correspondence between pinyin and characters often leads to numerous misidentified characters in the converted text. To identify these misidentified characters from the text, current methods focus on the character dimension for detection and automatic correction. However, since speech recognition involves conversion between two dimensions (pinyin and text), the accuracy of current methods based solely on the character dimension for misidentification detection is low, consequently resulting in low accuracy for automatic correction.

[0003] Improving the accuracy of text error detection and automatic correction in speech recognition applications is a problem that urgently needs to be solved in existing technologies. Summary of the Invention

[0004] To address the problems in the prior art, embodiments of this specification provide a text correction method, apparatus, computer device, and storage medium that perform text-level and pinyin-level error correction on the text to be corrected, thereby determining the corrected text and improving the accuracy of text error detection and automatic correction in speech recognition application scenarios.

[0005] To solve the above-mentioned technical problems, the specific technical solution in this specification is as follows:

[0006] On the one hand, embodiments of this specification provide a text correction method, including,

[0007] The text to be checked is processed by word segmentation and pinyin conversion to obtain multiple text texts to be checked and multiple pinyin texts to be checked.

[0008] Using a preset set of segmented text and a set of obfuscated text, at least one pre-error text is determined from the plurality of texts to be checked;

[0009] Using a preset set of pinyin and a set of confused pinyin, at least one pre-erroneous pinyin text is determined from the plurality of pinyin texts to be checked;

[0010] For each of the pre-error pinyin texts, perform inverse pinyin conversion to obtain the pre-converted error text; and

[0011] Using the obfuscated text set and the preset target text set, the pre-error text and the pre-conversion error text in the text to be checked are processed to determine the corrected text.

[0012] The pinyin conversion and the pinyin inverse conversion are mutually inverse conversions.

[0013] Furthermore, the text to be checked is processed by word segmentation and pinyin conversion, resulting in multiple text texts to be checked and multiple pinyin texts to be checked, which further include...

[0014] The text to be checked is segmented into words to obtain the plurality of texts to be checked; and

[0015] The multiple texts to be checked are processed by pinyin conversion to obtain the multiple pinyin texts to be checked.

[0016] The pinyin conversion process includes determining the plurality of pinyin texts to be checked based on the plurality of texts to be checked and the speech information corresponding to the texts to be checked for errors.

[0017] Furthermore, the text to be checked is processed by word segmentation and pinyin conversion, resulting in multiple text texts to be checked and multiple pinyin texts to be checked, which further include...

[0018] The text to be checked is processed by word segmentation to obtain the multiple texts to be checked.

[0019] The text to be checked for errors is converted to pinyin to obtain the pinyin text to be checked; and

[0020] The text to be checked for errors is processed by word segmentation to obtain multiple texts of the same pinyin.

[0021] The pinyin conversion process includes determining the pinyin text to be checked based on the text to be checked and the corresponding speech information.

[0022] Furthermore, the preset word segmentation text set includes a preset first text set and a preset second text set. The step of using the preset word segmentation text set and the obfuscated text set to determine at least one pre-error text from the plurality of texts to be checked further includes,

[0023] According to the order of the characters in the text to be checked, the multiple texts to be checked are sorted to obtain a sequence of texts to be checked;

[0024] Each pair of adjacent texts to be checked in the text sequence to be checked is taken as a sub-set of texts to be checked, resulting in multiple sub-sets of texts to be checked.

[0025] Using the preset first text set, at least one erroneous sub-text set is determined from the plurality of sub-text sets to be checked; and

[0026] Using the obfuscated text set and the preset second text set, the at least one pre-error text is determined from the at least one set of error sub-texts.

[0027] The preset first character set includes a two-character standard character set, and the preset second character set includes a three-character standard character set.

[0028] Furthermore, the method of determining at least one erroneous sub-text set from the plurality of sub-text sets to be checked using the preset first text set further includes,

[0029] Using the preset first text set, calculate two first probability values ​​corresponding to each of the plurality of sub-text sets to be checked; and

[0030] If it is determined that there is a target first probability value less than or equal to a preset first threshold among the two first probability values, the sub-text set to be checked corresponding to the target first probability value is determined as the erroneous sub-text set, and the at least one erroneous sub-text set is obtained.

[0031] Furthermore, the determination of the at least one pre-error text from the at least one erroneous sub-text set, in the context of the obfuscated text set and the preset second text set, further includes:

[0032] For each of the at least one set of erroneous subtext texts, a first erroneous text is determined;

[0033] Using the sub-obfuscated text set corresponding to the first erroneous text in the obfuscated text set, the first erroneous text in the erroneous sub-text set is replaced to obtain multiple first replacement texts corresponding to the erroneous sub-text set;

[0034] Using the preset second text set, calculate multiple second probability values ​​corresponding to the multiple first replacement texts; and

[0035] If it is determined that there is a target second probability value greater than or equal to a preset second threshold among the plurality of second probability values, a pre-error text is determined from the set of error sub-texts to obtain at least one pre-error text.

[0036] Furthermore, the preset pinyin set includes a preset first pinyin set and a preset second pinyin set. The step of using the preset pinyin set and the confused pinyin set to determine at least one pre-error pinyin text from the plurality of pinyin texts to be checked further includes...

[0037] According to the order of the characters in the text to be checked, the multiple texts to be checked are sorted to obtain a sequence of texts to be checked.

[0038] Each pair of adjacent pinyin texts in the pinyin text sequence to be checked is taken as a sub-pinyin text set to be checked, resulting in multiple sub-pinyin text sets to be checked;

[0039] Using the obfuscated pinyin set and the preset first pinyin set, at least one erroneous sub-pinyin text set is determined from the plurality of sub-pinyin text sets to be checked; and

[0040] Using the preset second pinyin set, determine the at least one pre-error pinyin text from the at least one set of error sub-pinyin texts.

[0041] The preset first pinyin set includes a set of standard characters with two-character pinyin, and the preset second pinyin set includes a set of standard characters with three-character pinyin.

[0042] Furthermore, the method of determining at least one erroneous sub-pinyin text set from the plurality of sub-pinyin text sets to be checked using the preset first pinyin set further includes:

[0043] Using the preset first pinyin set, calculate two third probability values ​​corresponding to each of the plurality of sub-pinyin text sets to be checked; and

[0044] If it is determined that there is a target third probability value less than or equal to a preset third threshold among the two third probability values, the sub-to-be-checked pinyin text set corresponding to the target third probability value is determined as the erroneous sub-pinyin text set, and the at least one erroneous sub-pinyin text set is obtained.

[0045] Furthermore, the method of determining the at least one pre-erroneous pinyin text from the at least one set of erroneous sub-pinyin texts using the confused pinyin set and the preset second pinyin set further includes:

[0046] For each of the at least one set of erroneous sub-pinyin texts, a second erroneous text is determined;

[0047] Using the sub-confused pinyin set corresponding to the second erroneous text in the confused pinyin set, the second erroneous text in the erroneous sub-pinyin text set is replaced to obtain multiple second replacement texts corresponding to the erroneous sub-pinyin text set;

[0048] Using the second pinyin set, calculate multiple fourth probability values ​​corresponding to the multiple second replacement texts; and

[0049] If it is determined that there is a target fourth probability value greater than or equal to a preset fourth threshold among the plurality of fourth probability values, a pre-error pinyin text is determined from the set of erroneous sub-pinyin texts to obtain at least one pre-error pinyin text.

[0050] Furthermore, the process of determining the preset pinyin set further includes,

[0051] For each preset segmented text in the preset segmented text set, a pinyin conversion process is performed to obtain multiple preset converted pinyin texts;

[0052] Remove duplicate preset converted pinyin texts from the plurality of preset converted pinyin texts to obtain a target preset converted pinyin text set; and

[0053] Using a digital processing model, the target preset conversion pinyin text set is processed to obtain the preset pinyin set.

[0054] Furthermore, the method utilizes the obfuscated text set and the preset target text set to process the pre-error text and the pre-conversion error text in the text to be checked for errors, and determines the corrected text, which further includes...

[0055] Based on the at least one pre-error text and the at least one pre-converted error text, determine the set of error texts to be checked;

[0056] Using multiple sub-obfuscated text sets in the obfuscated text set that correspond to the set of erroneous texts to be checked, typos are replaced in the text to be checked to obtain multiple candidate sentences;

[0057] Multiple candidate statements are processed using a preset target text set to obtain multiple fifth probability values;

[0058] From the plurality of fifth probability values, the largest fifth probability value is determined as the target fifth probability value;

[0059] The candidate statement corresponding to the fifth probability value of the target is determined as the corrected text.

[0060] The preset target text set includes a preset standard text set corresponding to multiple characters, wherein the multiple characters are more than three characters.

[0061] Furthermore, after determining the largest fifth probability value from the plurality of fifth probability values ​​as the target fifth probability value, the process further includes:

[0062] The target error text set is determined based on the candidate statements corresponding to the fifth probability value of the target and the error text to be detected.

[0063] On the other hand, embodiments of this specification also provide a text correction device, including,

[0064] The first processing unit is used to perform word segmentation and pinyin conversion on the text to be checked, resulting in multiple text texts to be checked and multiple pinyin texts to be checked.

[0065] The first determining unit is used to determine at least one pre-error text from the plurality of texts to be checked by utilizing a preset word segmentation text set and a confused text set;

[0066] The second determining unit is used to determine at least one pre-error pinyin text from the plurality of pinyin texts to be checked by utilizing a preset pinyin set and a confused pinyin set;

[0067] The second processing unit is configured to perform reverse pinyin conversion processing on each of the pre-error pinyin texts to obtain pre-converted error texts; and

[0068] The third processing unit is used to process the pre-error text and the pre-conversion error text in the text to be checked using the obfuscated text set and the preset target text set, and to determine the corrected text.

[0069] The pinyin conversion and the pinyin inverse conversion are mutually inverse conversions.

[0070] Furthermore, the text correction device further includes,

[0071] The third determining unit is used to determine the target error text set based on the target candidate statement corresponding to the target fifth probability value and the error text to be detected.

[0072] On the other hand, embodiments of this specification also provide a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-described method.

[0073] On the other hand, embodiments of this specification also provide a computer-readable storage medium having computer instructions stored thereon, which, when executed by a processor, implement the above-described method.

[0074] On the other hand, embodiments of this specification also provide a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of the above-described method.

[0075] Using the embodiments of this specification, at least one pre-error text is determined by utilizing a preset set of segmented texts and a set of confused texts for the text to be detected, and at least one pre-error pinyin text is determined by utilizing a preset set of pinyins and a set of confused pinyins for the pinyin dimensions of the text to be detected. Then, using the set of confused texts and a preset target text set, the pre-error texts in the text to be detected and the pre-conversion error texts corresponding to each pre-error pinyin text are processed to determine the corrected text. Thus, correction is performed based on both text and pinyin dimensions, improving the accuracy of text error detection and automatic correction in speech recognition application scenarios. Attached Figure Description

[0076] To more clearly illustrate the technical solutions in the embodiments or prior art of this specification, the drawings used in the description of the embodiments or prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0077] Figure 1 The diagram shown is a schematic representation of an implementation system for a text correction method according to an embodiment of this specification.

[0078] Figure 2 The diagram shown is a flowchart of a text correction method according to an embodiment of this specification;

[0079] Figure 3 The diagram shown is a flowchart of another embodiment of a text correction method in this specification;

[0080] Figure 4 The diagram shown is a flowchart of another embodiment of a text correction method in this specification;

[0081] Figure 5 The diagram shown is a flowchart of another embodiment of a text correction method in this specification;

[0082] Figure 6 The diagram shown is a flowchart of another embodiment of a text correction method in this specification;

[0083] Figure 7The diagram shown is a flowchart of another embodiment of a text correction method in this specification;

[0084] Figure 8 The diagram shown is a structural schematic of a text correction device according to an embodiment of this specification.

[0085] Figure 9 The diagram shown is a structural schematic of a text correction device according to another embodiment of this specification.

[0086] Figure 10 This is a schematic diagram of the structure of a computer device according to an embodiment of this specification.

[0087] [Explanation of Labels in the Attached Image]

[0088] 101. User terminal;

[0089] 102. Server;

[0090] 810. First processing unit;

[0091] 820. First Determined Unit;

[0092] 830. Second Determined Unit;

[0093] 840. Second processing unit;

[0094] 850. Third processing unit;

[0095] 954. First Sub-processing Unit;

[0096] 955. First sub-unit determination;

[0097] 950. The third unit to be determined;

[0098] 1002. Computer equipment;

[0099] 1004. Processing equipment;

[0100] 1006. Storage resources;

[0101] 1008. Drive mechanism;

[0102] 1010. Input / Output Module;

[0103] 1012. Input devices;

[0104] 1014. Output devices;

[0105] 1016. Presentation device;

[0106] 1018. Graphical User Interface;

[0107] 1020. Network interface;

[0108] 1022. Communication link;

[0109] 1024. Communication bus. Detailed Implementation

[0110] The technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this specification.

[0111] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, apparatus, product, or device that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices.

[0112] It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.

[0113] Figure 1The diagram illustrates an implementation system for a text correction method according to an embodiment of this specification. The system may include a user terminal 101 and a server 102, which communicate via a network. This network may include a Local Area Network (LAN), a Wide Area Network (WAN), the Internet, or a combination thereof, and is connected to a website, user equipment (e.g., a computing device), and a backend system. The user can send voice information or text information to be checked for errors to the server 102 using the user terminal 101. Upon receiving the voice information, the server 102 performs speech-to-text processing to determine the text to be checked for errors. Upon receiving the text information to be checked for errors, the server uses this text information as the text to be checked for errors. After determining the text to be checked for errors, the server performs error correction and modification on the text in both textual and phonetic dimensions to obtain the corrected text. Then, the corrected text is further processed, or it is sent back to the user terminal 101. Alternatively, server 102 may be a node of a cloud computing system (not shown in the figure), or each server 102 may be a separate cloud computing system comprising multiple computers interconnected by a network and operating as a distributed processing system.

[0114] In an optional embodiment, the user terminal 101 may include electronic devices, including but not limited to smartphones, data acquisition devices, desktop computers, tablets, laptops, smart speakers, digital assistants, augmented reality (AR) / virtual reality (VR) devices, smart wearable devices, and other similar electronic devices. Optionally, the operating system running on the electronic device may include, but is not limited to, Android, iOS, Linux, Windows, etc.

[0115] In addition, it should be noted that, Figure 1 The example shown is merely one application environment provided in this manual. In actual applications, it may include multiple user terminals 101, and this manual does not impose any restrictions.

[0116] Figure 2 The diagram shows a flowchart of a text correction method according to an embodiment of this specification. The text correction process is depicted in this figure, but may include more or fewer steps based on conventional or non-creative labor. The order of steps listed in the embodiment is merely one possible execution order among many and does not represent the only possible execution order. In actual system or device products, the method can be executed sequentially or in parallel according to the embodiment or the accompanying drawings. Specifically, as shown... Figure 2 As shown, the method may include:

[0117] S210, perform word segmentation and pinyin conversion on the text to be checked to obtain multiple text texts to be checked and multiple pinyin texts to be checked;

[0118] S220, using a preset word segmentation text set and a confused text set, determine at least one pre-error text from the plurality of texts to be checked;

[0119] S230, using a preset set of pinyin and a set of confused pinyin, determine at least one pre-error pinyin text from the plurality of pinyin texts to be checked;

[0120] S240, Perform reverse pinyin conversion on each pre-error pinyin text to obtain the pre-converted error text;

[0121] S250, using a confused text set and a preset target text set, processes the pre-error text and pre-conversion error text in the text to be checked in order to determine the corrected text.

[0122] Using the embodiments of this specification, at least one pre-error text is determined by utilizing a preset set of segmented texts and a set of confused texts for the text to be detected, and at least one pre-error pinyin text is determined by utilizing a preset set of pinyins and a set of confused pinyins for the pinyin dimensions of the text to be detected. Then, using the set of confused texts and a preset target text set, the pre-error texts in the text to be detected and the pre-conversion error texts corresponding to each pre-error pinyin text are processed to determine the corrected text. Thus, correction is performed based on both text and pinyin dimensions, improving the accuracy of text error detection and automatic correction in speech recognition application scenarios.

[0123] According to embodiments of this specification, word segmentation processing can be, for example, two-character word segmentation. That is, each of the multiple texts to be checked consists of two Chinese characters. Pinyin conversion processing involves converting the text into corresponding pinyin. It should be noted that the pinyin includes syllables and tones. Furthermore, when each text to be checked consists of two Chinese characters, each of the multiple pinyin texts to be checked is a pinyin text corresponding to the two Chinese characters.

[0124] For example, when the text to be checked is “patent application document”, after performing two-character word segmentation on the text to be checked, the resulting text includes “patent”, “benefit”, “application”, “application document” and “document”.

[0125] The preset word segmentation text set includes a multi-character dictionary, which contains preset standard texts of words with two or three characters. For example, the two-character dictionary includes a preset standard text set corresponding to words composed of two characters, and this standard text set includes multiple two-character words without spelling mistakes. Similarly, the three-character dictionary includes a preset standard text set corresponding to words composed of three characters.

[0126] Similarly, the preset pinyin set includes a multi-pinyin dictionary, which contains preset standard pinyins corresponding to words with two or three characters. For example, the two-pinyin dictionary includes a preset standard pinyin set corresponding to the pinyins of words composed of two characters, and this standard pinyin set includes multiple pinyins of two characters without spelling mistakes. Similarly, the three-pinyin dictionary includes a preset standard pinyin set corresponding to the pinyins of words composed of three characters. For example, for "patent", the corresponding pinyin is " ", which is the pinyin of two characters without spelling mistakes.

[0127] The confused text set includes multiple texts and multiple homophonic characters corresponding to each text. These multiple homophonic characters are used to replace the spelling mistakes of the text. For example, when it is determined that "利" is a spelling mistake, the multiple homophonic characters corresponding to "利" are used to replace "利" respectively. Similarly, the confused pinyin set includes multiple pinyins and multiple similar pinyins corresponding to each pinyin. These multiple similar pinyins are used to replace the misspelled pinyins of the pinyin.

[0128] According to the number of characters in the text to be checked after word segmentation, the corresponding preset word segmentation text set is determined. For example, if the word segmentation is two-character word segmentation, the preset word segmentation text set is the two-character dictionary. Each text to be checked in multiple texts to be checked is matched with each preset text in the preset word segmentation text set. When it is determined that there is a target preset text that matches the text to be checked, it is determined that the text to be checked is a correct text; otherwise, it is determined that the text to be checked is a mismatched text. The sub-confused text sets corresponding to each text in the mismatched text are used to replace the corresponding texts in the mismatched text respectively, and multiple replaced text texts are obtained. A correct rate calculation model is used to determine the correct rate of each replaced text text. This correct rate calculation model can be, for example, a trained neural network model. When it is determined that there is a target correct rate that meets the correct rate threshold among multiple correct rates, it is determined that the mismatched text is a pre-error text. Similarly, using the preset pinyin set and the confused pinyin set, multiple texts to be checked for pinyin are processed respectively to determine at least one pre-error pinyin text.

[0129] After identifying the pre-error pinyin text, a pinyin inverse conversion process is performed on each pre-error pinyin text to determine the pre-converted error text corresponding to that pre-error pinyin text. The pinyin inverse conversion process converts pinyin into its corresponding characters. This pinyin inverse conversion is the inverse of the pinyin conversion mentioned above, ensuring that the pre-converted error text matches one of the multiple texts to be checked. Specifically, determining the pre-converted error text can also involve performing a pinyin inverse conversion process on each pre-error pinyin text to obtain multiple pre-converted texts. A consistency match is then performed on each of the multiple texts to be checked within each pre-converted text, and the target pre-converted text that matches the target text among the multiple texts to be checked is identified as the pre-converted error text.

[0130] The preset target text set is a multi-element text dictionary, where the number of elements is greater than the number of elements in the preset word segmentation text set mentioned above. In other words, the preset target text set includes a preset standard text set corresponding to words composed of multiple characters, where the number of characters in the multiple characters is greater than the number of characters in the text to be checked.

[0131] Using sub-obfuscated text sets corresponding to each character in the pre-error text and pre-conversion error text from the obfuscated text set, the corresponding characters in the text to be checked are replaced respectively, resulting in multiple replaced texts. Each replaced text is then matched with each preset target text in the preset target text set, and the target replaced text that matches the target preset target text in the preset target text set is determined as the corrected text.

[0132] According to another embodiment of this specification, performing word segmentation and pinyin conversion processing on the text to be checked to obtain multiple text texts to be checked and multiple pinyin texts to be checked includes: performing word segmentation processing on the text to be checked to obtain multiple text texts to be checked; and performing pinyin conversion processing on the multiple text texts to be checked to obtain multiple pinyin texts to be checked, wherein the pinyin conversion processing includes determining multiple pinyin texts to be checked based on the multiple text texts to be checked and the speech information corresponding to the text to be checked.

[0133] The pinyin conversion process includes a first pinyin conversion process and a comparison and correction process. The first pinyin conversion process can be implemented using a pinyin module that can be based on software (Python).

[0134] The text to be checked is first converted into pinyin to obtain multiple pinyin texts to be checked; then, using the phonetic information corresponding to the text to be checked as the standard pinyin, the multiple pinyin texts to be checked are corrected to obtain the text to be checked.

[0135] According to another embodiment of this specification, performing word segmentation and pinyin conversion processing on the text to be checked to obtain multiple text texts to be checked and multiple pinyin texts to be checked includes: performing word segmentation processing on the text to be checked to obtain multiple text texts to be checked; performing pinyin conversion processing on the text to be checked to obtain pinyin texts to be checked; and performing pinyin word segmentation processing on the pinyin texts to be checked to obtain multiple pinyin texts to be checked, wherein the pinyin conversion processing includes determining the pinyin texts to be checked based on the text to be checked and the corresponding speech information.

[0136] Text segmentation can be performed, for example, by segmenting two characters into words.

[0137] The pinyin conversion process includes a first pinyin conversion process and a comparison and correction process. The first pinyin conversion process is implemented using a pinyin module that can be based on software (Python). It should be noted that during the pinyin conversion process, multiple pinyin syllables corresponding to multiple Chinese characters are connected with identifiers to facilitate pinyin word segmentation.

[0138] After obtaining the pinyin text to be checked, pinyin word segmentation is performed to obtain multiple pinyin texts to be checked. This pinyin word segmentation process can be, for example, a binary search method, where a word is segmented upon encountering two identifiers. Then, using a comparison and correction process, the phonetic information corresponding to the text to be checked is used as the standard pinyin to correct these multiple pinyin texts to obtain multiple pinyin texts to be checked.

[0139] Figure 3 The diagram shows a flowchart of another embodiment of a text correction method according to this specification. The text correction process is described in this figure, but may include more or fewer steps based on conventional or non-inventive labor. The order of steps listed in the embodiment is merely one possible execution order among many and does not represent the only possible execution order. In actual system or device products, the method can be executed sequentially or in parallel according to the embodiment or the accompanying drawings. Specifically, as shown... Figure 3 As shown, the method may include:

[0140] S321, Sort multiple texts to be checked according to the order of the characters in the text to be checked to obtain a sequence of texts to be checked;

[0141] S322, take each pair of adjacent texts to be checked in the text sequence to be checked as a sub-text set to be checked, and obtain multiple sub-text sets to be checked.

[0142] S323, using a preset first text set, determine at least one erroneous sub-text set from multiple sub-text sets to be checked;

[0143] S324. Use the confused text set and the preset second text set to determine at least one pre-error text from at least one error sub-text text set.

[0144] According to another embodiment of this specification, for example, when the text to be error-checked is "patent application document", after performing two-word segmentation on the text to be error-checked, the sequence of text to be checked obtained includes "patent", "benefit body", "body application", "application text", and "document". Aggregate the sequence of text to be checked to obtain multiple sub-text-to-be-checked text sets, namely "patent, benefit body", "benefit body, body application", "body application, application text", and "application text, document".

[0145] The preset first text set includes a two-word standard text set, and the preset second text set includes a three-word standard text set.

[0146] For each sub-text-to-be-checked text set, match the two texts to be checked in the sub-text-to-be-checked text set with the first text in the first text set. When it is determined that one of the texts to be checked in each text to be checked does not match the first text set, determine the sub-text-to-be-checked text set as the target sub-text-to-be-checked text set. For example, "benefit body, body application" is an error sub-text text set. Perform three-word processing on this error sub-text text set to obtain the target error phrase "benefit body application". Use the sub-confused text sets corresponding to "benefit", "body", and "application" in the confused text set to replace the corresponding characters in the target error phrase "benefit body application" with misspelled characters, and obtain the corresponding three-word phrase. Match each three-word phrase with the second text in the preset second text set. When it is determined that the three-word phrase matches the target preset second text in the preset second text set, determine the replaced character corresponding to the three-word phrase as the pre-error text, that is, the pre-error text is "body".

[0147] According to another embodiment of this specification, using the preset first text set to determine at least one error sub-text text set from multiple sub-text-to-be-checked text sets includes: using the preset first text set to calculate two first probability values respectively corresponding to each sub-text-to-be-checked text set in the multiple sub-text-to-be-checked text sets; and when it is determined that there is a target first probability value less than or equal to the preset first threshold among the two first probability values, determine the sub-text-to-be-checked text set corresponding to the target first probability value as the error sub-text text set, and obtain at least one error sub-text text set.

[0148] The specific method for calculating the first probability value of each text to be checked in the set of sub-texts to be checked is the Bayesian conditional probability model in the error detection method of multi-element matching (n-gram). If it is greater than the preset first probability value, it indicates that the text to be checked corresponding to this first probability value is the correct text, otherwise it is an incorrect text.

[0149] For example, for "patent" and "benefiting the body" in "patent, benefiting the body", the calculated first probability values are 0.97 and 0.98 respectively. When the preset first probability value is 0.95, it is determined that this set of sub-texts to be checked is a set of correct sub-texts, otherwise it is a set of incorrect sub-texts. Finally, it is determined that "benefiting the body, applying" is a set of incorrect sub-texts.

[0150] Figure 4 The figure shows a flowchart of a text correction method according to another embodiment of the present specification. The text correction process is described in this figure, but based on routine or non-creative labor, it may include more or fewer operation steps. The order of steps listed in the embodiment is only one way among the execution orders of numerous steps and does not represent the only execution order. When the actual system or device product is executed, it can be executed in the order of the method shown in the embodiment or the figure, or executed in parallel. Specifically, as Figure 4 shown, the method may include:

[0151] S4241, for each set of incorrect sub-texts in at least one set of incorrect sub-texts, respectively determine the first incorrect text;

[0152] S4242, using the set of sub-confusing texts corresponding to the first incorrect text in the set of confusing texts, replace the first incorrect text in the set of incorrect sub-texts, and obtain multiple first replacement texts corresponding to the set of incorrect sub-texts;

[0153] S4243, using the preset second set of texts, calculate multiple second probability values corresponding to the multiple first replacement texts;

[0154] S4244, when it is determined that there is a target second probability value greater than or equal to the preset second threshold among the multiple second probability values, determine the pre-incorrect text from the set of incorrect sub-texts, and obtain at least one pre-incorrect text.

[0155] According to another embodiment of the present specification, determine the second character of the first text to be checked and the first character of the second text to be checked in the set of incorrect sub-texts as the first incorrect text. For the set of incorrect sub-texts "benefiting the body, applying", determine the first incorrect text as "body".

[0156] From the set of confused texts, determine the subset of confused texts corresponding to the first incorrect text. Use each confused text in this subset of confused texts to replace the first incorrect text, obtaining a modified set of incorrect sub-texts, and perform three-character processing on this modified set of incorrect sub-texts to obtain the first replacement text.

[0157] The specific method for calculating the second probability value of the first replacement text is the Bayesian conditional probability model in the error detection method of multi-matching (n-gram). A value greater than the preset second probability value indicates that the first replacement text corresponding to this second probability value is the correct text, otherwise it is an incorrect text.

[0158] When the first replacement text is "li application", if the determined second probability value corresponding to this first replacement text is 0.99 and the preset second probability value is 0.92, then "shen" is determined as the pre-incorrect text.

[0159] Figure 5 The figure shows a flowchart of a text correction method according to another embodiment of the present specification. The text correction process is described in this figure, but based on routine or non-creative labor, it may include more or fewer operation steps. The order of steps listed in the embodiment is only one way among the execution orders of numerous steps and does not represent the only execution order. When the actual system or device product is executed, it can be executed in the order of the method shown in the embodiment or the figure, or executed in parallel. Specifically, as Figure 5 shown, the method may include:

[0160] S531, Sort the multiple texts of pinyin to be checked according to the order of the characters in the text to be error-detected, obtaining a sequence of texts of pinyin to be checked;

[0161] S532, Take every two adjacent texts of pinyin to be checked in the sequence of texts of pinyin to be checked as a subset of texts of pinyin to be checked, obtaining multiple subsets of texts of pinyin to be checked;

[0162] S533, Use the set of confused pinyin and the preset first set of pinyin to determine at least one subset of incorrect texts of pinyin from multiple subsets of texts of pinyin to be checked;

[0163] S534, Use the preset second set of pinyin to determine at least one pre-incorrect text of pinyin from at least one subset of incorrect texts of pinyin.

[0164] According to another embodiment of the present specification, the preset set of pinyin includes a preset first set of pinyin and a preset second set of pinyin. The preset first set of pinyin includes a set of standard texts of two-character pinyin, and the preset second set of pinyin includes a set of standard texts of three-character pinyin.

[0165] After identifying multiple texts containing pinyin to be checked, for example, a process similar to... Figure 3 The operation shown is to identify at least one pre-erroneous pinyin text.

[0166] According to another embodiment of this specification, the process of determining the preset pinyin set includes: performing pinyin conversion processing on each preset text in the preset word segmentation text set to obtain multiple preset converted pinyin texts; removing duplicate preset converted pinyin texts from the multiple preset converted pinyin texts to obtain a target preset converted pinyin text set; and using a digital processing model to process the target preset converted pinyin text set to obtain the preset pinyin set.

[0167] For each preset text in the preset word segmentation text set, the Chinese text is converted to Pinyin using the HanLP (Chinese Language Library) platform, resulting in multiple preset converted Pinyin texts. Then, duplicate preset converted Pinyin texts are removed from these multiple preset converted Pinyin texts to obtain a target preset converted Pinyin text set. For each target preset converted Pinyin text in the target preset converted Pinyin text set, a preset Pinyin set is constructed using the one-hot encoding method to ensure that each syllable and tone corresponds to a unique code, making the code for each Pinyin distinct. This results in a preset Pinyin set used for Pinyin-level error detection. The preset Pinyin set includes a first preset Pinyin set and a second preset Pinyin set. The only difference between the first and second preset Pinyin sets is the number of Chinese characters corresponding to each preset Pinyin in the first preset Pinyin set.

[0168] According to another embodiment of this specification, determining at least one erroneous sub-pinyin text set from multiple sub-pinyin text sets to be checked using a preset first pinyin set includes: using the preset first pinyin set, calculating two third probability values ​​corresponding to each of the multiple sub-pinyin text sets to be checked; and if it is determined that there is a target third probability value less than or equal to a preset third threshold among the two third probability values, determining the sub-pinyin text set to be checked corresponding to the target third probability value as an erroneous sub-pinyin text set, thereby obtaining at least one erroneous sub-pinyin text set.

[0169] After identifying multiple pinyin texts to be checked, an operation similar to "using a preset first character set, calculating two first probability values ​​corresponding to each of the multiple sub-text sets to be checked; and if it is determined that there is a target first probability value less than or equal to a preset first threshold among the two first probability values, determining the sub-text set to be checked corresponding to the target first probability value as an erroneous sub-text set, thereby obtaining at least one erroneous sub-text set" can be performed to identify at least one erroneous sub-pinyin text set.

[0170] According to another embodiment of this specification, determining at least one pre-erroneous pinyin text from at least one set of erroneous sub-pinyin texts using a confused pinyin set and a preset second pinyin set includes: determining a second erroneous text for each set of erroneous sub-pinyin texts in the at least one set of erroneous sub-pinyin texts; replacing the second erroneous text in the set of erroneous sub-pinyin texts using a sub-confused pinyin set corresponding to the second erroneous text in the confused pinyin set to obtain a plurality of second replacement texts corresponding to the set of erroneous sub-pinyin texts; calculating a plurality of fourth probability values ​​corresponding to the plurality of second replacement texts using the second pinyin set; and determining the pre-erroneous pinyin text from the set of erroneous sub-pinyin texts when a target fourth probability value greater than or equal to a preset fourth threshold is determined among the plurality of fourth probability values ​​to obtain at least one pre-erroneous pinyin text.

[0171] After identifying at least one set of erroneous sub-phonetic texts, for example, a process similar to... Figure 4 The operation shown is to identify at least one pre-erroneous pinyin text.

[0172] Figure 6 The diagram shows a flowchart of another embodiment of a text correction method according to this specification. The text correction process is described in this figure, but may include more or fewer steps based on conventional or non-inventive labor. The order of steps listed in the embodiment is merely one possible execution order among many and does not represent the only possible execution order. In actual system or device products, the method can be executed sequentially or in parallel according to the embodiment or the accompanying drawings. Specifically, as shown... Figure 6 As shown, the method may include:

[0173] S651, determine the set of error texts to be checked based on at least one pre-error text and at least one pre-conversion error text;

[0174] S652, using multiple sub-obfuscated text sets in the obfuscated text set that correspond to the error text set to be checked, perform typo replacement on the error text to be checked to obtain multiple candidate sentences;

[0175] S653, using a preset target text set to process multiple candidate sentences to obtain multiple fifth probability values;

[0176] S654, From multiple fifth probability values, determine the largest fifth probability value as the target fifth probability value;

[0177] S655, determine the candidate statement corresponding to the fifth probability value of the target as the corrected text.

[0178] According to another embodiment of this specification, the preset target text set includes a preset standard text set corresponding to multiple characters, where the number of characters is more than three.

[0179] The union of the pre-error text and at least one pre-converted error text is taken as the set of error texts to be checked. For each character in the error text to be checked, a corresponding sub-obfuscated text set is determined from the obfuscated text set. Using this sub-obfuscated text set, the corresponding characters in the error text to be checked are replaced with typos, resulting in multiple candidate sentences. It is important to note that the typo replacement is performed first on one corresponding character in the error text to be checked, then on two corresponding characters, and finally on multiple corresponding characters. The number of multiple corresponding characters is determined by the number of characters in the error text set to be checked. For example, the number of multiple corresponding characters is the same as the number of characters in the error text set to be checked.

[0180] The specific method for calculating the fifth probability value corresponding to the candidate statement is the Bayesian conditional probability model in the multivariate matching error detection method (n-gram).

[0181] After determining multiple fifth probability values, the largest fifth probability value is selected as the target fifth probability value, and the candidate statement corresponding to the target fifth probability value is used as the correction text.

[0182] Figure 7 The figure shows a flowchart of a text correction method according to another embodiment of this specification. The figure illustrates a parking garage charging process, but based on conventional or non-creative labor, it may include more or fewer operational steps. The order of steps listed in the embodiment is merely one possible execution order among many and does not represent the only possible execution order. In actual system or device products, the methods shown in the embodiments or figures can be executed sequentially or in parallel. Specifically, as shown... Figure 7 As shown, the method may include:

[0183] S754, from multiple fifth probability values, determine the largest fifth probability value as the target fifth probability value;

[0184] S755, determine the candidate statement corresponding to the fifth probability value of the target as the corrected text.

[0185] S760, determine the target error text set based on the candidate statements and the text to be checked that correspond to the fifth probability value of the target.

[0186] According to another embodiment of this specification, the target error text set includes typos in the error-detecting text obtained from the error detection.

[0187] To ensure the traceability of the error detection process and the ease of tracing, a target error text set is determined by comparing the candidate statements corresponding to the target fifth probability value with the text to be detected. This target error text set is then associated with and stored in conjunction with at least one of the text to be detected and the candidate statements corresponding to the target fifth probability value.

[0188] To further improve the accuracy of the corrected text, the target set of error texts can be sent to the user terminal so that the user can confirm whether the typos included in the target set are genuine typos. Upon receiving non-genuine typos from the user terminal within the target set of error texts, a candidate statement corresponding to the target fifth probability value is sent to the user terminal, allowing the user to manually modify the candidate statement and send the modified text to the server. Upon receiving the modified text, the server uses it as the target text corresponding to the text to be corrected. Furthermore, the target text, the corrected text, and the text to be corrected are stored together for use in updating the text correction method.

[0189] Figure 8 The diagram shown is a structural schematic of a text correction device according to an embodiment of this specification. Figure 8 As shown, including,

[0190] The first processing unit 810 is used to perform word segmentation and pinyin conversion on the text to be checked, and to obtain multiple text texts to be checked and multiple pinyin texts to be checked.

[0191] The first determining unit 820 is used to determine at least one pre-error text from the plurality of texts to be checked by using a preset word segmentation text set and a confused text set;

[0192] The second determining unit 830 is used to determine at least one pre-error pinyin text from the plurality of pinyin texts to be checked by utilizing a preset pinyin set and a confused pinyin set;

[0193] The second processing unit 840 is used to perform reverse pinyin conversion processing on each pre-error pinyin text to obtain the pre-converted error text.

[0194] The third processing unit 850 is used to process the pre-error text and pre-conversion error text in the text to be checked by using the obfuscated text set and the preset target text set, so as to determine the corrected text.

[0195] Since the principle of the above-mentioned device in solving the problem is similar to that of the above-mentioned method, the implementation of the above-mentioned device can refer to the implementation of the above-mentioned method, and the repeated parts will not be described again.

[0196] Figure 9 The diagram shown is a structural schematic of a text correction device according to another embodiment of this specification. Figure 9 As shown, including,

[0197] The first sub-processing unit 954 is used to determine the largest fifth probability value as the target fifth probability value from multiple fifth probability values;

[0198] The first sub-determination unit 955 is used to determine the candidate statement corresponding to the fifth probability value of the target as the corrected text;

[0199] The third determining unit 960 is used to determine the target error text set based on the target candidate statement and the error text to be detected, which correspond to the target fifth probability value.

[0200] Since the principle of the above-mentioned device in solving the problem is similar to that of the above-mentioned method, the implementation of the above-mentioned device can refer to the implementation of the above-mentioned method, and the repeated parts will not be described again.

[0201] like Figure 10 The diagram illustrates the structure of a computer device according to an embodiment of this specification. The apparatus described in this specification can be the computer device in this embodiment, executing the methods described above. The computer device 1002 may include one or more processing devices 1004, such as one or more central processing units (CPUs), each of which can implement one or more hardware threads. The computer device 1002 may also include any storage resource 1006 for storing information of any kind, such as code, settings, data, etc. Without limitation, for example, the storage resource 1006 may include any one or more combinations of the following: any type of RAM, any type of ROM, flash memory, hard disk, optical disk, etc. More generally, any storage resource can use any technology to store information. Furthermore, any storage resource can provide volatile or non-volatile retention of information. Further, any storage resource may represent a fixed or removable component of the computer device 1002. In one case, when the processing device 1004 executes associated instructions stored in any storage resource or combination of storage resources, the computer device 1002 can perform any operation of the associated instructions. The computer device 1002 also includes one or more drive mechanisms 1008 for interacting with any storage resource, such as hard disk drive mechanism, optical disk drive mechanism, etc.

[0202] Computer device 1002 may further include an input / output module 1010 (I / O) for receiving various inputs (via input device 1012) and providing various outputs (via output device 1014). A specific output mechanism may include a presentation device 1016 and an associated graphical user interface (GUI) 1018. In other embodiments, the input / output module 1010 (I / O), input device 1012, and output device 1014 may be omitted, and the device may function solely as a computer device within a network. Computer device 1002 may also include one or more network interfaces 1020 for exchanging data with other devices via one or more communication links 1022. One or more communication buses 1024 couple the components described above together.

[0203] The communication link 1022 can be implemented in any way, such as via a local area network, a wide area network (e.g., the Internet), a point-to-point connection, or any combination thereof. The communication link 1022 may include any combination of hardwired links, wireless links, routers, gateway functions, name servers, etc., governed by any protocol or combination of protocols.

[0204] This specification also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method.

[0205] This specification also provides a computer program product, which includes a computer program that, when executed by a processor, implements the above-described method.

[0206] Those skilled in the art will understand that embodiments of this specification can be provided as methods, systems, or computer program products. Therefore, this specification may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this specification may take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0207] This specification is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this specification. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create a machine for implementing the flowchart illustrations and / or block diagrams. Figure 1One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0208] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0209] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0210] The above specific embodiments further illustrate the purpose, technical solutions, and beneficial effects of this specification. It should be understood that the above are merely specific embodiments of this specification and are not intended to limit the scope of protection of this specification. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this specification should be included within the scope of protection of this specification.

Claims

1. A method of text correction, characterized by, Comprising: performing word segmentation and pinyin conversion on the text to be checked to obtain a plurality of text to be checked and a plurality of pinyin to be checked; determining at least one pre-error text from the plurality of text to be checked using a preset word segmentation text set and a confused text set; determining at least one pre-error pinyin text from the plurality of pinyin to be checked using a preset pinyin set and a confused pinyin set; performing pinyin reverse conversion on each of the pre-error pinyin text to obtain a pre-conversion error text; and processing the pre-error text and the pre-conversion error text in the text to be checked using the confused text set and a preset target text set to determine a corrected text, wherein the pinyin conversion and the pinyin reverse conversion are mutually inverse conversions; the preset word segmentation text set includes a preset first text set and a preset second text set, and the determining at least one pre-error text from the plurality of text to be checked using a preset word segmentation text set and a confused text set includes: sorting the plurality of text to be checked in the order of the text in the text to be checked to obtain a text to be checked sequence; taking every two adjacent text to be checked in the text to be checked sequence as a sub-text to be checked set to obtain a plurality of sub-text to be checked sets; determining at least one error sub-text set from the plurality of sub-text to be checked sets using the preset first text set; and determining the at least one pre-error text from the at least one error sub-text set using the confused text set and the preset second text set, wherein the preset first text set includes a two-character standard text set, and the preset second text set includes a three-character standard text set.

2. The method of claim 1, wherein, The word segmentation and pinyin conversion on the text to be checked to obtain a plurality of text to be checked and a plurality of pinyin to be checked includes: performing word segmentation on the text to be checked to obtain the plurality of text to be checked; and performing pinyin conversion on the plurality of text to be checked to obtain the plurality of pinyin to be checked, wherein the pinyin conversion includes determining the plurality of pinyin to be checked according to the plurality of text to be checked and voice information corresponding to the text to be checked.

3. The method of claim 1, wherein, The word segmentation and pinyin conversion on the text to be checked to obtain a plurality of text to be checked and a plurality of pinyin to be checked includes: performing word segmentation on the text to be checked to obtain the plurality of text to be checked; performing pinyin conversion on the text to be checked to obtain text to be checked pinyin; and performing pinyin segmentation on the text to be checked pinyin to obtain the plurality of pinyin to be checked, wherein the pinyin conversion includes determining the text to be checked pinyin according to the text to be checked and voice information corresponding to the text to be checked.

4. The method of claim 1, wherein, The determining at least one error sub-text set from the plurality of sub-text to be checked sets using the preset first text set includes: The two first probability values respectively corresponding to each of the plurality of sub-to-be-checked character text sets are calculated by using the preset first character set. In a case where it is determined that there is a target first probability value less than or equal to a preset first threshold value in the two first probability values, a sub-to-be-checked character text set corresponding to the target first probability value is determined as an error sub-character text set, and the at least one error sub-character text set is obtained.

5. The method of claim 1, wherein, The at least one pre-error character text is determined from the at least one error sub-character text set by using the confusion character set and the preset second character set, including: For each error sub-character text set in the at least one error sub-character text set, a first error text is determined respectively; The first error text in the error sub-character text set is replaced by using a sub-confusion character set in the confusion character set corresponding to the first error text, and a plurality of first replacement texts corresponding to the error sub-character text set are obtained; The plurality of second probability values corresponding to the plurality of first replacement texts are calculated by using the preset second character set; and In a case where it is determined that there is a target second probability value greater than or equal to a preset second threshold value in the plurality of second probability values, a pre-error character text is determined from the error sub-character text set, and the at least one pre-error character text is obtained.

6. The method of claim 1, wherein, The preset pinyin set includes a preset first pinyin set and a preset second pinyin set, and the at least one pre-error pinyin text is determined from the plurality of to-be-checked pinyin texts by using the preset pinyin set and the confusion pinyin set, including: The plurality of to-be-checked pinyin texts are sorted according to the order of characters in the to-be-checked text, and a to-be-checked pinyin text sequence is obtained; Each of adjacent two to-be-checked pinyin texts in the to-be-checked pinyin text sequence is taken as a sub-to-be-checked pinyin text set, and a plurality of sub-to-be-checked pinyin text sets are obtained; The at least one error sub-pinyin text set is determined from the plurality of sub-to-be-checked pinyin text sets by using the confusion pinyin set and the preset first pinyin set; and The at least one pre-error pinyin text is determined from the at least one error sub-pinyin text set by using the preset second pinyin set, The preset first pinyin set includes a two-character pinyin standard character set, and the preset second pinyin set includes a three-character pinyin standard character set.

7. The method of claim 6, wherein, The at least one error sub-pinyin text set is determined from the plurality of sub-to-be-checked pinyin text sets by using the preset first pinyin set, including: The two third probability values respectively corresponding to each of the plurality of sub-to-be-checked pinyin text sets are calculated by using the preset first pinyin set; and In a case where it is determined that there is a target third probability value less than or equal to a preset third threshold value in the two third probability values, a sub-to-be-checked pinyin text set corresponding to the target third probability value is determined as an error sub-pinyin text set, and the at least one error sub-pinyin text set is obtained.

8. The method of claim 6, wherein, The determining the at least one pre-error pinyin text from the at least one error sub-pinyin text set by using the confusion pinyin set and the preset second pinyin set comprises: For each error sub-pinyin text set in the at least one error sub-pinyin text set, a second error text is determined respectively; For the second error text in the error sub-pinyin text set, a sub-confusion pinyin set corresponding to the second error text in the confusion pinyin set is used for replacement, to obtain a plurality of second replacement texts corresponding to the error sub-pinyin text set; A plurality of fourth probability values corresponding to the plurality of second replacement texts are calculated by using the second pinyin set; and In a case where there is a target fourth probability value greater than or equal to a preset fourth threshold value in the plurality of fourth probability values, a pre-error pinyin text is determined from the error sub-pinyin text set, to obtain the at least one pre-error pinyin text.

9. The method of claim 6, wherein, The determination process of the preset pinyin set comprises: For each preset segmented character text in the preset segmented character set, a pinyin conversion process is performed to obtain a plurality of preset converted pinyin texts; Duplicate preset converted pinyin texts in the plurality of preset converted pinyin texts are removed to obtain a target preset converted pinyin text set; and By using a digital processing model, the target preset converted pinyin text set is processed to obtain the preset pinyin set.

10. The method of claim 1, wherein, The processing of the pre-error character text and the pre-converted error character text in the to-be-checked text by using the confusion character set and the preset target character set to determine a corrected text comprises: According to the at least one pre-error character text and the at least one pre-converted error character text, a to-be-inspected error text set is determined; For the to-be-checked text, a plurality of sub-confusion character sets corresponding to the to-be-inspected error text set in the confusion character set are used for error replacement to obtain a plurality of candidate sentences; The plurality of candidate sentences are processed by using a preset target character set to obtain a plurality of fifth probability values; From the plurality of fifth probability values, a maximum fifth probability value is determined as a target fifth probability value; A candidate sentence corresponding to the target fifth probability value is determined as a corrected text, The preset target character set comprises a preset standard character set corresponding to a plurality of characters, and the plurality of characters are greater than three characters.

11. The method of claim 10, wherein, After the maximum fifth probability value is determined as the target fifth probability value from the plurality of fifth probability values, the following steps are further included: According to the candidate sentence corresponding to the target fifth probability value and the to-be-checked text, a target error text set is determined.

12. A text correction apparatus characterized by comprising: Comprise: The first processing unit is configured to perform segmentation processing and pinyin conversion processing on the to-be-checked text to obtain a plurality of to-be-inspected character texts and a plurality of to-be-inspected pinyin texts; The first determination unit is configured to determine at least one pre-error character text from the plurality of to-be-inspected character texts by using a preset segmented character set and a confusion character set; The second determining unit is configured to determine at least one pre-error pinyin text from the plurality of pinyin texts to be checked by using the preset pinyin set and the confused pinyin set. The second processing unit is configured to perform pinyin reverse conversion processing on each of the pre-error pinyin texts to obtain a pre-conversion error character text. The third processing unit is configured to perform processing on the pre-error character text and the pre-conversion error character text in the text to be checked by using the confused character set and a preset target character set to determine a corrected text. The pinyin conversion and the pinyin reverse conversion are reciprocal conversions. The preset segmented character set includes a preset first character set and a preset second character set, and the determining at least one pre-error character text from the plurality of character texts to be checked by using the preset segmented character set and the confused character set includes: sequencing the plurality of character texts to be checked according to the order of characters in the text to be checked to obtain a character text sequence to be checked; taking each pair of adjacent character texts to be checked in the character text sequence to be checked as a sub-character text set to be checked to obtain a plurality of sub-character text sets to be checked; determining at least one error sub-character text set from the plurality of sub-character text sets to be checked by using the preset first character set; and determining the at least one pre-error character text from the at least one error sub-character text set by using the confused character set and the preset second character set. The preset first character set includes a two-character standard character set, and the preset second character set includes a three-character standard character set. The third determining unit is configured to determine a target error text set according to a target candidate sentence corresponding to a target fifth probability value and the text to be checked.

13. The apparatus of claim 12, wherein, The processor executes the computer program to implement the method of any one of claims 1-11. The computer readable storage medium stores a computer program, and the computer program is executed by the processor to implement the method of any one of claims 1-11.

14. A computer device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The computer program / instruction is executed by the processor to implement the steps of the method according to any one of claims 1-11.

15. A computer-readable storage medium, characterized in that, The computer program / instruction is executed by the processor to implement the steps of the method according to any one of claims 1-11.

16. A computer program product comprising computer programs / instructions, characterized in that, ​

Citation Information

Patent Citations

  • Voice text error correction method

    CN111985234A

  • Method and device for correcting the error of the text, electronic equipment and storage medium

    CN113051896A