Text error correction method and apparatus, electronic device, and storage medium

By correcting error correction text in a large-scale language model and generating an edit distance matrix, the existing text error correction methods are solved, and efficient and intuitive text error correction effects are achieved.

WO2025130404A1PCT designated stage expired Publication Date: 2025-06-26SHENZHEN INTELLIFUSION TECHNOLOGIES CO LTD

Patent Information

Application Number
PCT/CN2024/130226
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-21
Filing Date
2024-11-06
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

The existing text error correction methods are inefficient and cannot intuitively reflect the problem of text errors and the error correction methods.

Method used

By inputting the text to be corrected to be corrected for error correction, an edit distance matrix is ​​generated, and the error correction result is determined based on the edit operands in the matrix.

Benefits of technology

It improves the efficiency of text error correction, and displays the problem of text errors and the error correction methods through intuitive error correction results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024130226_26062025_PF_FP_ABST
    Figure CN2024130226_26062025_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the present invention provide a text error correction method. The method comprises: acquiring text waiting for error correction; inputting the text waiting for error correction into a preset large language model for text error correction, so as to obtain error corrected text; on the basis of the text waiting for error correction and the error corrected text, determining an edit distance matrix between the text waiting for error correction and the error corrected text, the edit distance matrix comprising the number of edit operations between each character in the text waiting for error correction and each character in the error corrected text; and on the basis of the number of edit operations in the edit distance matrix, determining an error correction result of the text waiting for error correction. By means of text error correction of the large language model, text errors in different error correction modes can be corrected in a unified manner, and the edit correlation between the number of edit operations and text is comprehensively taken into account, thereby achieving high efficiency during text error correction; and by means of the error correction result, specific text errors and error correction modes can be more visually displayed.
Need to check novelty before this filing date? Find Prior Art

Description

Text error correction method, device, electronic device and storage medium Technical Field

[0001] The present invention relates to the field of text processing technology, and in particular to a text error correction method, device, system, electronic device and storage medium.

[0002] This application claims priority to the Chinese patent application filed with the China Patent Office on December 21, 2023, with application number 202311773526.9 and invention name “Text Error Correction Method, Device, Electronic Device and Storage Medium”, the entire contents of which are incorporated by reference into this application. Background Art

[0003] Traditional text correction methods use separate training models based on the correction method. For example, a spelling correction model is trained based on the correction method for spelling errors, and a grammar correction model is trained based on the correction method for grammatical errors. However, the need to train different correction models based on the correction method results in low efficiency and a lack of intuitive representation of the text error's problem and the correction method. Therefore, a text correction method that can perform text correction efficiently and more intuitively represent the text error's problem and the correction method has become an urgent problem to be solved. Technical issues

[0004] The embodiment of the present invention provides a text error correction method, which aims to solve the problems of low efficiency of existing text error correction and inability to intuitively reflect the location of text errors and the method of error correction.

[0005] In a first aspect, an embodiment of the present invention provides a text error correction method, the method comprising the following steps:

[0006] Get the text to be corrected;

[0007] Inputting the text to be corrected into the preset large-scale language model for text correction processing to obtain a corrected text;

[0008] Determining, based on the text to be corrected and the corrected text, an edit distance matrix between the text to be corrected and the corrected text, the edit distance matrix including the number of edit operations between each character in the text to be corrected and each character in the corrected text;

[0009] Based on the number of edit operations in the edit distance matrix, a correction result of the text to be corrected is determined.

[0010] In a second aspect, an embodiment of the present invention further provides a text error correction device, the text error correction device comprising:

[0011] A first acquisition module is used to acquire the text to be corrected;

[0012] A text error correction module, configured to input the text to be corrected into the preset large-scale language model for text error correction processing to obtain a corrected text;

[0013] A first determining module is configured to determine, based on the text to be corrected and the corrected text, an edit distance matrix between the text to be corrected and the corrected text, wherein the edit distance matrix includes a number of edit operations between each character in the text to be corrected and each character in the corrected text;

[0014] The second determining module is used to determine the error correction result of the text to be corrected based on the editing operation number in the editing distance matrix.

[0015] In a third aspect, an embodiment of the present invention provides an electronic device comprising: a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor implements the steps of the text error correction method provided in the embodiment of the present invention when executing the computer program.

[0016] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of the text error correction method provided in the embodiment of the invention are implemented.

[0017] In an embodiment of the present invention, a text to be corrected is obtained; the text to be corrected is input into the preset large-scale language model for text correction processing to obtain a corrected text; based on the text to be corrected and the corrected text, an edit distance matrix between the text to be corrected and the corrected text is determined, wherein the edit distance matrix includes the edit operations between each character in the text to be corrected and each character in the corrected text; and based on the edit operations in the edit distance matrix, a correction result of the text to be corrected is determined. The text to be corrected is corrected using a preset large-scale language model to obtain a corrected text, an edit distance matrix is ​​generated based on the edit operations between the corrected text and the text to be corrected, and a correction result is determined based on the edit distance matrix. Text correction using a large-scale language model can uniformly correct text errors of different correction methods, comprehensively considers the editing correlation between the edit operations and each text, thereby achieving higher efficiency in text correction, and more intuitively displays the problem of the text error and the correction method through the correction result. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] FIG1 is a flow chart of a text error correction method provided by an embodiment of the present invention;

[0019] FIG2 is a schematic structural diagram of an edit distance matrix provided by an embodiment of the present invention;

[0020] FIG3 is a schematic structural diagram of an error correction path provided by an embodiment of the present invention;

[0021] FIG4 is a flow chart of another text error correction method provided by an embodiment of the present invention;

[0022] FIG5 is a schematic structural diagram of a text error correction device provided in an embodiment of the present invention;

[0023] FIG6 is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. Modes for Carrying Out the Invention

[0024] As shown in FIG1 , FIG1 is a flowchart of a text error correction method provided by an embodiment of the present invention, including:

[0025] 101. Obtain the text to be corrected.

[0026] In an embodiment of the present invention, the above-mentioned text correction method can be applied to a text correction platform. The above-mentioned text correction platform can be constructed using a server or server cluster. The above-mentioned server or server cluster can be any electronic device with functions such as text recognition, text processing, data transmission, and data storage. The above-mentioned text correction platform can be deployed with the above-mentioned preset large-scale language model or can call the interface of any large model capable of text correction. The above-mentioned text correction platform can respond to the text to be corrected input by the user, perform text correction and other processing, and obtain the corrected text corresponding to the text to be corrected and the correction result.

[0027] The aforementioned preset large-scale language model can be understood as a large-scale deep learning model for text tasks, typically with a large number of parameters and computational overhead, and can be used to process text data and perform text error correction. The aforementioned preset large-scale language model can be obtained by training a preset large model in a specific vertical domain (i.e., text error correction capabilities). The aforementioned preset large model can be a large-scale deep learning model with a large number of parameters and computational overhead, and can be used to process text data. Any of the aforementioned large models capable of text error correction can be, for example, the ChatGPT large model, the Wenxin Yiyan large model, the Yunque large model, the Baichuan large model, etc.

[0028] The text to be corrected can be Chinese or English text with text errors, such as spelling errors, grammatical errors, or semantic errors. It is understood that each text to be corrected can be a string of characters, each string can contain multiple characters, and each string can also contain at least one punctuation mark. The text to be corrected can be specified text entered by the user, or it can be text obtained from any electronic device via a data transmission function.

[0029] 102. Input the text to be corrected into a preset large-scale language model for text correction processing to obtain a corrected text.

[0030] In an embodiment of the present invention, the preset large-scale language model requires corresponding instructions when performing text error correction processing, and the instructions can be used to instruct the preset large-scale language model to perform specific operations. For example, when the specific operation is text error correction processing, the instruction can be "Please help me correct the errors in the input text", and when the specific operation is text content expansion, the instruction can be "Please help me expand the content of the input text". It can be understood that when different instructions are input to the preset large-scale language model, the preset large-scale language model will respond to the instructions and perform corresponding processing.

[0031] Specifically, based on the above instructions, the content of the above instructions can be enriched and rich information can be provided to assist the above preset large-scale language model to better perform the above text error correction processing to obtain the output expected by the user. For example, when the above specific operation is text error correction processing, the above instruction can be "You are a text error correction expert. Please check whether there are errors in the input text, such as word order errors, word errors, grammatical errors and punctuation errors. If there are errors, please correct the errors and only output the corrected sentences. If there are no errors, directly output the original sentence."

[0032] When text correction is required, the text to be corrected and the instruction can be simultaneously input into the preset large-scale language model. The preset large-scale language model can then process the text to be corrected according to the instruction and output the corrected text. It is understood that the input of the preset large-scale language model includes the instruction and the text to be corrected, and the output of the preset large-scale language model includes the corrected text.

[0033] For example, the input of the above-mentioned preset large-scale language model when performing the above-mentioned text correction processing can be an instruction: "You are a text correction expert. Please check whether there are errors in the input text, such as word order errors, word errors, grammatical errors, and punctuation errors. If there are errors, please correct the errors and only output the corrected sentences. If there are no errors, directly output the original sentences. Text to be corrected: I like to eat apples. The output of the above-mentioned preset large-scale language model when performing the above-mentioned text correction processing can be the answer: "I like to eat apples."

[0034] 103. Based on the text to be corrected and the corrected text, determine an edit distance matrix between the text to be corrected and the corrected text.

[0035] In an embodiment of the present invention, the edit distance matrix includes the number of edit operations between each character in the text to be corrected and each character in the corrected text. The edit operation number can be used to measure the edit distance between the character strings composed of the characters in the two texts. The edit operation number can be the number of edit operations required to convert one string into another string. The edit operation can be insertion, deletion, or substitution. The edit distance can be understood as a metric for measuring the difference between two strings. More specifically, the edit distance measures the similarity between two strings by calculating the minimum number of edit operations between the two strings. It should be noted that the greater the number of edit operations, the greater the edit distance, and vice versa.

[0036] The above-mentioned number of edit operations can be calculated using a dynamic programming algorithm. This algorithm can be simply understood as splitting a large string into smaller strings. By finding a recursive relationship between the number of edit operations of the large string and the number of edit operations of the smaller strings, the edit operation numbers of the smaller strings are solved one by one, ultimately obtaining the number of edit operations of the larger string.

[0037] It should be noted that the order of characters in the above-mentioned character string is consistent with the order of characters in the above-mentioned text to be corrected or the corrected text. The above-mentioned character string may include all the characters in the above-mentioned text to be corrected and the above-mentioned text to be corrected, or the reverse order of the characters in the above-mentioned text to be corrected and the corrected text may be determined based on the above-mentioned character order, and the characters in the above-mentioned text to be corrected or the corrected text may be reduced accordingly based on the above-mentioned reverse order of the characters.

[0038] For example, if the corrected text includes the character "sitting", the character string corresponding to the corrected text may include the character string "sitting", the character string "sittin", the character string "sitti", the character string "sitt", the character string "sit", the character string "si", and the character string "s". If the text to be corrected includes the character "kitten", the character string corresponding to the text to be corrected may include the character string "kitten", the character string "kitte", the character string "kitt", the character string "kit", the character string "ki", and the character string "k".

[0039] Based on the character strings corresponding to the above-mentioned text to be corrected and the character strings corresponding to the above-mentioned corrected text, the number of editing operations between each character in the above-mentioned text to be corrected and each character in the text to be corrected is determined. Based on the above-mentioned editing operations, according to the character order of the above-mentioned text to be corrected and the character order of the corrected text, corresponding cells are generated to obtain the above-mentioned editing distance matrix.

[0040] For example, to convert "kitten" to "sitting," we need to replace "k" with "s," "e" with "i," and insert "g" at the end. The edit operations between "sitting" and "kitten" include replace, replace, and insert, so the number of edit operations is 3.

[0041] For example, to convert the string "k" to "sitting," replace "k" with "s," insert "i" after "s," insert "t" after "i," insert "t" after "t," insert "i" after "t," and insert "n" after "i." The edit operations between "k" and "sittin" include replace, insert, insert, insert, insert, and insert, resulting in a total of 6 edit operations.

[0042] 104. Determine the error correction result of the text to be corrected based on the number of edit operations in the edit distance matrix.

[0043] In an embodiment of the present invention, the correction result of the text to be corrected may include a correction method, a correction position, and a correction character. The correction method of the text to be corrected may be determined based on the edit operation number in the edit distance matrix. The correction position of the text to be corrected may be determined based on the cell position in the edit distance matrix corresponding to the correction method. The correction character to be corrected in the text to be corrected may be determined based on the correction position of the text to be corrected.

[0044] More specifically, the correction path of the text to be corrected can be determined based on the number of edit operations in the edit distance matrix. According to the correction path, the correction method of each correction cell in the correction path can be determined. According to the characters of the text to be corrected corresponding to each correction cell, the correction position and correction character of the text to be corrected can be determined. According to the correction method, correction position and correction text of each correction cell, the correction result of the text to be corrected can be determined.

[0045] It should be noted that the above edit distance matrix includes cells arranged in the above character sequence, and each cell corresponds to an edit operation number.

[0046] In an embodiment of the present invention, a text to be corrected is obtained; the text to be corrected is input into the preset large-scale language model for text correction processing to obtain a corrected text; based on the text to be corrected and the corrected text, an edit distance matrix between the text to be corrected and the corrected text is determined, wherein the edit distance matrix includes the edit operations between each character in the text to be corrected and each character in the corrected text; and based on the edit operations in the edit distance matrix, a correction result of the text to be corrected is determined. The text to be corrected is corrected using a preset large-scale language model to obtain a corrected text, an edit distance matrix is ​​generated based on the edit operations between the corrected text and the text to be corrected, and a correction result is determined based on the edit distance matrix. Text correction using a large-scale language model can uniformly correct text errors of different correction methods, comprehensively considers the editing correlation between the edit operations and each text, thereby achieving higher efficiency in text correction, and more intuitively displays the problem of the text error and the correction method through the correction result.

[0047] It is understandable that in the specific implementation of this application, related data such as text to be corrected and text that has been corrected are involved. When the embodiments in this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of relevant data and the use of large models need to comply with relevant laws, regulations and standards of relevant countries and regions.

[0048] Optionally, before the step of obtaining the text to be corrected, a preset large model and sample data can also be obtained; the sample data is preprocessed to obtain training data; the preset large model is adjusted and trained based on the training data to obtain a large-scale language model to be deployed; the large-scale language model to be deployed is deployed according to a preset deployment method, and the model parameters of the large-scale language model to be deployed are adjusted to obtain a preset large-scale language model.

[0049] In an embodiment of the present invention, the sample data includes sample text and labeled text, and the labeled text is obtained after error correction processing based on the sample text. The preset large model can be a large deep learning model with a large number of parameters and computing power that can be used to process text data, or it can be any existing large model capable of text error correction, such as the ChatGPT large model, the Wenxin Yiyan large model, the Yunque large model, the Baichuan large model, etc. The sample data can be open source text error correction data, for example, it can be error correction data including spelling check: sighan13, sighan14, sighan15, etc., error correction data including grammatical correction: NLPCC, CGED, etc., and CSED, etc., including semantic correction. Alternatively, the relevant text error correction data can be generated according to certain rules. For example, the sample text can be obtained by replacing the font or pinyin of a public text such as a journal or a news speech, or reversing the order of words and sentences in the public text.

[0050] The preprocessing may include proofreading, a first filtering process, a screening process, and a second filtering process. Specifically, the encoding format of the sample data may be proofread, and text outside a preset encoding range may be discarded to obtain proofread sample data. Poor text quality in the proofread sample data is filtered to obtain sample data after the first filtering process. The poor text quality sample data may include, for example, missing punctuation between different sentences or containing a large number of emoticons.

[0051] The sample data after the first filtering process is filtered out of any text that does not conform to human values, thereby obtaining filtered sample data. The text that does not conform to human values ​​may be text that contains speech related to terrorism, fraud, violence, etc. The sample data after the filtering process is filtered in both Chinese and English. For example, when Chinese text error correction is required, the English text length exceeding half of the total text length in the sample data after the filtering process needs to be filtered out. When English text error correction is required, the Chinese text length exceeding half of the total text length in the sample data after the filtering process needs to be filtered out, thereby obtaining cleaned sample data.

[0052] The cleaned sample data may be grammatical correction data, semantic correction data or spelling correction data. Each cleaned sample data may include a sample text and a label text.

[0053] For example, if the cleaned sample data is spelling-corrected data, the sample text might be "I like eating apples." The label text might be "I like eating apples." If the cleaned sample data is grammatically corrected data, the sample text might be "Three seriously injured soldiers are undergoing surgery." The label text might be "Three seriously injured soldiers are undergoing surgery." If the cleaned sample data is semantically corrected data, the sample text might be "He drove very fast and rushed to the hospital." The label text might be "He drove very fast and rushed to the hospital."

[0054] After obtaining the cleaned sample data, the cleaned sample data can be converted into the training data according to the instructions and a preset format. Specifically, for example, if the cleaned sample data is spelling correction data, and the sample text is "I like eating apples," and the label text is "I like eating apples," and the instruction is "You are a text correction expert. Please check the input text for errors, such as errors in word order, word errors, grammatical errors, and punctuation. If errors exist, please correct them and only output the corrected sentence. If no errors exist, please output the original sentence directly.", then the training data can be {"Question" (i.e., the instruction): "You are a text correction expert. Please check the input text for errors, such as errors in word order, word errors, grammatical errors, and punctuation. If errors exist, please correct them and only output the corrected sentence. If no errors exist, please output the original sentence directly. Input: I like eating apples. Output: "," "Answer": "I like eating apples"}.

[0055] After obtaining the aforementioned training data, a portion of the training data can be randomly extracted to obtain verification data. The aforementioned large model is then trained iteratively using the training data to ensure that the large model follows the instructions and outputs in the format expected by the instructions. During the training iterations, the accuracy of the large model in text correction is continuously verified using the verification data. When the accuracy reaches a preset value, the training iterations are terminated, resulting in the aforementioned large-scale language model to be deployed.

[0056] After obtaining the above-mentioned large-scale language model to be deployed, the above-mentioned large-scale language model to be deployed can be deployed according to a preset deployment method, and the model parameters of the large-scale language model to be deployed can be adjusted to obtain the above-mentioned preset large-scale language model.

[0057] The preset deployment method can be to deploy the large-scale language model as an interface using a backend framework, or to deploy the large-scale language model using an inference framework. The backend framework can include fastapi, Django, etc., and the inference framework can include TGI (text-generation-inference), LMDeploy, vLLM, etc.

[0058] For example, after deploying the large-scale language model using the aforementioned inference framework TGI (text-generation-inference), a request URL is generated. The general format of the request URL is: http: / / ip address:port / generate or generate_stream. The `generate` method returns the complete output of the deployed large-scale language model, while the `generate_stream` method returns the output of the deployed large-scale language model in a streaming format, returning a portion of the results at a time. When sending a request to the deployed large-scale language model using the aforementioned sending URL, you can adjust certain parameters to ensure the consistency of the results generated by the deployed large-scale language model. These parameters may include topk, do_sample, and topP. If the text correction processing is for Chinese text, do_sample is set to False and topK is set to 1 to ensure a consistent format for the output of the deployed large-scale language model. This approach is called greedy decoding. After adjusting these parameters, the preset large-scale language model is obtained.

[0059] Specifically, greedy decoding means that when generating text, the model predicts the probability distribution of the next token and selects the token with the highest probability as the next token. This decoding method only considers the token with the highest probability, hence the name greedy decoding. This ensures that the token selected each time is the most likely token in the current probability distribution, resulting in more coherent and accurate text.

[0060] In Chinese text error correction tasks, since the model's output may vary depending on the input text, greedy decoding is often used to ensure consistency. Setting do_sample to False ensures that the model only selects the token with the highest probability as the next token, thus avoiding generating inconsistent results. Furthermore, setting topK to 1 limits the model to considering only the token with the highest probability, further ensuring consistency in the output.

[0061] Optionally, in the step of determining an edit distance matrix between the text to be corrected and the corrected text, where the text to be corrected includes a first character string and the corrected text includes a second character string, based on the text to be corrected and the corrected text, the minimum number of edit operations required to convert each character in the first character string into each character in the second character string can also be calculated; and the edit distance matrix can be generated according to the correspondence between the minimum number of edit operations and each character in the first character string and each character in the second character string.

[0062] In an embodiment of the present invention, the minimum number of editing operations required to convert each character in the first string into each character in the second string can be calculated by the above-mentioned dynamic programming algorithm. According to the correspondence between the minimum number of editing operations and each character in the first string and each character in the second string, a cell in the editing distance matrix is ​​generated, and the above-mentioned cells are arranged in the character order of the above-mentioned correspondence to obtain the above-mentioned editing distance matrix. It should be noted that the above-mentioned first string and the above-mentioned second string can be a single string or multiple strings, and the number of the above-mentioned first strings is positively correlated with the number of characters in the above-mentioned text to be corrected, and the number of the above-mentioned second strings is positively correlated with the number of characters in the above-mentioned text that has been corrected.

[0063] The minimum number of editing operations may be understood as the minimum number of editing operations required to convert one string into another string. The editing operations may be replacement processing, insertion processing, and deletion processing.

[0064] Specifically, the above edit distance matrix can be illustrated by a structural diagram of an edit distance matrix as shown in Figure 2. In Figure 2, the text to be corrected is "strong sensational effect", and the corrected text is "strong repercussion". The first string includes "strong", "strongly", "strong's", "strong's sensational", "strong's sensational effect", "strong's sensational effect", and "strong's sensational effect", and the second string includes "strong", "strongly", "strong's", "strong's reaction", and "strong's repercussion". The cells corresponding to the first column of the second row and the first row of the second column in the matrix are empty characters. It can be seen that the edit operation numbers of the cells in the second row increase by 1 successively, and the edit operation number of the last cell in the second row is 7, indicating the minimum number of edit operations required to gradually change from an empty character to the first string. The edit operation numbers of the cells in the second column increase by 1 successively, and the edit operation number of the last cell in the second column is 5, indicating the minimum number of edit operations required to gradually evolve from an empty character to the second string.

[0065] For example, the minimum edit operation number corresponding to the cell in the fifth column of the third row in Figure 2 is 2, which indicates that the minimum number of edit operations required to convert the first string "strong's" to the second string "strong" is 2. It can be seen that the first edit operation to convert the first string "strong's" to the second string "strong" is to delete the character "ly" in the first string, and the second edit operation is to delete the character "s" in the first string.

[0066] Or take the cell in the sixth column of the third row in Figure 2 as an example. The minimum edit operation number corresponding to the cell in the sixth column of the third row is 3, which indicates that the minimum number of edit operations required to convert the first string "strong's sensational" to the second string "strong" is 3. It can be seen that the first edit operation to convert the first string "strong's sensational" to the second string "strong" is to delete the character "ly" in the first string, the second edit operation is to delete the character "s" in the first string, and the third edit operation is to delete the character "sensational" in the first string. In Figure 2, the generation methods of the remaining cells are the same as those in the above examples, and the effects are the same. To avoid duplication and redundancy, they will not be repeated here.

[0067] Optionally, in the step of determining the correction result of the text to be corrected based on the edit operation numbers in the edit distance matrix, the starting cell can also be determined according to the reverse order of the characters of the first string and the second string; starting from the starting cell, traverse the cells in the preset direction in the edit distance matrix based on the edit operation numbers, and determine the correction path in the edit distance matrix; determine the correction result of the text to be corrected based on the correction path.

[0068] In an embodiment of the present invention, the error correction path includes error correction cells in a preset direction, and the starting cell may be a cell corresponding to the last character in the first character string and the last character in the second character string. During the traversal process, starting from the starting cell, the edit operation counts of the adjacent cells in the preset direction of the currently processed error correction cell are traversed in reverse order of the characters. Based on the edit operation counts of the adjacent cells in the preset direction and the current error correction cell, the next error correction cell is determined, and the error correction path is determined based on the edit distance matrix of all error correction cells.

[0069] It can be understood that, in the above error correction path, the position of the terminating cell may be the cell corresponding to the first character in the above first character string and the first character in the above second character string.

[0070] Optionally, starting from the starting cell, cells in a preset direction are traversed in the edit distance matrix based on the edit operation number, and in the step of determining the error correction path in the edit distance matrix, adjacent cells of the current error correction cell in the preset direction can also be determined in the edit distance matrix, the adjacent cells including a first adjacent cell in a first preset direction, a second adjacent cell in a second preset direction, and a third adjacent cell in a third preset direction; when the edit operation number of the first adjacent cell is less than the edit operation number of the current error correction cell, and the first character and the second character corresponding to the current error correction cell are different, the first adjacent cell is used as the next error correction cell; when the edit operation number of the second adjacent cell is less than the edit operation number of the current error correction cell, the second adjacent cell is used as the next error correction cell; when the edit operation number of the third adjacent cell is less than the edit operation number of the current error correction cell, the third adjacent cell is used as the next error correction cell; when the edit operation number of the first adjacent cell is equal to the edit operation number of the current cell, the first adjacent cell is used as the next error correction cell; based on all error correction cells, the error correction path is determined in the edit distance matrix.

[0071] In an embodiment of the present invention, the first character string includes each first character, the second character string includes each second character, the preset direction may include a first preset direction, a second preset direction and a third preset direction, the first preset direction may be a left diagonal direction, the second preset direction may be a left direction, and the third preset direction may be an upward direction.

[0072] It should be noted that when judging the error-correcting cell in the next step, traversal can be performed according to the order of the above preset directions. That is, in the first judgment, it is judged whether the number of editing operations of the first adjacent cell in the first preset direction is less than the number of editing operations of the current error-correcting cell. In the second judgment, the second adjacent cell in the second preset direction is judged. In the third judgment, the third adjacent cell in the third preset direction is judged. In the fourth judgment, it is judged whether the number of editing operations of the first adjacent cell in the first preset direction is equal to the number of editing operations of the current error-correcting cell.

[0073] When the number of editing operations of the above first adjacent cell is less than the number of editing operations of the above current error-correcting cell, then the above first adjacent cell is the error-correcting cell in the next step. At this time, there is no need to judge the second adjacent cell in the second preset direction and the third adjacent cell in the third direction. When the number of editing operations of the above first adjacent cell is not less than the number of editing operations of the above current error-correcting cell, at this time, it can be judged whether the number of editing operations of the above second adjacent cell is less than the number of editing operations of the current error-correcting cell. When the number of editing operations of the above second adjacent cell is less than the number of editing operations of the current error-correcting cell and less than the number of editing operations of the current error-correcting cell, then the above second adjacent cell is the error-correcting cell in the next step. At this time, there is no need to judge the above third adjacent cell. When the number of editing operations of the above third adjacent cell is less than the number of editing operations of the current error-correcting cell, then the above third adjacent cell is the error-correcting cell in the next step. At this time, there is no need to judge whether the number of editing operations of the above first adjacent cell is less than the number of editing operations of the current error-correcting cell.

[0074] Specifically, the above error-correcting path can be illustrated by a structural schematic diagram of an error-correcting path as shown in Figure 3. In Figure 3, the current error-correcting cell in the first step is the starting cell (i.e., the cell in the ninth column of the seventh row in the figure). The number of editing operations of the cell in its left diagonal direction is 3, which is less than the number of editing operations of the cell in the ninth column of the seventh row, which is 4. And the first character "应" corresponding to the cell in the ninth column of the seventh row is different from the second character "响". Then the error-correcting cell in the next step is the cell in the eighth column of the sixth row.

[0075] The current error-correcting cell in the second step is the cell in the eighth column of the sixth row. The number of editing operations of the cell in its left diagonal direction is 2, which is less than the number of editing operations of the cell in the eighth column of the sixth row, which is 3. And the first character "效" corresponding to the cell in the eighth column of the sixth row is different from the second character "反". Then the error-correcting cell in the next step is the cell in the seventh column of the fifth row.

[0076] The current error correction cell in the third step is the cell in the fifth row and seventh column. The editing operation number of the cell in the left diagonal direction is 2, which is not less than the editing operation number 2 of the cell in the fifth row and seventh column. Then the editing operation number of the cell in the left direction is 1, which is less than the editing operation number 2 of the cell in the fifth row and seventh column. Then the next error correction cell is the cell in the fifth row and sixth column.

[0077] The current error correction cell in the fourth step is the cell in the fifth row and sixth column. The number of cell editing operations in the left diagonal direction is 1, which is not less than the number of editing operations 1 of the cell in the fifth row and sixth column. Then the number of editing operations of the cell in the left direction is 0, which is less than the number of editing operations 1 of the cell in the fifth row and sixth column. Then the next error correction cell is the cell in the fifth row and fifth column.

[0078] The current error correction cell in the fifth step is the cell in the fifth row and fifth column. The editing operation number of the cell in the left diagonal direction is 0, which is not less than the editing operation number 0 of the cell in the fifth row and fifth column. Then the editing operation number of the cell in the left direction is 1, which is not less than the editing operation number 0 of the cell in the fifth row and fifth column. Then the editing operation number of the cell in the upward direction is 1, which is not less than the editing operation number 0 of the cell in the fifth row and fifth column. Then the editing operation number of the cell in the left diagonal direction is 0, which is equal to the editing operation number 0 of the cell in the fifth row and fifth column. Then the error correction cell for the next step is the cell in the fourth row and fourth column.

[0079] Subsequently, the current error correction cell in the sixth step is used to determine the current error correction cell in the seventh step. The steps of determining the current error correction cell in the eighth step through the current error correction cell in the seventh step are consistent with the steps in the fifth step above and have the same effect. To avoid repetition and redundancy, they will not be repeated here.

[0080] Based on the current error correction cells obtained from the first to eighth steps above, the vectors from each step to the next step (i.e., the arrows in FIG3 ) are determined, and the error correction path is obtained according to the order of the vectors.

[0081] Optionally, in the step of determining the error correction result of the text to be error-corrected based on the error correction path, when the next error correction cell corresponding to the current error correction cell is the first adjacent cell, and the first character corresponding to the current error correction cell is different from the second character, the error correction method for the current error correction cell is replacement processing; when the next error correction cell corresponding to the current error correction cell is the second adjacent cell, the error correction method for the current error correction cell is deletion processing; when the next error correction cell corresponding to the current error correction cell is the third adjacent cell, the error correction method for the current error correction cell is insertion processing; when the next error correction cell corresponding to the current error correction cell is the first adjacent cell, and the edit operation count of the current error correction cell is equal to the edit operation count of the first adjacent cell, the error correction method for the current error correction cell is not to perform processing; based on the error correction methods of each error correction cell in the error correction path, determine the error correction result of the text to be error-corrected.

[0082] In an embodiment of the present invention, when making the above determination of the error correction method, it may not be determined according to the above order. That is, when making the determination of the error correction method, when the current error correction cell meets the conditions of any one of the error correction methods of replacement processing, insertion processing, deletion processing, and not performing processing, the error correction method of the current error correction cell can be directly determined.

[0083] Specifically, the first to fourth determinations of the above error correction method can be illustrated by taking a structural schematic diagram of an error correction path as shown in Figure 3. If the current error correction cell is the starting cell (i.e., the cell in the ninth column of the seventh row in the figure), and its next error correction cell is the cell in the eighth column of the sixth row, which is a cell in the left diagonal direction, and the first character "should" corresponding to the cell in the ninth column of the seventh row is different from the second character "ring", then the error correction method for the current error correction cell is replacement processing.

[0084] If the current error correction cell is the cell in the seventh column of the fifth row, and its next error correction cell is the cell in the sixth column of the fifth row, which is a cell in the left direction, then the error correction method for the current error correction cell is deletion.

[0085] If the current error correction cell is the cell in the fifth column of the fifth row, and its next error correction cell is the cell in the fourth column of the fourth row, which is a cell in the left diagonal direction, and the edit operation count of the cell in the fifth column of the fifth row is 0, which is equal to the edit operation count 0 of the cell in the fourth column of the fourth row, then the error correction method for the current error correction cell is not to perform processing.

[0086] It can be seen that the error correction method for the cell in the seventh row and ninth column in Figure 3 is replacement processing. Specifically, it can be to replace the first character '应' with the second character '响' corresponding to the cell in the seventh row and ninth column. The error correction method for the cell in the sixth row and eighth column can be to replace the first character '效' with the second character '反' corresponding to the cell in the sixth row and eighth column. The error correction method for the cell in the fifth row and seventh column is deletion processing. Specifically, it can be to delete the first character '动' corresponding to the cell in the fifth row and seventh column. The error correction method for the cell in the fifth row and sixth column is deletion processing. Specifically, it can be to delete the first character '轰' corresponding to the cell in the fifth row and sixth column. The error correction cells that do not undergo processing on the remaining error correction paths can be not included in the above error correction results.

[0087] Finally, according to the above error correction methods, it can be determined that the error correction result is to replace '应' in the text to be error corrected with '响', replace '效' with '反', delete '动', and delete '轰', thereby obtaining the above error corrected text.

[0088] More specifically, assume that it has currently run to the cell in the i-th row and j-th column. If the first character corresponding to the cell in the i-th row and j-th column is the same as the second character, the operation to be performed currently is d = 0; otherwise, d = 1.

[0089] Traverse and judge in reverse starting from the starting position of the edit distance matrix in Figure 3 (i.e., the cell in the seventh row and ninth column):

[0090] The first judgment: If dp[i - 1][j - 1] + 1 is equal to dp[i][j], and word1[i] is not equal to word2[j], it means that word1[i] is to be replaced with word2[j], and then i is decremented by 1 and j is decremented by 1.

[0091] Specifically, the above first judgment means that if the edit operation count of the cell dp[i - 1][j - 1] plus one is equal to the edit operation count of the cell dp[i][j], and the first character corresponding to the cell dp[i][j] is not equal to the second character, it means that the first character corresponding to the cell dp[i][j] in the text to be error corrected needs to be replaced with the second character. The next cell to be error corrected is dp[i - 1][j - 1].

[0092] The second judgment: If dp[i][j - 1] + 1 is equal to dp[i][j], it means that word1[j] is to be deleted, and then j is decremented by 1.

[0093] Specifically, the second judgment is that if the number of edit operations for the cell dp[i][j-1] plus one equals the number of edit operations for the cell dp[i][j], then the first character corresponding to the cell dp[i][j] in the text to be corrected needs to be deleted. The next cell to be corrected is dp[i][j-1].

[0094] The third judgment: If dp[i-1][j]+1 is equal to dp[i][j], it means that word2[i] is to be inserted into word1[j], and then i is reduced by 1.

[0095] Specifically, the third judgment is expressed as follows: if the edit operation count of the cell dp[i-1][j]+1 plus one equals the edit operation count of the cell dp[i][j], then it means that the second character needs to be inserted at the position of the first character corresponding to the cell dp[i][j] in the text to be corrected. The next cell to be corrected is dp[i-1][j].

[0096] Fourth judgment: If dp[i-1][j-1] is equal to dp[i][j], it means that no operation is required, then i is reduced by 1 and j is reduced by 1.

[0097] Specifically, the fourth judgment above means that if the number of edit operations of dp[i-1][j-1] is equal to the number of edit operations of dp[i][j], then no operation is required. The next cell to be corrected is dp[i-1][j-1].

[0098] Among them, the above d represents the operation that needs to be performed currently (that is, whether the judgment operation needs to be performed), the above p represents the cell, the above i represents the cell in the i-th row, the above j represents the cell in the j-th column, the above i-1 represents the cell in the i-1th row, the above j-1 represents the cell in the j-1th column, and the above +1 represents the operand plus one.

[0099] Optionally, in the step of determining the correction result of the text to be corrected based on the correction method of each correction cell in the correction path, it is also possible to determine the correction position corresponding to the correction method in the text to be corrected based on the correction cell corresponding to the correction method; determine the correction character of the text to be corrected in the text to be corrected based on the correction position; and determine the correction result of the text to be corrected based on the correction method, correction position and correction character.

[0100] In an embodiment of the present invention, based on the error correction cells corresponding to the above error correction methods, the row position and column position of the error correction cells in the edit distance matrix can be determined. According to the column position of the error correction cells in the edit distance matrix, the error correction position corresponding to the error correction method (i.e., the column position of the text to be error corrected) can be determined. According to the column position of the text to be error corrected, the characters of the text to be error corrected at the column position can be determined. According to the error correction methods, error correction positions, and error correction characters corresponding to each error correction method, the error correction result of the text to be error corrected can be determined.

[0101] It should be noted that the error correction result of the text to be error corrected includes each sub-error correction result. Each sub-error correction result includes an error correction method (i.e., it can be a replacement process, an insertion process, or a deletion process), an error correction position, and an error correction character.

[0102] Specifically, the above error correction result can be illustrated by a structural schematic diagram of an error correction path shown in FIG. 3. In FIG. 3, the error correction methods that need to be performed include a deletion process and a replacement process. The error correction result corresponding to the deletion process is [3, 4, 'hōng dòng', 'delete'], and the error correction result corresponding to the replacement process is [5, 6, 'fǎn xiǎng','replace']. Then the error correction result of the text to be error corrected is [[3, 4, 'hōng dòng', 'delete'], [5, 6, 'fǎn xiǎng','replace']].

[0103] It should be noted that each error correction method corresponds to a sub-error correction result. The first digit of each sub-error correction result represents the starting position of the error correction method, the second digit represents the ending position of the error correction method, the third digit represents the text to be operated on, and the fourth digit represents the operation method. Among them, the position 0 to be operated on in the error correction result represents the first character "qiáng" of the text to be error corrected in the figure from left to right, the position 1 represents the second character "liè" of the text to be error corrected in the figure from left to right, the position 2 represents the third character "de" of the text to be error corrected in the figure from left to right, the position 3 represents the fourth character "hōng" of the text to be error corrected in the figure from left to right, the position 4 represents the fifth character "dòng" of the text to be error corrected in the figure from left to right, the position 5 represents the sixth character "xiào" of the text to be error corrected in the figure from left to right, and the position 6 represents the seventh character "yǐng" of the text to be error corrected in the figure from left to right. The above [3, 4, 'hōng dòng', 'delete'] means deleting the characters "hōng dòng" at positions 3 and 4 in the text to be error corrected, and the above [5, 6, 'fǎn xiǎng','replace'] means replacing the characters at positions 5 and 6 in the above text to be error corrected with "fǎn xiǎng".

[0104] As shown in Figure 4, an embodiment of the present invention also provides a flowchart of another text correction method. As can be seen from Figure 4, before starting text correction, data collection (i.e., sample data) is required. Then, the sample data is preprocessed to obtain training data. The model is trained and verified based on the training data to obtain a model to be deployed. The model to be deployed is then deployed to obtain a model interface. When text correction begins, the text to be corrected is obtained from the user, and the format of the text to be corrected is converted to meet the input requirements of the model and achieve the user's desired output format. The converted text to be corrected is input into the model interface, and prediction is performed using the model to obtain a prediction result (i.e., the corrected text). Errors are located based on the corrected text and the text to be corrected, and a correction method is determined, ending the text correction process.

[0105] As shown in FIG5 , an embodiment of the present invention further provides a text error correction device, comprising:

[0106] The first acquisition module 501 is used to acquire the text to be corrected;

[0107] The text error correction module 502 is used to input the text to be corrected into the preset large-scale language model for text error correction processing to obtain a corrected text;

[0108] A first determining module 503 is configured to determine an edit distance matrix between the text to be corrected and the corrected text based on the text to be corrected and the corrected text, wherein the edit distance matrix includes a number of edit operations between each character in the text to be corrected and each character in the corrected text;

[0109] The second determining module 504 is configured to determine the error correction result of the text to be corrected based on the number of edit operations in the edit distance matrix.

[0110] Optionally, the text error correction device further includes:

[0111] A second acquisition module is used to acquire a preset large model and sample data, wherein the sample data includes sample text and label text, and the label text is obtained after error correction processing is performed on the sample text;

[0112] A preprocessing module, used to preprocess the sample data to obtain training data;

[0113] An adjustment training module is used to adjust and train the preset large model based on the training data to obtain a large-scale language model to be deployed;

[0114] The deployment module is used to deploy the large-scale language model to be deployed according to a preset deployment method, and adjust the model parameters of the large-scale language model to be deployed to obtain the preset large-scale language model.

[0115] Optionally, the first determining module 503 includes:

[0116] a calculation submodule, configured to calculate a minimum number of editing operations required to convert each character in the first character string into each character in the second character string;

[0117] A generating submodule is configured to generate the edit distance matrix according to the corresponding relationship between the minimum number of edit operations and each character in the first character string and each character in the second character string.

[0118] Optionally, the second determining module 504 includes:

[0119] A first determining submodule, configured to determine a starting cell according to the reverse order of characters of the first character string and the second character string;

[0120] a traversal submodule, configured to start from the starting cell and traverse cells in a preset direction in the edit distance matrix based on the edit operand, and determine an error correction path in the edit distance matrix, wherein the error correction path includes error correction cells in the preset direction;

[0121] The second determining submodule is used to determine the error correction result of the text to be corrected based on the error correction path.

[0122] Optionally, the traversal submodule includes:

[0123] A first determining unit is configured to determine, in the edit distance matrix, adjacent cells of a current error correction cell in a preset direction, wherein the adjacent cells include a first adjacent cell in a first preset direction, a second adjacent cell in a second preset direction, and a third adjacent cell in a third preset direction;

[0124] a first comparing unit, configured to use the first adjacent cell as a next error correction cell when the number of edit operations of the first adjacent cell is less than the number of edit operations of the current error correction cell, and the first character and the second character corresponding to the current error correction cell are different;

[0125] a second comparing unit, configured to use the second adjacent cell as a next error correction cell when the number of edit operations of the second adjacent cell is less than the number of edit operations of the current error correction cell;

[0126] a third comparing unit, configured to use the third adjacent cell as a next error correction cell when the number of edit operations of the third adjacent cell is less than the number of edit operations of the current error correction cell;

[0127] a fourth comparing unit, configured to use the first adjacent cell as a next error correction cell when the number of edit operations of the first adjacent cell is equal to the number of edit operations of the current cell;

[0128] The second determining unit is configured to determine the error correction path in the edit distance matrix based on all error correction cells.

[0129] Optionally, the second determining submodule includes:

[0130] a fifth comparing unit, configured to, when the next error correction cell corresponding to the current error correction cell is the first adjacent cell, and the first character and the second character corresponding to the current error correction cell are different, perform a replacement process on the error correction method of the current error correction cell;

[0131] a sixth comparing unit, configured to, when the next error correction cell corresponding to the current error correction cell is the second adjacent cell, perform an error correction mode of the current error correction cell as deletion;

[0132] a seventh comparing unit, configured to, when the next error correction cell corresponding to the current error correction cell is the third adjacent cell, use an insertion process as the error correction method for the current error correction cell;

[0133] an eighth comparing unit, configured to, when the next correcting cell corresponding to the current correcting cell is the first adjacent cell, and the number of edit operations of the current correcting cell is equal to the number of edit operations of the first adjacent cell, select no processing for the correcting cell;

[0134] The third determining unit is used to determine the error correction result of the text to be corrected based on the error correction method of each error correction unit in the error correction path.

[0135] Optionally, the third determining unit includes:

[0136] A first determining subunit is configured to determine, based on the error correction cell corresponding to the error correction method, an error correction position corresponding to the error correction method in the text to be corrected;

[0137] A second determining subunit is configured to determine, based on the error correction position, an error correction character of the text to be corrected in the text to be corrected;

[0138] The third determining subunit is used to determine the error correction result of the text to be corrected based on the error correction method, error correction position and error correction character.

[0139] As shown in FIG6 , an embodiment of the present invention further provides an electronic device, characterized in that it includes a processor, and the processor can execute any one of the above-mentioned text error correction methods.

[0140] Specifically, it includes a processor 601 and a memory 602, and a computer program for executing the text error correction method stored in the memory 602 and capable of running on the processor 601, wherein: the processor 601 runs the computer program for executing the text error correction method stored in the memory 602. The specific execution steps of the text error correction method refer to the text error correction method in the above embodiment, which will not be repeated here.

[0141] Optionally, the determining of the error correction result of the to-be-corrected text based on the edit operation number in the edit distance matrix performed by the processor 601 includes:

[0142] Determine a starting cell according to the reverse order of characters of the first character string and the second character string;

[0143] Starting from the starting cell, traversing cells in a preset direction in the edit distance matrix based on the edit operand, and determining an error correction path in the edit distance matrix, the error correction path including the error correction cells in the preset direction;

[0144] The error correction result of the text to be corrected is determined based on the error correction path.

[0145] Optionally, the first string executed by the processor 601 includes each first character, the second string includes each second character, the preset direction includes a first preset direction, a second preset direction, and a third preset direction, and starting from the starting cell, traversing cells in the preset direction in the edit distance matrix based on the edit operand, and determining the error correction path in the edit distance matrix includes:

[0146] In the edit distance matrix, determining adjacent cells of a current error correction cell in a preset direction, the adjacent cells including a first adjacent cell in a first preset direction, a second adjacent cell in a second preset direction, and a third adjacent cell in a third preset direction;

[0147] When the number of edit operations of the first adjacent cell is less than the number of edit operations of the current error correction cell, and the first character and the second character corresponding to the current error correction cell are different, the first adjacent cell is used as the next error correction cell;

[0148] When the number of edit operations of the second adjacent cell is less than the number of edit operations of the current error correction cell, the second adjacent cell is used as the next error correction cell;

[0149] When the number of edit operations of the third adjacent cell is less than the number of edit operations of the current error correction cell, the third adjacent cell is used as the next error correction cell;

[0150] When the number of edit operations of the first adjacent cell is equal to the number of edit operations of the current cell, the first adjacent cell is used as the next error correction cell;

[0151] The error correction path is determined in the edit distance matrix based on all error correction cells.

[0152] Optionally, the first character string executed by the processor 601 includes each first character, the second character string includes each second character, and determining the error correction result of the text to be corrected based on the error correction path includes:

[0153] When the next error correction cell corresponding to the current error correction cell is the first adjacent cell, and the first character and the second character corresponding to the current error correction cell are different, the error correction method of the current error correction cell is replacement processing;

[0154] When the next error correction cell corresponding to the current error correction cell is the second adjacent cell, the error correction method of the current error correction cell is deletion processing;

[0155] When the next error correction cell corresponding to the current error correction cell is the third adjacent cell, the error correction method of the current error correction cell is insertion processing;

[0156] When the next error correction cell corresponding to the current error correction cell is the first adjacent cell, and the number of edit operations of the current error correction cell is equal to the number of edit operations of the first adjacent cell, the error correction mode of the current error correction cell is no processing;

[0157] Based on the error correction method of each error correction cell in the error correction path, the error correction result of the text to be corrected is determined.

[0158] Optionally, the determining of the error correction result of the to-be-corrected text based on the error correction mode of each error correction cell in the error correction path performed by the processor 601 includes:

[0159] Based on the error correction cell corresponding to the error correction method, determining the error correction position corresponding to the error correction method in the text to be corrected;

[0160] Determining the error correction character of the text to be corrected in the text to be corrected based on the error correction position;

[0161] Based on the error correction method, error correction position and error correction character, the error correction result of the text to be corrected is determined.

[0162] An embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the computer program implements the various processes of the text error correction method or the application-side text error correction method provided in the embodiment of the present invention, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.

Claims

1. A text error correction method, characterized in that: The method comprises the following steps: Get the text to be corrected; Inputting the text to be corrected into the preset large-scale language model for text correction processing to obtain a corrected text; Based on the text to be corrected and the corrected text, determining an edit distance matrix between the text to be corrected and the corrected text, wherein the edit distance matrix includes the number of edit operations between each character in the text to be corrected and each character in the corrected text; Based on the number of edit operations in the edit distance matrix, a correction result of the text to be corrected is determined.

2. The text error correction method according to claim 1, characterized in that: Before obtaining the text to be corrected, the method further includes: Obtaining a preset macro model and sample data, wherein the sample data includes sample text and label text, wherein the label text is obtained after error correction processing is performed on the sample text; Preprocessing the sample data to obtain training data; Adjusting and training the preset large model based on the training data to obtain a large-scale language model to be deployed; The large-scale language model to be deployed is deployed according to a preset deployment mode, and the model parameters of the large-scale language model to be deployed are adjusted to obtain the preset large-scale language model.

3. The text error correction method according to claim 1 or 2, characterized in that: The text to be corrected includes a first character string, the corrected text includes a second character string, and determining the edit distance matrix between the text to be corrected and the corrected text based on the text to be corrected and the corrected text includes: Calculating the minimum number of editing operations required to convert each character in the first string to each character in the second string; The edit distance matrix is ​​generated according to the correspondence between the minimum number of edit operations and each character in the first character string and each character in the second character string.

4. The text error correction method according to claim 3, characterized in that: The step of determining the error correction result of the text to be corrected based on the number of edit operations in the edit distance matrix comprises: Determine a starting cell according to the reverse order of characters of the first character string and the second character string; Starting from the starting cell, traversing cells in a preset direction in the edit distance matrix based on the edit operand, determining an error correction path in the edit distance matrix, the error correction path including the error correction cells in the preset direction; The error correction result of the text to be corrected is determined based on the error correction path.

5. The text error correction method according to claim 4, characterized in that: The first character string includes each first character, the second character string includes each second character, the preset direction includes a first preset direction, a second preset direction, and a third preset direction, and starting from the starting cell, traversing cells in the preset direction in the edit distance matrix based on the edit operand, and determining an error correction path in the edit distance matrix includes: In the edit distance matrix, determining adjacent cells of a current error correction cell in a preset direction, wherein the adjacent cells include a first adjacent cell in a first preset direction, a second adjacent cell in a second preset direction, and a third adjacent cell in a third preset direction; When the number of edit operations of the first adjacent cell is less than the number of edit operations of the current error correction cell, and the first character and the second character corresponding to the current error correction cell are different, the first adjacent cell is used as the next error correction cell; When the number of edit operations of the second adjacent cell is less than the number of edit operations of the current error correction cell, the second adjacent cell is used as the next error correction cell; When the number of edit operations of the third adjacent cell is less than the number of edit operations of the current error correction cell, the third adjacent cell is used as the next error correction cell; When the number of edit operations of the first adjacent cell is equal to the number of edit operations of the current cell, the first adjacent cell is used as the next error correction cell; The error correction path is determined in the edit distance matrix based on all error correction cells.

6. The text error correction method according to claim 5, characterized in that: The step of determining the error correction result of the text to be corrected based on the error correction path includes: When the next error correction cell corresponding to the current error correction cell is the first adjacent cell, and the first character and the second character corresponding to the current error correction cell are different, the error correction method of the current error correction cell is replacement processing; When the next error correction cell corresponding to the current error correction cell is the second adjacent cell, the error correction method of the current error correction cell is deletion processing; When the next error correction cell corresponding to the current error correction cell is the third adjacent cell, the error correction method of the current error correction cell is insertion processing; When the next error correction cell corresponding to the current error correction cell is the first adjacent cell, and the number of edit operations of the current error correction cell is equal to the number of edit operations of the first adjacent cell, the error correction mode of the current error correction cell is not to perform processing; Based on the error correction method of each error correction cell in the error correction path, the error correction result of the text to be corrected is determined.

7. The text error correction method according to claim 6, characterized in that: The step of determining the error correction result of the text to be corrected based on the error correction mode of each error correction cell in the error correction path includes: Based on the error correction cell corresponding to the error correction method, determining the error correction position corresponding to the error correction method in the text to be corrected; Based on the error correction position, determining the error correction character of the text to be corrected in the text to be corrected; Based on the error correction method, error correction position and error correction character, the error correction result of the text to be corrected is determined.

8. A text error correction device, characterized in that: The text error correction device comprises: A first acquisition module is used to acquire the text to be corrected; A text error correction module, used for inputting the text to be corrected into the preset large-scale language model for text error correction processing to obtain a corrected text; A first determination module is used to determine an edit distance matrix between the text to be corrected and the corrected text based on the text to be corrected and the corrected text, wherein the edit distance matrix includes the number of edit operations between each character in the text to be corrected and each character in the corrected text; The second determination module is used to determine the error correction result of the text to be corrected based on the editing operation number in the editing distance matrix.

9. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the text error correction method according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the text error correction method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Error correction method and device for voice recognition text

    CN104464736A

  • Text processing method and device, computer equipment and storage medium

    CN114386406A

  • Training method of voice transfer text error correction model and computer equipment

    CN115293139A

  • Text wrongly written character detection method and device

    CN115759076A

  • Text error correction method and device, electronic equipment and storage medium

    CN117709335A

Cited By

  • Student problem solving state evaluation method based on knowledge graph

    CN122114402A