Text error correction processing method and device, electronic equipment, medium and program product

This text correction method, which combines multiple error correction models with error correction performance data for hierarchical screening, solves the problems of inaccurate error correction and false correction in existing technologies, and achieves higher recall and precision, especially with stable error correction performance in complex semantics and novel error scenarios.

CN121413612APending Publication Date: 2026-01-27TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511471000.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-15
Publication Date
2026-01-27

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately identify erroneous characters and effectively avoid miscorrection in text correction processing, resulting in low recall and precision, especially under complex semantics and novel error scenarios.

Method used

Multiple error correction models are used to perform initial error correction on the text to be corrected. The error correction performance data is combined to perform hierarchical filtering to obtain target error correction characters, error correction characters to be restored, or error correction characters to be verified. Error correction processing is carried out based on the hierarchical filtering results to ensure the accuracy and stability of error correction.

Benefits of technology

It improves the recall and precision of text correction, effectively avoids false corrections, and enhances the accuracy and rationality of correction, especially showing greater stability in complex semantics and novel error scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121413612A_ABST
    Figure CN121413612A_ABST
Patent Text Reader

Abstract

The invention discloses a text error correction processing method and device, electronic equipment, a medium and a program product. The method comprises the steps that an initial error correction result of at least one error correction model for at least one error correction position in a to-be-corrected text is acquired; determining first error correction performance data corresponding to the same initial error correction character at the same error correction position in the initial error correction result; based on the first error correction performance data, grading and screening at least one initial error correction character in the initial error correction result to obtain a grading and screening result, the grading and screening result comprising at least any one of a screened error correction character or an error correction character to be restored, the screened error correction character comprises at least any one of a target error correction character or an error correction character to be verified; and performing error correction processing on the to-be-corrected text based on a classification screening result to obtain a target text after error correction. By utilizing the technical scheme provided by the invention, error correction can be effectively avoided while error characters are accurately identified, and the recall rate and the accuracy rate of text error correction are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a text error correction processing method, apparatus, electronic device, medium, and program product. Background Technology

[0002] In the era of self-media, article quality has become increasingly difficult to control. Due to the complexity of language, the diversity of expression, the flexibility of grammar, dialects, and internet slang, issues such as typos, missing words, and extra words are quite prominent. Therefore, how to accurately and effectively correct text errors has become an urgent problem to be solved. Summary of the Invention

[0003] This application provides a text correction processing method, apparatus, device, storage medium, and computer program product, which can accurately identify erroneous characters while effectively avoiding miscorrection, greatly improving the recall and precision of text correction, and thus enhancing the accuracy and rationality of text correction.

[0004] On the one hand, this application provides a text error correction processing method, the method comprising: Obtain initial error correction results from at least one error correction model for at least one error correction position in the text to be corrected; the initial error correction results include at least one initial error correction character, and each initial error correction character corresponds to one of the error correction positions; Determine the first error correction performance data corresponding to the same initial error correction character at the same error correction position in the initial error correction results; Based on the first error correction performance data, at least one initial error correction character in the initial error correction result is subjected to hierarchical filtering to obtain hierarchical filtering results. The hierarchical filtering results include at least one of the filtered error correction characters or the error correction characters to be restored. The filtered error correction characters include at least one of the target error correction characters or the error correction characters to be verified. Based on the hierarchical screening results, the text to be corrected is processed to obtain the corrected target text.

[0005] On the other hand, this application provides a text correction processing apparatus, the apparatus comprising: The initial error correction result acquisition module is configured to acquire the initial error correction result of at least one error correction model for at least one error correction position in the text to be corrected; the initial error correction result includes at least one initial error correction character, and any initial error correction character corresponds to one error correction position; The error correction performance data determination module is configured to determine the first error correction performance data corresponding to the same initial error correction character at the same error correction position in the initial error correction result. The hierarchical filtering module is configured to perform hierarchical filtering on at least one initial error correction character in the initial error correction result based on the first error correction performance data, and obtain hierarchical filtering results. The hierarchical filtering results include at least one of the filtered error correction characters or error correction characters to be restored. The filtered error correction characters include at least one of the target error correction character or error correction characters to be verified. The error correction processing module is configured to perform error correction processing on the text to be corrected based on the hierarchical filtering results, so as to obtain the corrected target text.

[0006] In an optional embodiment, the error correction processing module includes: The first error correction processing unit is configured to perform error correction processing on the text to be corrected based on the filtered error correction characters when the hierarchical filtering result includes filtered error correction characters and the filtered error correction characters include the error correction characters to be verified, so as to obtain the text to be verified. The error correction and verification unit is configured to input the text to be verified and the preset error correction instruction information into the large language model for error correction and verification, so as to obtain the target text.

[0007] In an optional embodiment, the error correction verification unit is further configured to perform error correction verification by inputting the text to be corrected, the text to be verified, and the preset error correction indication information into a large language model to obtain the target text.

[0008] In an optional embodiment, the apparatus further includes: The preset text pair set acquisition module is configured to acquire a preset text pair set, which includes a first text pair set and a second text pair set; each preset text pair in the first text pair set includes two correct texts; each preset text pair in the second text pair set includes one incorrect text and the correct text corresponding to the incorrect text. The similarity determination module is configured to determine the similarity between the current text pair and each preset text pair in the preset text pair set, wherein the current text pair is a text pair consisting of the text to be corrected and the text to be verified. The similar text pair determination module is configured to determine similar text pairs of the current text pair from the preset text pair set based on the similarity. The error correction and verification unit is further configured to perform error correction and verification by inputting the similar text pairs, the current text pairs, and the preset error correction indication information into the large language model to obtain the target text.

[0009] In an optional embodiment, the error correction processing module includes: The second error correction processing unit is configured to perform error correction processing on the text to be corrected based on the target error correction character when the hierarchical filtering result includes the filtered error correction character, the filtered error correction character includes the target error correction character, and the filtered error correction character does not include the error correction character to be verified, thereby obtaining the target text.

[0010] In an optional embodiment, when the at least one error correction model includes multiple error correction models, any one of the at least one error correction position corresponds to one or at least two initial error correction characters; the apparatus further includes: The candidate error correction character determination module is configured to determine the candidate error correction character for each of the at least one error correction positions based on the first error correction performance data corresponding to the initial error correction character of each of the at least one error correction positions. The hierarchical filtering module is further configured to perform hierarchical filtering on the candidate error correction characters corresponding to the at least one error correction position based on the first error correction performance data corresponding to the candidate error correction characters at each error correction position, and obtain the hierarchical filtering result.

[0011] In an optional embodiment, the hierarchical screening module includes at least one of the following units: The first character determination unit is configured to select the candidate error correction characters whose first error correction performance data is less than or equal to a first preset threshold and greater than or equal to a second preset threshold from the candidate error correction characters corresponding to the at least one error correction position as the error correction characters to be verified. The second character determination unit is configured to execute the selection of candidate error correction words whose first error correction performance data is greater than the first preset threshold from the candidate error correction characters corresponding to the at least one error correction position as the target error correction character. The third character determination unit is configured to select the candidate error correction character whose first error correction performance data is less than the second preset threshold from the candidate error correction characters corresponding to the at least one error correction position as the error correction character to be restored.

[0012] In an optional embodiment, the apparatus further includes: Based on the candidate error correction characters at each error correction position, the text to be corrected is processed to obtain the corrected initial text. The error correction processing module includes: The third error correction processing unit is configured to perform error correction processing on the initial text based on the hierarchical filtering results to obtain the target text.

[0013] In an optional embodiment, the third error correction processing unit includes: The fourth error correction processing unit is configured to use the initial text as the target text when the hierarchical filtering result includes the post-filtering error correction character and the post-filtering error correction character includes characters that do not include the error correction character to be verified. or, The fifth error correction processing unit is configured to, when the hierarchical filtering result includes the filtered error correction character and the filtered error correction character includes the error correction character to be verified, input the initial text and preset error correction indication information into the large language model for error correction verification to obtain the target text.

[0014] In an optional embodiment, the third error correction processing unit further includes: The restoration processing unit is configured to perform restoration processing on the initial text based on the character to be restored when the hierarchical filtering result includes the character to be restored and corrected, so as to obtain the restored text. The fourth error correction processing unit is also configured to use the restored text as the target text. The fifth error correction processing unit is further configured to input the restored text and preset error correction instruction information into a large language model for error correction verification to obtain the target text.

[0015] In an optional embodiment, the error correction performance data determination module includes: The first error correction performance data acquisition unit is configured to acquire model error correction performance data of each error correction model in the at least one error correction model and second error correction performance data of each error correction model for each corresponding error correction position. The first error correction performance data determination unit is configured to perform the following: determine the third error correction performance data of each error correction model for the corresponding error correction position based on the model error correction performance data of each error correction model and the second error correction performance data of each error correction model for the corresponding error correction position. The second error correction performance data unit is configured to execute at least one third error correction performance data based on the same initial error correction character corresponding to the same error correction position in the initial error correction result, and determine the first error correction performance data corresponding to the same initial error correction result at the same error correction position.

[0016] In an optional embodiment, when the at least one error correction model is multiple error correction models, the error correction performance data determination module includes: The second error correction performance data acquisition unit is configured to acquire the model error correction performance data of each of the plurality of error correction models; The target error correction model determination unit is configured to execute the determination of the target error correction model corresponding to the same initial error correction character at each error correction position from the plurality of error correction models; The third error correction performance data determination unit is configured to execute the model error correction performance data of the target error correction model corresponding to the same initial error correction character at each error correction position, and determine the first error correction performance data corresponding to the same initial error correction character at each error correction position.

[0017] In an optional embodiment, the hierarchical screening module includes at least one of the following units: The fourth character determination unit is configured to execute the initial error correction character whose first error correction performance data is less than or equal to a first preset threshold and greater than or equal to a second preset threshold among the at least one initial error correction character as the error correction character to be verified. The fifth character determination unit is configured to execute the initial error correction character whose first error correction performance data is greater than the first preset threshold among the at least one initial error correction character as the target error correction character; The sixth character determination unit is configured to select the initial error correction character whose first error correction performance data is less than the second preset threshold from the at least one initial error correction character as the error correction character to be restored.

[0018] On the other hand, an electronic device is provided, including: a processor; Memory used to store the processor's executable instructions; The processor is configured to execute the instructions to implement any of the text correction processing methods described above.

[0019] On the other hand, a computer-readable storage medium is provided, which, when the instructions in the storage medium are executed by the processor of an electronic device, enables the electronic device to perform any of the text correction processing methods described above.

[0020] On the other hand, a computer program product or computer program is provided, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform any of the text correction processing methods provided in the various alternative implementations described above.

[0021] The text correction processing method, apparatus, device, storage medium, and computer program product provided in this application have the following technical effects: This application, by combining at least one error correction model to correct the text to be corrected, and obtaining the initial error correction result of at least one error correction model for at least one error correction position in the text to be corrected, determines the first error correction performance data corresponding to the same initial error correction character at the same error correction position in the initial error correction result. Based on the first error correction performance data, it performs hierarchical filtering on at least one initial error correction character in the initial error correction result, obtaining a hierarchical filtering result including at least one of the target error correction character, the error correction character to be restored, or the error correction character to be verified. This enables refined filtering of the initial error correction result. Based on the hierarchical filtering result, the text to be corrected is processed to obtain the target error correction text corresponding to the text to be corrected. This can accurately identify erroneous characters while effectively avoiding false corrections, greatly improving the recall and precision of text correction, and thus improving the accuracy and rationality of text correction. Attached Figure Description

[0022] To more clearly illustrate the technical solutions and advantages in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0023] Figure 1 This is a schematic diagram of an application environment provided in an embodiment of this application; Figure 2 This is a flowchart illustrating a text correction processing method provided in an embodiment of this application; Figure 3 This is a schematic diagram of an application interface for text correction processing provided in an embodiment of this application; Figure 4 This is a schematic diagram of a text correction process provided in an embodiment of this application; Figure 5 This is a schematic diagram of the structure of a text correction processing device provided in an embodiment of this application; Figure 6 This is a block diagram of an electronic device for text error correction processing provided in an embodiment of this application; Figure 7 This is a block diagram of another electronic device for text error correction processing provided in the embodiments of this application. Detailed Implementation

[0024] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0025] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar concepts and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or server that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices.

[0026] In the embodiments of this application, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.

[0027] Please see Figure 1 , Figure 1 This is a schematic diagram of an application environment provided in an embodiment of this application. The application environment may include at least a terminal device 101 and a server 102.

[0028] The terminal device 101 and the server 102 are connected via a wireless or wired network. The terminal device 101 can be used to provide text correction services to users; the server 102 can provide background support to the terminal device 101 based on the text correction processing method provided in the embodiments of this application.

[0029] Optionally, the terminal device 101 may include, but is not limited to, electronic devices such as smartphones, desktop computers, tablets, laptops, smart speakers, digital assistants, augmented reality (AR) / virtual reality (VR) devices, smart wearable devices, in-vehicle terminals, and smart TVs; it may also be software running on the aforementioned electronic devices, such as applications or mini-programs. The operating system running on the electronic device in this embodiment may include, but is not limited to, Android, iOS, Linux, and Windows systems.

[0030] Optionally, the server 102 can be an independent physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server that provides cloud computing services.

[0031] In addition, it should be noted that, Figure 1 The example shown is merely one application environment provided by this disclosure, and the embodiments of this application are not limited thereto.

[0032] Those skilled in the art should understand that the terminal device 101 and server 102 described above are merely illustrative examples. Other existing or future terminal devices or servers that are applicable to this application should also be included within the scope of protection of this application, and are hereby incorporated by reference.

[0033] In related technologies, the following methods are mainly used for text correction: First, rule-driven text correction. This method scans text according to preset rules (such as homophone / speech replacement, collocation, grammatical structure, etc.) through string matching or pattern recognition, and directly replaces or provides candidates after finding errors. However, this method has high accuracy only for errors within the rule coverage area, and low recall, making it difficult to handle complex semantics and novel errors. Second, statistical model-based text correction. This method uses statistical features (such as co-occurrence probability) obtained from training on large-scale corpora to judge the rationality of text and find errors. However, this method is limited by the size and quality of the corpus, and its ability to correct low-frequency words, new words, and complex semantics is weak, resulting in low recall and precision. Third, traditional deep learning-based text correction. This method is based on large-scale labeled parallel corpora, training an end-to-end neural network to automatically learn the mapping between errors and correct text, detecting and correcting errors. However, the model's performance depends on high-quality big data, and its performance is often unstable for rare or novel errors, making it impossible to guarantee the recall and precision of text correction.

[0034] In this embodiment, at least one error correction model is first used to correct the text to be corrected, resulting in initial error correction results for at least one error correction position in the text. This approach effectively handles complex semantics and novel errors without being limited by the size and quality of the corpus, effectively ensuring error correction capabilities for low-frequency words, new words, and complex semantics. Next, first error correction performance data is determined for the same initial error correction character at the same error correction position in the initial error correction results. Based on this first error correction performance data, at least one initial error correction character in the initial error correction results is hierarchically filtered to obtain a result including the target error correction character and the error correction to be restored. The hierarchical filtering results, including at least one of the characters or the characters to be verified and corrected, can achieve fine-grained filtering of the initial error correction results of the error correction model. Based on the hierarchical filtering results, the text to be corrected is processed to obtain the target error correction text corresponding to the text to be corrected. Further text error correction processing can be carried out through the hierarchical filtering results, including at least one of the target error correction characters, the error correction characters to be restored, or the error correction characters to be verified. This effectively ensures the stability of error correction for rare or novel errors. While accurately identifying erroneous characters, it effectively avoids false corrections, greatly improves the recall and precision of text error correction, and thus improves the accuracy and rationality of text error correction.

[0035] The following describes a text error correction method proposed in this application. Figure 2 This is a flowchart illustrating a text correction processing method provided in an embodiment of this application. This method can be executed by a computer device, which can be a terminal device, a server, or an interaction between the terminal device and the server. Figure 2 As shown, the method may include: S201: Obtain the initial error correction result of at least one error correction model for at least one error correction position in the text to be corrected.

[0036] In one specific embodiment, the text to be corrected can be text that requires error correction, such as a single sentence or a paragraph in an article. Optionally, the text to be corrected can be generated from real-time editing or pre-generated; specifically, the at least one error correction model can be one or more error correction models (i.e., at least two error correction models); specifically, the error correction model can be an artificial intelligence model for text correction, such as an LLM (Large Language Model) or NLP (Natural Language Processing) model. When at least one error correction model is multiple error correction models, the multiple error correction models can correspond to the same artificial intelligence model or different artificial intelligence models.

[0037] In a specific embodiment, the initial error correction result may include at least one initial error correction character, and each initial error correction character corresponds to an error correction position. Specifically, the text to be corrected may include multiple original characters; optionally, the original character may be the smallest text unit, or it may be a word (the smallest semantic unit when an artificial intelligence model processes text), and each of the multiple original characters corresponds to a character position in the text to be corrected (i.e., indicating the position of the corresponding original character in the text to be corrected, such as the 5th character in the text to be corrected); the error correction position may be the character position where the error-correcting model identifies the character (original character) with an error during the error correction process; any initial error correction character may be the character after the error-correcting model corrects the character with an error during the error correction process (i.e., the corrected character). Optionally, when at least one error correction model is multiple error correction models, the characters with errors identified by different error correction models during the error correction process of the text to be corrected may be the same or at least partially different; correspondingly, the corresponding corrected characters (initial error correction characters) may also be the same or at least partially different.

[0038] In a specific embodiment, obtaining the initial error correction result of at least one error correction model for at least one error correction position in the text to be corrected may include inputting the text to be corrected into at least one error correction model for error correction processing, thereby obtaining an error correction set for each error correction model for the text to be corrected. Optionally, the error correction set for any error correction model for the text to be corrected may include: at least one original character (a character in the text to be corrected that is identified as having an error by the error correction model, i.e., the character to be corrected), the error correction position of each original character (i.e., the character position of each original character in the text to be corrected), and the initial error correction character at each error correction position of the original character (i.e., the corrected character for each original character recalled by the error correction model). Optionally, if no character with an error is identified during the error correction processing of the text to be corrected by any error correction model, the corresponding error correction set may be empty.

[0039] S203: Determine the first error correction performance data corresponding to the same initial error correction character at the same error correction position in the initial error correction result.

[0040] In a specific embodiment, the first error correction performance data corresponding to the same initial error correction character at the same error correction position can characterize the reasonableness of correcting the original character at the error correction position to the initial error correction character, that is, the probability that the original character at the error correction position in the text to be corrected is incorrect and the correct character is the initial error correction character.

[0041] In an optional embodiment, when at least one error correction model is multiple error correction models, the first error correction performance data corresponding to the same initial error correction character at the same error correction position in the initial error correction result may include: Obtain the model error correction performance data for each of the multiple error correction models; Determine the target error correction model corresponding to the same initial error correction character at each error correction position from multiple error correction models; Based on the model error correction performance data of the target error correction model corresponding to the same initial error correction character at each error correction position, the first error correction performance data corresponding to the same initial error correction character at each error correction position is determined.

[0042] In one specific embodiment, the model error correction performance data for each error correction model can be data characterizing the error correction performance (accuracy and recall) of that model; for example, the model error correction performance data can be the F1 score, i.e., the harmonic mean of precision and recall. In one specific embodiment, the F1 score of each error correction model can be determined by combining test text pairs; specifically, the test text pairs can include multiple erroneous samples and multiple correct samples; any erroneous sample can include an erroneous text (i.e., text containing at least one erroneous character) and the correct text of that erroneous text. Correspondingly, the erroneous text in the erroneous samples can be input into the error correction model for error correction processing, obtaining the correct text recalled by the error correction model for that erroneous text. The recall rate (the proportion of all real-world targets (such as errors in text) successfully identified by the error correction model) is then calculated by comparing the recalled correct text with the corresponding correct text in the erroneous samples (i.e., the recall result of the error correction model for the erroneous samples). For example, recall rate = TP / (TP + FN), where TP represents the number of erroneous samples correctly identified and labeled as "erroneous" by the error correction model, and FN represents the number of erroneous samples missed and not labeled as "erroneous" by the error correction model. Furthermore, in the precision calculation process, in addition to combining the recall results of the error correction model for erroneous samples, it is also necessary to combine the recall results of the error correction model for correct samples. Specifically, any correct sample can include a correct text and that correct text. Correspondingly, the correct text is input into the error correction model for error correction processing, obtaining the correct text recalled by the error correction model for the correct text. The precision is then calculated by combining the comparison between the recalled correct text and the corresponding correct text in the correct samples (the recall result of the error correction model for the correct samples) and the recall result of the error correction model for the erroneous samples. Specifically, precision = TP / (TP + FP), where TP represents the number of "true errors" correctly identified by the error correction model in the erroneous samples; FP represents the number of "correct content" samples that the error correction model misclassified as errors in the correct samples. Specifically, the F1 score can be calculated by combining precision and recall using the following formula: F1 = 2 × (Precision × Recall) / (Precision + Recall).

[0043] In a specific embodiment, during the process of multiple error correction models correcting text to be corrected, more than one error correction model may recall the same initial error correction character for the same error correction position. Accordingly, the target error correction model corresponding to the same initial error correction character at each error correction position can be determined from the multiple error correction models. Specifically, the target error correction model corresponding to any initial error correction character at any error correction position can be the error correction model whose initial error correction character is the corrected character recalled after error correction processing at that error correction position. Specifically, determining the first error correction performance data corresponding to the same initial error correction character at each error correction position based on the model error correction performance data of the target error correction model corresponding to the same initial error correction character at each error correction position can include adding the model error correction performance data of the target error correction model corresponding to the same initial error correction character at each error correction position to obtain the first error correction performance data corresponding to the initial error correction character at that error correction position.

[0044] In the above embodiments, by using multiple error correction models to perform error correction processing on the text to be corrected, the error correction performance data of the target error correction model corresponding to the same initial error correction character at each error correction position can be combined to determine the first error correction performance data corresponding to the same initial error correction character at each error correction position. This ensures the effectiveness of measuring the error correction performance corresponding to each initial error correction character at each error correction position, reduces the risk of misjudgment by a single model, and improves the error correction recall and precision.

[0045] In an optional embodiment, the first error correction performance data corresponding to the same initial error correction character at the same error correction position in the initial error correction result may include: Obtain model error correction performance data for each error correction model in at least one error correction model and second error correction performance data for each error correction model for each corresponding error correction position; Based on the model error correction performance data of each error correction model and the second error correction performance data of each error correction model for the corresponding error correction position, determine the third error correction performance data of each error correction model for the corresponding error correction position. Based on at least one third error correction performance data of the same initial error correction character corresponding to the same error correction position in the initial error correction results, determine the first error correction performance data corresponding to the same initial error correction result at the same error correction position.

[0046] In a specific embodiment, the detailed model error correction performance data for each error correction model can be found in the relevant descriptions above, and will not be repeated here.

[0047] In a specific embodiment, the second error correction performance data for each error correction model at each corresponding error correction position can characterize the confirmation probability of the error correction result (initial error correction character) of the error correction model at that error correction position. For example, the second error correction performance data can be determined by combining the PPL (Perplexity) of the error correction model at that error correction position in the text to be corrected—that is, the PPL of the corrected text obtained after replacing the original character at that error correction position in the text to be corrected with the initial error correction character of the error correction model at that error correction position. For example, the error correction model for any error correction position in the text to be corrected... The PPL (i.e., the PPL of the corrected text) can be determined by combining a pre-trained perplexity prediction model or by combining a preset PPL calculation formula. Specifically, since a lower PPL indicates higher sentence (text) quality and more fluent grammar, PPL is negatively correlated (inversely proportional) with the second error correction performance data. A preset inverse proportional function can be used to convert PPL into the corresponding second error correction performance data. Specifically, the preset inverse proportional function can be set according to the actual application. To avoid excessive weight due to a small PPL, we use a smoother exponential form. For example, the preset inverse proportional function is as follows:

[0048] in, This represents the second error correction performance data of the i-th error correction model for the k-th original character (error correction position). This represents the PPL of the i-th error correction model for the k-th original character (error correction position). The coefficients are preset, and n represents the total number of error correction models; This indicates the number of original characters in the text to be corrected.

[0049] In one specific embodiment, during the error correction process of any error correction model on the text to be corrected, the confidence level corresponding to each error correction position can also be output. The confidence level corresponding to each error correction position represents the confirmation probability of the error correction result (initial error correction character) of the error correction model at that error correction position. Optionally, the confidence level corresponding to each error correction position can be used as the second error correction performance data of the error correction model for that error correction position.

[0050] In a specific embodiment, determining the third error correction performance data of each error correction model for the corresponding error correction position based on the model error correction performance data of each error correction model and the second error correction performance data of each error correction model for the corresponding error correction position may include multiplying the model error correction performance data of each error correction model by the second error correction performance data of each error correction model for the corresponding error correction position to obtain the third error correction performance data of each error correction model for the corresponding error correction position; that is, the third error correction performance data of each error correction model for the corresponding error correction position may be the product of the model error correction performance data of the error correction model and the second error correction performance data of the error correction model for the corresponding error correction position.

[0051] In a specific embodiment, in a scenario with multiple error correction models, different error correction models may recall the same initial error correction character for the same error correction position. Accordingly, the same initial error correction character corresponding to the same error correction position may correspond to at least one third error correction performance data (the third error correction performance data corresponding to each error correction model that recalls the initial error correction character at the error correction position). Accordingly, the above-mentioned determination of the first error correction performance data corresponding to the same initial error correction result at the same error correction position based on at least one third error correction performance data of the same initial error correction character corresponding to the same error correction position in the initial error correction result may include: taking the sum of at least one third error correction performance data of the same initial error correction character corresponding to the same error correction position as the first error correction performance data corresponding to the initial error correction result at the error correction position.

[0052] In the above embodiments, by combining the model error correction performance data of each error correction model and the second error correction performance data of each error correction model for the corresponding error correction position, the third error correction performance data of each error correction model for the corresponding error correction position is determined. Then, by combining at least one third error correction performance data of the same initial error correction character corresponding to the same error correction position in the initial error correction result, the first error correction performance data corresponding to the same initial error correction result at the same error correction position is determined. Based on the global error correction performance data of the error correction model itself (model error correction performance data) and the error correction performance data of the initial error correction character itself (at least one third error correction performance data), a more reliable error correction model and a higher quality initial error correction character gain greater influence, thereby improving the precision and recall of error correction processing.

[0053] S205: Based on the first error correction performance data, at least one initial error correction character in the initial error correction result is subjected to hierarchical filtering to obtain hierarchical filtering results.

[0054] In a specific embodiment, the above-mentioned hierarchical filtering result includes at least one of the filtered correction character or the correction character to be restored. That is, the hierarchical filtering result may include the filtered correction character, or the hierarchical filtering result may include the correction character to be restored, or the hierarchical filtering result may include both the correction character to be restored and the filtered correction character. Specifically, the above-mentioned filtered correction character may include at least one of the target correction character or the correction character to be verified. That is, the filtered correction character may include the target correction character, or the filtered correction character may include the correction character to be verified; or the filtered correction character may include both the target correction character and the correction character to be verified. Specifically, the target correction character may be an acceptable correction character (i.e., determining that the original character at the corresponding correction position in the text to be corrected will be modified to the target correction character); the correction character to be verified may be a correction character that needs further verification to determine whether it is acceptable; the correction character to be restored may be a correction character that is not acceptable.

[0055] In an optional embodiment, when at least one error correction model is used, each error correction position corresponds to an initial error correction character; correspondingly, the above-mentioned hierarchical filtering of at least one initial error correction character in the initial error correction result based on the first error correction performance data to obtain the hierarchical filtering result may include at least one of the following steps: The initial error correction character whose first error correction performance data is less than or equal to a first preset threshold and greater than or equal to a second preset threshold is selected as the error correction character to be verified. The initial error correction character whose first error correction performance data is greater than the first preset threshold is taken as the target error correction character. The initial error correction character whose first error correction performance data is less than the second preset threshold is selected as the error correction character to be restored.

[0056] In a specific embodiment, the first preset threshold and the second preset threshold can be set according to actual application requirements, and the first preset threshold is greater than the second preset threshold. For example, the first preset threshold is 0.6 and the second preset threshold is 0.35. Specifically, if there is an initial error correction character among at least one initial error correction character whose corresponding first error correction performance data is less than or equal to the first preset threshold and greater than or equal to the second preset threshold, the initial error correction character whose corresponding first error correction performance data is less than or equal to the first preset threshold and greater than or equal to the second preset threshold can be used as the error correction character to be verified; if there is an initial error correction character among at least one initial error correction character whose corresponding first error correction performance data is greater than the first preset threshold, the initial error correction character whose corresponding first error correction performance data is greater than the first preset threshold can be used as the target error correction character; if there is an initial error correction character among at least one initial error correction character whose corresponding first error correction performance data is less than the second preset threshold, the initial error correction character whose corresponding first error correction performance data is less than the second preset threshold can be used as the error correction character to be restored.

[0057] In the above embodiments, by setting a first preset threshold and a second preset threshold, at least one initial error correction character is subjected to hierarchical screening, which can achieve fine-grained screening of the initial error correction results. Initial error correction characters with high error correction performance data (high confidence) are directly adopted; initial error correction characters with low error correction performance data (low confidence) are not corrected to avoid false corrections; and initial error correction characters with medium confidence are used as error correction characters to be verified for secondary verification. This can effectively avoid false corrections while accurately identifying erroneous characters, greatly improving the recall and accuracy of text error correction.

[0058] In an optional embodiment, when at least one error correction model includes multiple error correction models, any one of the at least one error correction position may correspond to one or at least two initial error correction characters; correspondingly, the above method may further include: Based on the first error correction performance data corresponding to the initial error correction character of each error correction position in at least one error correction position, determine the candidate error correction character for each error correction position in at least one error correction position. Accordingly, based on the first error correction performance data, the above-mentioned hierarchical filtering of at least one initial error correction character in the initial error correction result can include: Based on the first error correction performance data corresponding to the candidate error correction character at each error correction position, the candidate error correction characters corresponding to at least one error correction position are sorted and filtered in a hierarchical manner to obtain the hierarchical filtering results.

[0059] In a specific embodiment, the initial error correction character with the largest first error correction performance data among the initial error correction characters at each error correction position can be used as the candidate error correction character for that error correction position. In an optional embodiment, the process of hierarchically filtering candidate error-correcting characters corresponding to at least one error-correcting position based on the first error-correcting performance data corresponding to the candidate error-correcting characters at each error-correcting position to obtain the hierarchical filtering result may include at least one of the following steps: Candidate error correction characters whose first error correction performance data is less than or equal to a first preset threshold and greater than or equal to a second preset threshold are selected as error correction characters to be verified. The candidate error correction word whose first error correction performance data is greater than a first preset threshold is selected as the target error correction character from the candidate error correction characters corresponding to at least one error correction position. Candidate error correction characters whose first error correction performance data is less than a second preset threshold are selected as error correction characters to be restored.

[0060] In a specific embodiment, if there is a candidate error correction character whose first error correction performance data is less than or equal to a first preset threshold and greater than or equal to a second preset threshold among the candidate error correction characters corresponding to at least one error correction position, the candidate error correction character whose first error correction performance data is less than or equal to the first preset threshold and greater than or equal to the second preset threshold can be used as the error correction character to be verified; if there is a candidate error correction character whose first error correction performance data is greater than the first preset threshold among the candidate error correction characters corresponding to at least one error correction position, the candidate error correction character whose first error correction performance data is greater than the first preset threshold can be used as the target error correction character; if there is a candidate error correction character whose first error correction performance data is less than the second preset threshold among the candidate error correction characters corresponding to at least one error correction position, the candidate error correction character whose first error correction performance data is less than the second preset threshold can be used as the error correction character to be restored.

[0061] In the above embodiments, in scenarios where multiple error correction models are used for error correction, one error correction position may correspond to one or more initial error correction characters. Therefore, by combining the first error correction performance data corresponding to the initial error correction character at each error correction position, a candidate error correction character can be selected for each error correction position, which can better ensure the accuracy of subsequent error correction results. Furthermore, by setting a first preset threshold and a second preset threshold, the candidate error correction characters corresponding to at least one error correction position can be hierarchically filtered, which can achieve fine-grained filtering of candidate error correction characters. Candidate error correction characters with high error correction performance data (high confidence) are directly adopted; candidate error correction characters with low error correction performance data (low confidence) are not corrected to avoid false corrections; and candidate error correction characters with medium confidence are used as error correction characters to be verified for secondary verification. This can effectively avoid false corrections while accurately identifying erroneous characters, greatly improving the recall and accuracy of text error correction.

[0062] S207: Based on the hierarchical screening results, perform error correction processing on the text to be corrected to obtain the corrected target text.

[0063] In one specific embodiment, the target text can be the text after error correction processing (corrected text) of the text to be corrected. In an optional embodiment, the above-mentioned error correction processing of the text to be corrected based on the hierarchical screening results to obtain the corrected target text may include: If the hierarchical filtering results include the filtered correction characters, and the filtered correction characters include the target correction characters, but the filtered correction characters do not include the correction characters to be verified, then the text to be corrected is corrected based on the target correction characters to obtain the target text.

[0064] In one specific embodiment, based on the target error correction character, the text to be corrected is processed to obtain the target text. This may include replacing the original character at the error correction position (character position) corresponding to the target error correction character in the text to be corrected with the target error correction character to obtain the target text.

[0065] In the above embodiments, when the filtered correction characters include the target correction characters but do not include the correction characters to be verified, the text to be corrected is directly combined with the target correction characters to obtain the corrected target text, which can effectively ensure the accuracy and recall of text correction processing.

[0066] In an optional embodiment, the above-mentioned error correction processing of the text to be corrected based on the hierarchical screening results, to obtain the corrected target text, may include: If the hierarchical filtering results include post-filtered error correction characters, and the post-filtered error correction characters include error correction characters to be verified, then based on the post-filtered error correction characters, the text to be corrected is processed to obtain the text to be verified. The text to be verified and the preset error correction instructions are input into the large language model for error correction and verification to obtain the target text.

[0067] In a specific embodiment, when the hierarchical filtering result includes post-filtering correction characters and the post-filtering correction characters include correction characters to be verified, the post-filtering correction characters may also include target correction characters; optionally, they may not include target correction characters; specifically, based on the post-filtering correction characters, performing error correction processing on the text to be corrected to obtain the text to be verified may include replacing the original character at the correction position (character position) corresponding to the post-filtering correction characters in the text to be corrected with the post-filtering correction characters to obtain the aforementioned text to be verified.

[0068] In one specific embodiment, the preset error correction instruction information can be information used to instruct the large language model to perform error correction and verification on the input text. For example, please identify whether the input text is correct and give the correct result.

[0069] In an optional embodiment, the above-mentioned inputting the text to be verified and the preset error correction indication information into the large language model for error correction verification, to obtain the target text, may include: The text to be corrected, the text to be verified, and the preset error correction instructions are input into the large language model for error correction and verification to obtain the target text.

[0070] In one specific embodiment, when the text to be corrected is input into the large language model together, the preset error correction indication information can also be used to indicate the relationship between the input texts, such as which is the original text to be corrected and which is the corrected text (the text to be verified). For example, please identify whether the input original text S (the identifier of the text to be corrected) and the corrected text S' (the identifier of the text to be verified) are correct and give the correct result.

[0071] In an optional embodiment, if the hierarchical filtering results include the filtered correction characters, and the filtered correction characters include the correction characters to be verified but do not include the target correction characters, and the large language model analyzes that the text to be corrected is also correct, the text to be corrected can be directly used as the target text.

[0072] In the above embodiments, the text to be corrected and the text to be verified are simultaneously sent to a large language model for secondary verification, so that the differences and correlations between the text before and after correction are fully utilized, which can better improve the recall and precision of text correction.

[0073] In one specific embodiment, the above method may further include: Get the preset set of text pairs; Determine the similarity between the current text pair and each preset text pair in the preset text pair set. The current text pair is a text pair consisting of the text to be corrected and the text to be verified. Based on similarity, similar text pairs of the current text pair are determined from a preset set of text pairs; Accordingly, the above-mentioned input of the text to be corrected, the text to be verified, and the preset error correction instructions into the large language model for error correction and verification, the resulting target text may include: The similar text pairs, the current text pair, and the preset error correction instructions are input into the large language model for error correction and verification to obtain the target text.

[0074] In a specific embodiment, the aforementioned preset text pair set includes a first text pair set and a second text pair set; each preset text pair in the first text pair set includes two correct texts, specifically, the two correct texts in any preset text pair are two different expressions of the same semantics; each preset text pair in the second text pair set includes one incorrect text and the correct text corresponding to the incorrect text.

[0075] In one specific embodiment, the similarity between text pairs can be determined by extracting the feature vectors corresponding to the text pairs and combining the distance between the feature vectors, such as cosine distance, Euclidean distance, etc.; optionally, a preset similarity recognition model can also be used to recognize the similarity between text pairs.

[0076] In an optional embodiment, determining similar text pairs of the current text pair from a preset set of text pairs based on similarity may include: selecting text pairs in the preset set whose similarity is greater than or equal to a preset similarity threshold as similar text pairs of the current text pair; optionally, determining similar text pairs of the current text pair from a preset set of text pairs based on similarity may include: selecting text pairs in the preset set that are sorted in descending order of similarity and ranked at the top of a preset number of positions as similar text pairs of the current text pair. Specifically, the preset similarity threshold and the preset number can be set according to actual application requirements.

[0077] In a specific embodiment, suppose the text to be corrected, S, is "Covering the Sky and Obscuring the Sun! XXX (character name) delivered 4 blocks in 4 minutes of playing time"; and the text to be verified, S', is "Covering the Sky and Obscuring the Sun!". XXX (character name) delivered 4 blocks in 4 minutes of playing time; similar text pairs include {"s":"Want to play an alley-oop in front of me? Character name 1, Blood Hat, Character name 2","s'":"Want to play an alley-oop in front of me? Character name 1, Blood Hat, Character name 2"}; the preset error correction instruction is that you are a typo recognition expert, responsible for identifying whether the original sentence S and the alternative result S' are correct, and giving the correct result (if S and S' are both incorrect, please give the corrected result); input the above text to be corrected, the text to be verified, the similar text pairs, and the preset error correction instruction into the large language model for error correction processing. The large language model can output: {"S":"No typos","S'":"No typos","reason":"Blood Hat and Big Hat have similar meanings, both are applicable in this scenario, the result is S","res":"Covering the sky! XXX (character name) delivered 4 blocks in 4 minutes of playing time"}, correspondingly, the target text can be Covering the sky! XXX (character name) delivered 4 blocks in 4 minutes of playing time.

[0078] In a specific embodiment, assume the text to be corrected, S, is a person who practices dancing, and the image of a dashing young man is brought to life; the text to be verified, S', is also a person who practices dancing, and the image of a dashing young man is brought to life; similar text pairs include {"s":"The dashing young man is brought to life at this moment!! I was completely stunned when Xiao Yu appeared","s'":"The dashing young man is brought to life at this moment!! I was completely stunned when Xiao Yu appeared"}; the preset correction instruction information is that you are a typo recognition expert, responsible for identifying whether the original sentence S and the alternative result S' are correct. The system confirms the correctness and provides the correct result. The above text to be corrected, the text to be verified, similar text pairs, and the preset error correction instructions are input into the large language model for error correction processing. The large language model can output: [{"S":"There is a typo","S'":"No typo","reason":"The word 'concrete' in the original sentence S should be 'concrete'. 'Concrete' refers to abstract things becoming concrete and vivid. The context requires 'image' instead of 'like'. The word choice in the text to be verified S' is correct.","res":"XXX (person's name) is a dancer, a dashing young man in fine clothes and riding a spirited horse, which has been materialized."}

[0079] In an optional embodiment, during the error correction and verification process using a large language model, additional input can be made regarding the judgment rules for the text to be corrected and the text to be verified (rules indicating whether the text to be corrected and the text to be verified are correct), to better adapt to different usage scenarios. For example, when making a judgment, the logic of the entire sentence is considered comprehensively. If S does not affect readability, it is considered to be without errors. For instance, internet slang, colloquialisms, abbreviations, ambiguous words, and similar expressions are all considered to be without errors. However, the wording of S' needs to be logically consistent and conform to the context. Common nouns do not require correction and are considered to be without errors.

[0080] In the above embodiments, before combining with the large language model for verification, similar text pairs are matched from a preset set of text pairs and input together into the large language model. This allows the large language model to better learn the correlation between the text to be verified and the text to be corrected, thereby improving the verification effect and the overall error correction effect.

[0081] In addition, it should be noted that in scenarios where error correction verification is performed in conjunction with a large language model, the text to be verified may include the target error correction character. If the target error correction character is modified back to the original character during the large language model verification process, the original character can be restored to the target error correction character to better ensure the accuracy of the error correction result.

[0082] In the above embodiments, when the hierarchical filtering results include filtered error correction characters and the filtered error correction characters include error correction characters to be verified, the text to be corrected is first processed based on the filtered error correction characters to obtain the text to be verified. Then, the text to be verified and the preset error correction indication information are input into the large language model for error correction verification to obtain the target text. Based on the error correction results of the error correction model, the powerful semantic understanding and reasoning capabilities of the large language model can be further utilized to improve the accuracy and recall of text correction. At the same time, it can also avoid the problem of over-correction or change of the original meaning that is easily caused by directly combining the large language model for text correction in related technologies.

[0083] In an optional embodiment, the above method may further include: Based on the candidate error correction characters at each error correction position, the text to be corrected is processed to obtain the initial text after error correction. Accordingly, based on the hierarchical screening results, the text to be corrected is processed to obtain the corrected target text, which includes: Based on the hierarchical screening results, the initial text is corrected to obtain the target text.

[0084] In a specific embodiment, the above-mentioned error correction processing of the text to be corrected based on the candidate error correction characters at each error correction position, and the resulting corrected initial text may include replacing the original character at each error correction position in the text to be corrected with the candidate error correction character to obtain the initial text.

[0085] In an optional embodiment, the above-described error correction process for the initial text based on the hierarchical screening results to obtain the target text may include: If the tiered filtering results include post-filtered error correction characters, and the post-filtered error correction characters include or do not include the error correction characters to be verified, the initial text will be used as the target text. or, If the hierarchical filtering results include post-filtered error correction characters, and the post-filtered error correction characters include characters to be verified for error correction, the initial text and preset error correction indication information are input into the large language model for error correction verification to obtain the target text.

[0086] In an optional embodiment, the above-mentioned error correction processing of the initial text based on the hierarchical screening results to obtain the target text may further include: If the graded filtering results include characters to be restored and corrected, the initial text is restored based on the characters to be restored and corrected to obtain the restored text. Accordingly, using the initial text as the target text can include using the restored text as the target text. Accordingly, the above-mentioned inputting the initial text and preset error correction instructions into the large language model for error correction verification to obtain the target text may include: inputting the restored text and preset error correction instructions into the large language model for error correction verification to obtain the target text.

[0087] In one specific embodiment, inputting the restored text and preset error correction instruction information into a large language model for error correction verification to obtain the target text may include: inputting the text to be corrected, the restored text, and preset error correction instruction information into a large language model for error correction verification to obtain the target text; In a specific embodiment, the text to be corrected, the restored text, and the preset error correction instruction information are input into a large language model for error correction verification to obtain the target text, which may include the text to be corrected, the restored text, the corresponding similar text pairs (similar text pairs formed by the text to be corrected and the restored text), and the preset error correction instruction information.

[0088] In one specific embodiment, the restored text and preset error correction indication information are input into a large language model for error correction verification to obtain a detailed breakdown of the relevant refinement steps for the target text. This can be referred to in the above-described detailed breakdown of the relevant refinement steps for inputting the text to be verified and preset error correction indication information into a large language model for error correction verification to obtain the target text, which will not be repeated here.

[0089] In the above embodiments, in the scenario of text correction using multiple error correction models, the performance index data corresponding to the initial error correction characters in the initial error correction results of multiple error correction models for the text to be corrected are combined. After selecting a candidate error correction result for each error correction position, the text to be corrected is first processed in combination with the candidate error correction results. Then, after hierarchical filtering, the initial text after correction can be directly processed in combination with the hierarchical filtering results, which can better ensure the accuracy of subsequent error correction results. Furthermore, when the hierarchical filtering results include filtered error correction characters, and the filtered error correction characters include error correction characters to be verified, the initial text and preset error correction indication information are input into a large language model for error correction verification to obtain the target text. Utilizing the powerful semantic understanding and reasoning capabilities of the large language model, the accuracy and recall of text correction are further improved. And when the hierarchical filtering results include error correction characters to be restored, the error correction characters to be restored in the initial text are restored first, which can effectively avoid false corrections and further improve the accuracy of text correction.

[0090] In practical applications, the text correction processing method provided in this application can perform text correction processing in real-time text editing and input scenarios; it can also perform text correction processing on already edited text. In a specific embodiment, when creating an article (manually entering or copying articles from other sources), the terminal device can extract the text to be corrected and transmit it to the server's correction service for text correction processing, and return the correction result to the terminal device. The terminal device prompts the user for typos in the main text and sidebar, and the user can directly accept or ignore the typos in the main text or sidebar. Figure 3 As shown, Figure 3 This is a schematic diagram of an application interface for text error correction processing provided in an embodiment of this application; combined with Figure 3 As can be seen, the text correction processing method provided in this application embodiment can display the characters to be corrected in the currently edited text (text to be corrected) and the corrected characters in the target text when an error occurs; optionally, it can also display an "Accept" control corresponding to the corrected characters, and the user can use the "Accept" control to correct the characters to be corrected in the currently edited text to the correct characters; in the case of multiple errors, all characters to be corrected in the currently edited text can be corrected to the correct characters with one click.

[0091] In a specific embodiment, such as Figure 4 As shown, Figure 4 This is a schematic diagram of a text correction processing procedure provided in an embodiment of this application; specifically, in conjunction with... Figure 4 As can be seen, the input text to be corrected is fed into multiple error correction models (n error correction models) for error correction processing, resulting in corresponding recall results (at least one initial error correction character for at least one error correction position in the text to be corrected). Further, for recall results combined with the same error correction position, voting can be performed using the corresponding third error correction performance data. During the voting process, the third error correction performance data corresponding to the same initial error correction character at each error correction position can be added together to determine the first error correction performance data for the same initial error correction character at each error correction position. The largest first error correction performance data in the recall results corresponding to each error correction position is then used as the candidate error correction character for that position. Next, based on the first error correction performance data corresponding to the candidate error correction character at each error correction position, the candidate error correction characters corresponding to at least one error correction position are further analyzed. In the hierarchical screening process, candidate error-correcting words with corresponding first error-correcting performance data greater than a first preset threshold are designated as target error-correcting characters (i.e., misspelled words); candidate error-correcting characters with corresponding first error-correcting performance data less than a second preset threshold are designated as error-correcting characters to be restored (i.e., not misspelled words); and candidate error-correcting characters with corresponding first error-correcting performance data less than or equal to the first preset threshold and greater than or equal to the second preset threshold are designated as error-correcting characters to be verified. Then, the corresponding text to be verified can be determined by combining the characters to be verified, and this text can be combined with the text to be corrected to form a current text pair. The current text pair, similar text pairs matched from the preset text pair set, and preset error-correcting indication information are then input into the large language model for error correction verification, resulting in the target text as the final corrected text.

[0092] In a specific embodiment, as shown in Table 1, the differences in performance between the text correction method of this application and three existing error correction methods (BERT-based fine-tuning, mixed-element fine-tuning, and statistical model-based methods) on a news domain test set are illustrated:

[0093] Table 1 As can be seen from Table 1, the text correction method provided in this application embodiment has significant improvements in recall, precision, and F1 score.

[0094] As can be seen from the technical solutions provided in the embodiments of this specification above, this specification, when combining at least one error correction model to correct the text to be corrected, and obtaining the initial error correction result of at least one error correction model for at least one error correction position in the text to be corrected, determines the first error correction performance data corresponding to the same initial error correction character at the same error correction position in the initial error correction result, and combines the first error correction performance data to perform hierarchical filtering on at least one initial error correction character in the initial error correction result, obtaining a hierarchical filtering result including at least one of the target error correction character, the error correction character to be restored, or the error correction character to be verified. This can achieve fine filtering of the initial error correction result, and based on the hierarchical filtering result, the text to be corrected is processed to obtain the target error correction text corresponding to the text to be corrected. This can accurately identify erroneous characters while effectively avoiding false corrections, greatly improving the recall and precision of text correction, and thus improving the accuracy and rationality of text correction.

[0095] This application also provides a text correction processing device, such as... Figure 5 As shown, the above-mentioned device includes: The initial correction result acquisition module 510 is configured to acquire the initial correction result of at least one correction model for at least one correction position in the text to be corrected; the initial correction result includes at least one initial correction character, and each initial correction character corresponds to a correction position; The error correction performance data determination module 520 is configured to determine the first error correction performance data corresponding to the same initial error correction character at the same error correction position in the initial error correction result. The hierarchical filtering module 530 is configured to perform hierarchical filtering on at least one initial error correction character in the initial error correction result based on the first error correction performance data, and obtain hierarchical filtering results. The hierarchical filtering results include at least one of the filtered error correction characters or the error correction characters to be restored. The filtered error correction characters include at least one of the target error correction character or the error correction character to be verified. The error correction processing module 540 is configured to perform error correction processing on the text to be corrected based on the hierarchical filtering results, and obtain the corrected target text.

[0096] In an optional embodiment, the error correction processing module 540 includes: The first error correction processing unit is configured to perform error correction processing on the text to be corrected based on the filtered error correction characters when the hierarchical filtering result includes filtered error correction characters and the filtered error correction characters include error correction characters to be verified, thereby obtaining the text to be verified. The error correction and verification unit is configured to perform error correction and verification by inputting the text to be verified and the preset error correction instruction information into the large language model to obtain the target text.

[0097] In an optional embodiment, the error correction and verification unit is further configured to perform error correction and verification by inputting the text to be corrected, the text to be verified, and preset error correction indication information into the large language model to obtain the target text.

[0098] In an optional embodiment, the above-described apparatus further includes: The preset text pair set acquisition module is configured to acquire the preset text pair set, which includes a first text pair set and a second text pair set; each preset text pair in the first text pair set includes two correct texts; each preset text pair in the second text pair set includes one incorrect text and the correct text corresponding to the incorrect text. The similarity determination module is configured to determine the similarity between the current text pair and each preset text pair in the preset text pair set. The current text pair is a text pair consisting of the text to be corrected and the text to be verified. The similar text pair determination module is configured to perform a similarity-based determination of similar text pairs from a preset set of text pairs for the current text pair. The error correction and verification unit is also configured to perform error correction and verification by inputting similar text pairs, the current text pair, and preset error correction indication information into the large language model to obtain the target text.

[0099] In an optional embodiment, the error correction processing module 540 includes: The second error correction processing unit is configured to perform error correction processing on the text to be corrected based on the target error correction character when the hierarchical filtering result includes the filtered error correction character, the filtered error correction character includes the target error correction character, and the filtered error correction character does not include the error correction character to be verified, thereby obtaining the target text.

[0100] In an optional embodiment, when at least one error correction model includes multiple error correction models, any one of the at least one error correction position corresponds to one or at least two initial error correction characters; the above apparatus further includes: The candidate error correction character determination module is configured to determine the candidate error correction character for each of the at least one error correction positions based on the first error correction performance data corresponding to the initial error correction character of each of the at least one error correction positions. The hierarchical filtering module 530 is also configured to perform hierarchical filtering on candidate error correction characters corresponding to at least one error correction position based on the first error correction performance data corresponding to each error correction position, and obtain hierarchical filtering results.

[0101] In an optional embodiment, the hierarchical screening module 530 includes at least one of the following units: The first character determination unit is configured to perform the following: select candidate error correction characters whose first error correction performance data is less than or equal to a first preset threshold and greater than or equal to a second preset threshold from the candidate error correction characters corresponding to at least one error correction position as error correction characters to be verified. The second character determination unit is configured to select candidate error correction words whose first error correction performance data is greater than a first preset threshold from the candidate error correction characters corresponding to at least one error correction position as target error correction characters; The third character determination unit is configured to select candidate error correction characters whose first error correction performance data is less than a second preset threshold from the candidate error correction characters corresponding to at least one error correction position as the error correction characters to be restored.

[0102] In an optional embodiment, the above-described apparatus further includes: Based on the candidate error correction characters at each error correction position, the text to be corrected is processed to obtain the initial text after error correction. Error correction processing module 540 includes: The third error correction processing unit is configured to perform error correction processing on the initial text based on the hierarchical filtering results to obtain the target text.

[0103] In an optional embodiment, the third error correction processing unit includes: The fourth error correction processing unit is configured to use the initial text as the target text when the hierarchical filtering result includes the error correction characters after filtering, and the error correction characters after filtering include or do not include the error correction characters to be checked. or, The fifth error correction processing unit is configured to input the initial text and preset error correction indication information into the large language model for error correction verification when the hierarchical filtering result includes the filtered error correction characters and the filtered error correction characters include the error correction characters to be verified, so as to obtain the target text.

[0104] In an optional embodiment, the third error correction processing unit further includes: The restoration processing unit is configured to perform restoration processing on the initial text based on the characters to be restored and corrected when the hierarchical filtering results include characters to be restored and corrected, so as to obtain the restored text. The fourth error correction processing unit is also configured to use the restored text as the target text; The fifth error correction processing unit is also configured to input the restored text and preset error correction instruction information into the large language model for error correction verification to obtain the target text.

[0105] In an optional embodiment, the error correction performance data determination module 520 includes: The first error correction performance data acquisition unit is configured to acquire model error correction performance data of each error correction model in at least one error correction model and second error correction performance data of each error correction model for each corresponding error correction position. The first error correction performance data determination unit is configured to determine the third error correction performance data of each error correction model for the corresponding error correction position based on the model error correction performance data of each error correction model and the second error correction performance data of each error correction model for the corresponding error correction position. The second error correction performance data unit is configured to execute at least one third error correction performance data based on the same initial error correction character corresponding to the same error correction position in the initial error correction result, and determine the first error correction performance data corresponding to the same initial error correction result at the same error correction position.

[0106] In an optional embodiment, when at least one error correction model is multiple error correction models, the error correction performance data determination module 520 includes: The second error correction performance data acquisition unit is configured to acquire the model error correction performance data of each of the multiple error correction models. The target error correction model determination unit is configured to execute the determination of the target error correction model corresponding to the same initial error correction character at each error correction position from multiple error correction models; The third error correction performance data determination unit is configured to determine the first error correction performance data corresponding to the same initial error correction character at each error correction position based on the model error correction performance data of the target error correction model corresponding to the same initial error correction character at each error correction position.

[0107] In an optional embodiment, the hierarchical screening module 530 includes at least one of the following units: The fourth character determination unit is configured to execute the initial error correction character whose first error correction performance data is less than or equal to a first preset threshold and greater than or equal to a second preset threshold as the error correction character to be verified. The fifth character determination unit is configured to execute the initial error correction character whose first error correction performance data is greater than a first preset threshold among at least one initial error correction character as the target error correction character; The sixth character determination unit is configured to execute the initial error correction character whose first error correction performance data is less than a second preset threshold as the error correction character to be restored.

[0108] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.

[0109] Figure 6This is a block diagram of an electronic device for text error correction processing provided in an embodiment of this application. The electronic device can be a terminal, and its internal structure diagram can be as follows. Figure 6 As shown, the electronic device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage medium. The network interface is used to communicate with external terminals via a network connection. When the computer program is executed by the processor, it implements a text error correction method. The display screen can be a liquid crystal display (LCD) or an e-ink display. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad mounted on the device's casing, or an external keyboard, touchpad, or mouse. Figure 7 This is a block diagram of another electronic device for text error correction processing provided in the embodiments of this application. The electronic device can be a server, and its internal structure diagram can be as follows. Figure 7 As shown, this electronic device includes a processor, memory, and a network interface connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The network interface is used to communicate with external terminals via a network connection. When the computer program is executed by the processor, it implements a text error correction processing method. Those skilled in the art will understand that Figure 6 or Figure 7 The structure shown is merely a block diagram of a portion of the structure related to the present disclosure and does not constitute a limitation on the electronic device to which the present disclosure is applied. A specific electronic device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements. In an exemplary embodiment, an electronic device is also provided, including: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to execute the instructions to implement the text correction processing method as described in the embodiments of this disclosure.

[0110] In an exemplary embodiment, a computer-readable storage medium is also provided, which, when the instructions in the storage medium are executed by a processor of an electronic device, enables the electronic device to perform the text correction processing method of the present disclosure embodiments. In an exemplary embodiment, a computer program product or computer program is also provided, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the text correction processing methods provided in the various optional implementations described above.

[0111] It is understood that in the specific implementation of this application, user-related data is involved. When the above embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0112] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM), etc.

[0113] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the claims.

[0114] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.

Claims

1. A text error correction processing method, characterized in that, The method includes: Obtain initial error correction results from at least one error correction model for at least one error correction position in the text to be corrected; the initial error correction results include at least one initial error correction character, and each initial error correction character corresponds to one of the error correction positions; Determine the first error correction performance data corresponding to the same initial error correction character at the same error correction position in the initial error correction results; Based on the first error correction performance data, at least one initial error correction character in the initial error correction result is subjected to hierarchical filtering to obtain hierarchical filtering results. The hierarchical filtering results include at least one of the filtered error correction characters or the error correction characters to be restored. The filtered error correction characters include at least one of the target error correction characters or the error correction characters to be verified. Based on the hierarchical screening results, the text to be corrected is processed to obtain the corrected target text.

2. The method according to claim 1, characterized in that, The step of performing error correction processing on the text to be corrected based on the hierarchical screening results, to obtain the corrected target text, includes: If the hierarchical filtering result includes post-filter correction characters, and the post-filter correction characters include the correction characters to be verified, then the text to be corrected is corrected based on the post-filter correction characters to obtain the text to be verified. The text to be verified and the preset error correction indication information are input into the large language model for error correction and verification to obtain the target text.

3. The method according to claim 2, characterized in that, The step of inputting the text to be verified and the preset error correction indication information into the large language model for error correction and verification to obtain the target text includes: The text to be corrected, the text to be verified, and the preset error correction indication information are input into a large language model for error correction and verification to obtain the target text.

4. The method according to claim 3, characterized in that, The method further includes: Obtain a preset text pair set, which includes a first text pair set and a second text pair set; each preset text pair in the first text pair set includes two correct texts; each preset text pair in the second text pair set includes one incorrect text and the correct text corresponding to the incorrect text. Determine the similarity between the current text pair and each preset text pair in the preset text pair set, wherein the current text pair is a text pair consisting of the text to be corrected and the text to be verified; Based on the similarity, similar text pairs of the current text pair are determined from the preset text pair set; The step of inputting the text to be corrected, the text to be verified, and the preset error correction indication information into a large language model for error correction and verification to obtain the target text includes: The similar text pairs, the current text pairs, and the preset error correction indication information are input into the large language model for error correction and verification to obtain the target text.

5. The method according to claim 1, characterized in that, The step of performing error correction processing on the text to be corrected based on the hierarchical screening results, to obtain the corrected target text, includes: If the hierarchical filtering result includes the filtered correction character, and the filtered correction character includes the target correction character, and the filtered correction character does not include the correction character to be verified, then the text to be corrected is corrected based on the target correction character to obtain the target text.

6. The method according to any one of claims 1 to 5, characterized in that, When the at least one error correction model includes multiple error correction models, any one of the at least one error correction position corresponds to one or at least two initial error correction characters; the method further includes: Based on the first error correction performance data corresponding to the initial error correction character of each of the at least one error correction position, determine the candidate error correction character for each of the at least one error correction position. The step of performing hierarchical filtering on at least one initial error correction character in the initial error correction result based on the first error correction performance data, to obtain hierarchical filtering results, includes: Based on the first error correction performance data corresponding to the candidate error correction character at each error correction position, the candidate error correction characters corresponding to the at least one error correction position are subjected to hierarchical filtering to obtain the hierarchical filtering result.

7. The method according to claim 6, characterized in that, The step of performing hierarchical filtering on the candidate error correction characters corresponding to the candidate error correction characters at each error correction position based on the first error correction performance data corresponding to the candidate error correction characters at each error correction position, and obtaining the hierarchical filtering result, includes at least one of the following steps: The candidate error correction characters whose first error correction performance data is less than or equal to a first preset threshold and greater than or equal to a second preset threshold among the candidate error correction characters corresponding to the at least one error correction position are taken as the error correction characters to be verified. The candidate error correction word whose first error correction performance data is greater than the first preset threshold is taken as the target error correction character from the candidate error correction characters corresponding to the at least one error correction position. The candidate error correction character whose first error correction performance data is less than the second preset threshold among the candidate error correction characters corresponding to the at least one error correction position is taken as the error correction character to be restored.

8. The method according to claim 6, characterized in that, The method further includes: Based on the candidate error correction characters at each error correction position, the text to be corrected is processed to obtain the corrected initial text. The step of performing error correction processing on the text to be corrected based on the hierarchical screening results, to obtain the corrected target text, includes: Based on the hierarchical screening results, the initial text is corrected to obtain the target text.

9. The method according to claim 8, characterized in that, The step of performing error correction processing on the initial text based on the hierarchical screening results to obtain the target text includes: If the graded filtering result includes the filtered correction character, and the filtered correction character includes characters that do not include the correction character to be verified, then the initial text is taken as the target text. or, If the graded filtering result includes the filtered error correction character, and the filtered error correction character includes the error correction character to be verified, the initial text and the preset error correction indication information are input into the large language model for error correction verification to obtain the target text.

10. The method according to claim 9, characterized in that, The step of correcting the initial text based on the hierarchical screening results to obtain the target text further includes: If the graded filtering results include the characters to be restored and corrected, the initial text is restored based on the characters to be restored and corrected to obtain the restored text; The step of using the initial text as the target text includes: using the restored text as the target text; The step of inputting the initial text and preset error correction indication information into a large language model for error correction verification to obtain the target text includes: inputting the restored text and preset error correction indication information into a large language model for error correction verification to obtain the target text.

11. The method according to any one of claims 1 to 5, characterized in that, The first error correction performance data corresponding to the same initial error correction character at the same error correction position in the initial error correction result includes: Obtain the model error correction performance data of each error correction model in the at least one error correction model and the second error correction performance data of each error correction model for each corresponding error correction position; Based on the model error correction performance data of each error correction model and the second error correction performance data of each error correction model for the corresponding error correction position, the third error correction performance data of each error correction model for the corresponding error correction position is determined. Based on at least one third error correction performance data of the same initial error correction character corresponding to the same error correction position in the initial error correction results, the first error correction performance data corresponding to the same initial error correction result at the same error correction position is determined.

12. The method according to any one of claims 1 to 5, characterized in that, When the at least one error correction model is multiple error correction models, determining the first error correction performance data corresponding to the same initial error correction character at the same error correction position in the initial error correction result includes: Obtain the model error correction performance data for each of the multiple error correction models; Determine the target error correction model corresponding to the same initial error correction character at each error correction position from the plurality of error correction models; Based on the model error correction performance data of the target error correction model corresponding to the same initial error correction character at each error correction position, the first error correction performance data corresponding to the same initial error correction character at each error correction position is determined.

13. The method according to any one of claims 1 to 5, characterized in that, The step of performing hierarchical filtering on at least one initial error correction character in the initial error correction result based on the first error correction performance data to obtain the hierarchical filtering result includes at least one of the following steps: The initial error correction character whose first error correction performance data is less than or equal to a first preset threshold and greater than or equal to a second preset threshold is taken as the error correction character to be verified. The initial error correction character whose first error correction performance data is greater than the first preset threshold is taken as the target error correction character. The initial error correction character whose first error correction performance data is less than the second preset threshold is selected as the error correction character to be restored.

14. A text error correction processing device, characterized in that, The device includes: The initial error correction result acquisition module is configured to acquire the initial error correction result of at least one error correction model for at least one error correction position in the text to be corrected; the initial error correction result includes at least one initial error correction character, and any initial error correction character corresponds to one error correction position; The error correction performance data determination module is configured to determine the first error correction performance data corresponding to the same initial error correction character at the same error correction position in the initial error correction result. The hierarchical filtering module is configured to perform hierarchical filtering on at least one initial error correction character in the initial error correction result based on the first error correction performance data, and obtain hierarchical filtering results. The hierarchical filtering results include at least one of the filtered error correction characters or error correction characters to be restored. The filtered error correction characters include at least one of the target error correction character or error correction characters to be verified. The error correction processing module is configured to perform error correction processing on the text to be corrected based on the hierarchical filtering results, so as to obtain the corrected target text.

15. An electronic device, characterized in that, include: processor; Memory used to store the processor's executable instructions; The processor is configured to execute the instructions to implement the text correction processing method as described in any one of claims 1 to 13.

16. A computer-readable storage medium, characterized in that, When the instructions in the storage medium are executed by the processor of the electronic device, the electronic device is able to perform the text correction processing method as described in any one of claims 1 to 13.

17. A computer program product, characterized in that, The computer program product includes a computer program stored in a computer-readable storage medium, and a processor reads from and executes the computer program to implement the text correction processing method as described in any one of claims 1 to 13.