Text error correction method and device, electronic equipment and storage medium
Through semantic boundary segmentation and domain knowledge-driven error correction methods on text, multiple error types interleaving and dynamic context adaptability are solved, and efficient text error correction is achieved, suitable for high-risk areas such as medical and legal.
Patent Information
- Application Number
- CN202510806686.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-17
AI Technical Summary
Existing text error correction techniques have performed poorly in dealing with multi-error type interleaving, dynamic context adaptability, and domain-specific knowledge integration, especially in high-risk areas such as medical care and law.
By obtaining the semantic boundaries of the target text, segmenting them into multiple clauses and assigning sequence IDs, semantic analysis is performed by combining the language model and the target domain knowledge base, identifying and adjusting the error correction path, and reorganizing the clauses to generate error correction results, including error location, type and correction suggestions.
It realizes collaborative detection and correction of grammatical, semantic and logical errors, improves the accuracy and efficiency of text error correction technology in complex scenarios, adapts to dynamic context changes, integrates domain knowledge, and supports industry-oriented implementation.
Smart Images

Figure CN120336535A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing, and in particular, to a text error correction method, device, electronic device and storage medium. Background Art
[0002] As one of the core technologies in the field of natural language processing, text error correction has an irreplaceable value in improving information accuracy and ensuring text standardization. With the acceleration of the digitalization process, the scale of text data in scenarios such as legal documents, academic literature, and medical records has increased exponentially, and incorrect texts may lead to semantic ambiguity or even decision-making errors (such as miswriting in medical records resulting in diagnostic deviations).
[0003] Traditional error correction methods mainly rely on explicit rule bases (such as regular expression matching) or static language models (such as n-gram). Although they can handle basic grammar errors, the recognition rate for deep logical errors (such as term expression standardization errors and cross-sentence reference errors) is less than 40%. In recent years, deep learning technologies (such as Seq2Seq models and Transformers) have significantly improved the error correction accuracy, but still face bottlenecks in complex semantic reasoning, dynamic context adaptation, and collaborative processing of multiple error types.
[0004] The current text intelligent error correction technology faces three challenges. First, the problem of intertwined multiple error types is significant. For example, grammar errors, semantic contradictions, and logical fallacies often coexist in the same text, and it is necessary to meet both the requirements of grammar standardization and semantic rationality. Second, the dynamic context adaptability is insufficient. Existing models are difficult to adjust the reasoning path in real time according to the context, resulting in semantic breaks during the processing of long texts. In addition, it is difficult to integrate domain-specific knowledge. For example, the use of professional terms in medical records and the update of sensitive word libraries require building targeted error correction strategies by combining domain knowledge. These problems jointly restrict the application efficiency of traditional methods in complex scenarios, especially in high-risk fields such as medicine and law, where minor errors may cause serious consequences. Summary of the Invention
[0005] The present invention provides a text error correction method, device, electronic device and storage medium to solve the defects of the existing text error correction technology in the problems of intertwined multiple error types, dynamic context adaptability, and domain-specific knowledge integration, realize the collaborative detection and correction of grammar, semantic, and logical errors, effectively cope with the challenges of dynamic context, and provide a solution for the industrial implementation of text intelligent error correction technology.
[0006] The present invention provides a text error correction method, including: Obtain a target text to be error-corrected; Perform semantic boundary recognition on the target text to obtain all semantic boundaries of the target text; Based on the above-mentioned all semantic boundaries, segment the target text into multiple clauses, and assign an ID to each clause according to the order in which each clause appears in the target text, generating structured data with sequence IDs; Input the structured data of each clause into a language model, and call the target domain knowledge base corresponding to the target text as a constraint condition to perform semantic parsing on the structured data, and identify the error types in the structured data; Adjust the error correction path based on the error types, and re-identify the error information in the structured data based on the adjusted error correction path, where the error information includes error content, error type, and error location; Recombine the multiple clauses based on the sequence IDs and the error content, error type, and error location, generating an error correction result including error location, error type, and correction suggestions.
[0007] In a possible implementation manner, the method further includes: Perform a confidence score on each error information in the error correction result; Integrate the error information for each type of error in the error correction result based on the confidence score, where the error information integration includes streamlining error prompts and filtering error information with a confidence score lower than a threshold.
[0008] In a possible implementation manner, the method further includes: Identify the punctuation marks of the target text through a semantic boundary recognition algorithm; Segment the target text into multiple first short sentences based on the punctuation marks; Identify the semantic features of the multiple first short sentences, and identify semantic boundaries based on the semantic features, obtaining all semantic boundaries of the target text.
[0009] In a possible implementation manner, the method further includes: Based on the above-mentioned all semantic boundaries, segment the target text into multiple second short sentences; Extract the semantic embedding vectors of each second short sentence, and calculate the cosine similarity of adjacent second short sentences based on the semantic embedding vectors; Determine whether there is a dependency relationship between pairs of second short sentences with a cosine similarity greater than a threshold; If it is determined that there is a dependency relationship between the pairs of second short sentences, then merge the pairs of second short sentences into independent clauses, obtaining multiple clauses into which the target text is segmented.
[0010] In a possible implementation manner, the method further includes: Generate an inference task for each clause based on a preset model inference service engine, and each inference task corresponds to a thread; Perform multi-threaded parallel processing based on the occupancy of the processor.
[0011] In a possible implementation, the method further includes: Input the structured data of each clause into a language model, and obtain the semantic vector of each clause through model encoding; Retrieve the corresponding knowledge semantic vectors in the target domain knowledge base corresponding to the target text based on the semantic vectors of each clause; Compare the similarity between the semantic vectors of each clause and the knowledge semantic vectors, and determine the knowledge entries most relevant to each clause; Use the knowledge entries as constraint conditions to perform semantic parsing on the structured data of each clause, and identify the error types in the structured data.
[0012] In a possible implementation, the method further includes: If there are multiple error messages in each type of error, sort the multiple error messages in descending order based on the confidence score from high to low; Retain the target error message with the highest confidence in each type of error; Generate a correction suggestion for the target error message retained in each type of error.
[0013] The present invention also provides a text error correction device, including the following modules: An acquisition module, configured to acquire a target text to be corrected; A boundary recognition module, configured to perform semantic boundary recognition on the target text to obtain all semantic boundaries of the target text; A text segmentation module, configured to segment the target text into multiple clauses based on the all semantic boundaries, and assign an ID to each clause according to the order in which each clause appears in the target text, and generate structured data with sequence IDs; An error recognition module, configured to input the structured data of each clause into a language model, and call the target domain knowledge base corresponding to the target text as a constraint condition to perform semantic parsing on the structured data, and identify the error types in the structured data; The error recognition module is further configured to adjust the error correction path based on the error type, and re-identify the error information in the structured data based on the adjusted error correction path, where the error information includes error content, error type, and error location; An error correction module for reorganizing the multiple clauses based on the sequence ID, the error content, the error type, and the error location to generate an error correction result including the error location, the error type, and a correction suggestion.
[0014] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the text error correction method described in any one of the above is implemented.
[0015] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the text error correction method described in any one of the above is implemented.
[0016] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the text error correction method described in any one of the above is implemented.
[0017] The text error correction method, device, electronic device, and storage medium provided by the present invention obtain a target text to be error-corrected; perform semantic boundary recognition on the target text to obtain all semantic boundaries of the target text; based on the all semantic boundaries, segment the target text into multiple clauses, and assign an ID to each clause according to the order in which each clause appears in the target text to generate structured data with a sequence ID; input the structured data of each clause into a language model, and call the target domain knowledge base corresponding to the target text as a constraint condition to perform semantic parsing on the structured data to identify the error type in the structured data; adjust the error correction path based on the error type, and re-identify the error information in the structured data based on the adjusted error correction path, where the error information includes error content, error type, and error location; reorganize the multiple clauses based on the sequence ID, the error content, the error type, and the error location to generate an error correction result including the error location, the error type, and a correction suggestion. Compared with the defects of the existing text error correction technology in the problems of interweaving of multiple error types, dynamic context adaptability, and integration of domain-specific knowledge, in this solution, the input text is standardized and cleaned and semantic unit division is performed through the text preprocessing layer, and then multi-dimensional error detection and correction are realized by combining the domain knowledge base and the dynamic path optimization algorithm in the inference error correction layer. Finally, the error correction results are integrated through the result integration layer and a structured report is generated according to the priority, realizing the collaborative correction of grammar, semantics, and logical errors, and providing a solution for the industrial implementation of the text error correction technology. Description of the Drawings
[0018] To more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0019] Figure 1 It is one of the schematic flowcharts of the text error correction method provided by the present invention.
[0020] Figure 2 It is the second schematic flowchart of the text error correction method provided by the present invention.
[0021] Figure 3 It is the schematic diagram of the error correction categories supported by the text error correction method provided by the present invention.
[0022] Figure 4 It is the architecture diagram of the text error correction based on implicit thought chain verification and dynamic reasoning path optimization provided by the present invention.
[0023] Figure 5 It is the schematic structural diagram of the text error correction device provided by the present invention.
[0024] Figure 6 It is the schematic structural diagram of the electronic device provided by the present invention. Detailed implementation manners
[0025] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments in the present invention belong to the scope of protection of the present invention.
[0026] To facilitate the understanding of the embodiments of the present invention, the following will further explain with specific embodiments in conjunction with the accompanying drawings. The embodiments do not constitute a limitation to the embodiments of the present invention.
[0027] Figure 1 It is one of the schematic flowcharts of the text error correction method provided by the present invention. As Figure 1 shown, the method includes the following: S11. Obtain the target text to be corrected.
[0028] The embodiment of the present invention provides a text error correction method for solving the problems of low efficiency in collaborative correction of multiple error types and difficulty in maintaining semantic coherence in complex text scenarios.
[0029] Specifically, first obtain the initial text to be corrected, and then eliminate text noise through operations such as traditional-simplified conversion, full-width and half-width unification, and special character filtering, and filter special characters, irrelevant symbols (e.g., HTML tags), etc., to obtain the target text after standardization processing.
[0030] S12. Perform semantic boundary recognition on the target text to obtain all semantic boundaries of the target text.
[0031] Based on a semantic boundary recognition algorithm (e.g., a sentence segmentation model that combines punctuation rules and a pre-trained language model), perform semantic boundary recognition on the target text to obtain all semantic boundaries of the target text.
[0032] S13. Based on all the semantic boundaries, segment the target text into multiple clauses, and assign an ID to each clause according to the order in which each clause appears in the target text, generating structured data with sequence IDs.
[0033] The long text is segmented into independent clauses according to semantic boundaries, and at the same time, adjacent clauses with strong dependencies are merged according to logical relevance analysis. The segmented clauses are assigned unique sequence IDs and then encapsulated into structured data units containing sequence IDs and original position information. This design not only ensures the efficiency of parallel reasoning but also avoids the semantic discontinuity problem caused by traditional clause-splitting methods by retaining semantic relevance, providing high-quality input units for subsequent implicit thought chain construction.
[0034] S14. Input the structured data of each clause into a language model, and call the target domain knowledge base corresponding to the target text as a constraint condition to perform semantic parsing on the structured data and identify the error types in the structured data.
[0035] Use a pre-trained language model to perform in-depth semantic parsing on the text, combine a custom knowledge base to construct a logical reasoning chain, perform semantic parsing on the structured data, and identify the error types in the structured data. Professional terms in a specific knowledge domain can be combined. For example, when correcting non-standard medical terms, the corresponding knowledge domain is the medical field.
[0036] S15. Adjust the error correction path based on the error types, and re-identify the error information in the structured data based on the adjusted error correction path. Among them, the error information includes error content, error type, and error location.
[0037] Adaptive selection of error correction paths according to error types and context features. If the preliminary judgment result shows that a certain type of error is not involved in the text, the system will adaptively adjust the error correction path and skip the inspection for that type of error, so as to alleviate the problem of overcorrection and improve the processing efficiency of the entire system. For example, if the model determines that the text does not involve professional term expressions or only involves conventional expressions, then the system will not start the inspection process for knowledge errors and allocate more resources to aspects where errors may exist. In this way, while ensuring the accuracy of error correction, the dynamic strategy optimization module realizes the efficient utilization of resources and the improvement of system performance.
[0038] Furthermore, based on the adjusted error correction path, re-identify the error information corresponding to the above-determined error types in the structured data, where the error information includes error content, error type, and error location. The error location can be determined by the sequence ID.
[0039] S16. Recombine the multiple clauses based on the sequence ID, the error content, the error type, and the error location to generate an error correction result including the error location, the error type, and the correction suggestion.
[0040] Recombine the clauses according to the sequence ID to generate a structured result including the error location, type, and correction suggestion, and determine the priority and content filtering through a confidence scoring mechanism.
[0041] The confidence scoring mechanism comprehensively considers factors such as the severity of the error type, the matching degree of the error description, and the consistency of the context to score the confidence of each error. At the same time, according to the confidence scoring and error situation, the structured result can be presented in a personalized way, such as sorting by error type, severity, etc., and filtering out repeatedly pointed-out errors and errors with too low scores.
[0042] Finally, the result integration module will sort the filtered and organized error correction results in descending order of the error hazard level, and attach the correction suggestion, the original text fragment, and the confidence star rating to each error, and present them to the user in a clear and structured manner.
[0043] The text error correction method provided by the present invention includes: obtaining a target text to be error-corrected; identifying semantic boundaries of the target text to obtain all semantic boundaries of the target text; based on all the semantic boundaries, splitting the target text into multiple clauses, and assigning an ID to each clause according to the order in which each clause appears in the target text to generate structured data with sequence IDs; inputting the structured data of each clause into a language model, and calling the target domain knowledge base corresponding to the target text as a constraint condition to perform semantic parsing on the structured data to identify the error types in the structured data; adjusting the error correction path based on the error types, and re-identifying the error information in the structured data based on the adjusted error correction path, where the error information includes error content, error type, and error location; reorganizing the multiple clauses based on the sequence IDs and the error content, error type, and error location to generate an error correction result including error location, error type, and correction suggestions. Compared with the existing text error correction technologies that perform poorly in problems such as the interweaving of multiple error types, dynamic context adaptability, and integration of domain-specific knowledge, in this method, through the text preprocessing layer, the input text is standardized and cleaned and semantic units are divided, and then in the inference and error correction layer, combined with the domain knowledge base and the dynamic path optimization algorithm, multi-dimensional error detection and correction are realized. Finally, through the result integration layer, the error correction results are integrated and a structured report is generated according to the priority, realizing the collaborative correction of grammar, semantics, and logic errors, and providing a solution for the industrial implementation of text error correction technology.
[0044] Figure 2 It is the second flow diagram of the text error correction method provided by the present invention. As Figure 2 shown, the method includes the following: S21. Obtain a target text to be error-corrected, and identify the punctuation marks of the target text through a semantic boundary recognition algorithm.
[0045] S22. Based on the punctuation marks, split the target text into multiple first short sentences, identify the semantic features of the multiple first short sentences, and based on the semantic features, identify semantic boundaries to obtain all semantic boundaries of the target text.
[0046] An embodiment of the present invention provides a text error correction method for solving the problems of low efficiency of collaborative correction of multiple error types and difficulty in maintaining semantic coherence in complex text scenarios.
[0047] Specifically, in combination with Figure 4A detailed description is provided for the text error correction architecture diagram based on implicit thought chain verification and dynamic inference path optimization. The text error correction architecture implemented in the present invention consists of a text preprocessing layer, an inference and error correction layer, and a result integration layer. In the text preprocessing layer, the original text is sequentially processed by a text preprocessing module and a multi-granularity slicing module to form a structured sentence subset, and dynamic thread pool scheduling is performed in a parallel inference module. Entering the inference and error correction layer, the content after parallel inference is first subjected to real-time constraint injection by an implicit thought chain construction module in combination with a domain knowledge base, and then processed by a dynamic policy optimization module to generate a preliminary parsing result. Finally, in the result integration layer, the preliminary parsing result is sequentially subjected to confidence scoring and custom filtering to eliminate redundant information such as low confidence, and the final accurate error correction result is output.
[0048] Specifically, standard cleaning is performed through the text preprocessing module to provide high-quality input for subsequent error correction. First, irrelevant symbols such as HTML tags are filtered, and the tags are directly removed using regular expressions to avoid interfering with semantic analysis. Next, enter the character-level processing stage: when traversing each character, if a full-width number is found, it will be converted to a half-width through ASCII code offset to ensure unified number format. For half-width English, numbers, and specific symbols (such as @#¥!%, etc.), they are filtered according to requirements rather than retained, and this step is achieved by maintaining a set containing special characters for quick judgment.
[0049] Punctuation processing is a key link: convert half-width punctuation (such as!@#$%^& ()) to full-width, while retaining the half-width commas in numerical values (such as 1,000) and dashes (such as 2023-03-28). Specifically, when implementing, it is determined whether to convert by checking whether the characters before and after the punctuation are numbers. For example, when encountering a comma, it will be judged whether the characters before and after are numbers. If so, the half-width is maintained, otherwise it is converted to full-width. This strategy balances format standardization and numerical readability.
[0050] Traditional and simplified conversion is implemented using the OpenCC tool to convert traditional Chinese text to simplified Chinese. Exceptions are captured through a try-except block during the conversion process to ensure that the system can still run stably when encountering unsupported characters. After character processing, consecutive blank characters (including line breaks and tab characters) are merged, and multiple spaces are replaced with a single space using regular expressions, while removing redundant spaces at the beginning and end of the text.
[0051] To meet the need for unified digital formats, a thousands separator processing function has been added. For example, when encountering a numerical value like 1,000, the comma is removed through regular expression matching and unified into the format of 1000. This optimization ensures the consistency of numerical values in subsequent processing and avoids logical errors caused by format differences. The entire preprocessing process effectively eliminates text noise through multi-stage filtering and rule verification, providing a reliable foundation for subsequent implicit thought chain construction and dynamic reasoning.
[0052] Furthermore, the multi-granularity slicing module takes semantic boundary recognition and logical relevance analysis as the core, and combines the dual mechanisms of rules and models to achieve intelligent segmentation and merging of text. First, a sentence segmentation model based on BERT is used to initially segment the text. Based on punctuation rules, this model optimizes the segmentation boundary by learning context semantic features (such as logical connectives like "therefore" and "however") to obtain all semantic boundaries of the target text. For example, when it detects a causal relationship between the clause before the comma and the subsequent clause, it will preferentially retain the complete semantic unit.
[0053] S23. Based on all the semantic boundaries, segment the target text into multiple second short sentences, extract the semantic embedding vectors of each second short sentence, and calculate the cosine similarity between adjacent second short sentences based on the semantic embedding vectors.
[0054] S24. Determine whether there is a dependency relationship between the pairs of second short sentences with a cosine similarity greater than the threshold. If it is determined that there is a dependency relationship between the pairs of second short sentences, then merge the pairs of second short sentences into independent clauses to obtain multiple clauses into which the target text is segmented.
[0055] Based on all the semantic boundaries, enter the logical relevance analysis stage. Use Sentence-BERT to generate the semantic embedding vectors of each second short sentence, and calculate the cosine similarity between adjacent second short sentences. For pairs of second short sentences with a similarity higher than the threshold (such as 0.85), further determine whether there is a strong dependency relationship through dependency syntax analysis. For example, if the subject of clause B is the object of clause A, it is determined to be strongly associated and the merging operation is triggered.
[0056] For complex text scenarios, a dynamic threshold adjustment mechanism is introduced. In legal documents, since legal clauses often contain multi-layer nested structures, the threshold can be increased to 0.9 to retain more complete semantic units; while in ordinary texts, it can be reduced to 0.8 to improve the flexibility of segmentation.
[0057] S25. Assign an ID to each clause according to the order in which each clause appears in the target text, and generate structured data with sequence IDs.
[0058] Finally, the sliced clauses are encapsulated into structured data units containing sequence IDs and original position information. This design not only ensures the efficiency of parallel inference but also avoids the semantic discontinuity problem caused by traditional clause splitting methods by preserving semantic relevance, providing high-quality input units for subsequent implicit chain of thought construction.
[0059] S26. Input the structured data of each clause into the language model and obtain the semantic vector of each clause through model encoding.
[0060] In the embodiments of the present invention, a model inference service engine is also used to generate an inference task for each clause, and each inference task corresponds to a thread; multi-threaded parallel processing is performed based on the occupancy of the processor.
[0061] Specifically, on the basis of combining VLLM to manage the large language model (LLM), the parallel inference module uses the acceleration mechanism of VLLM to perform efficient and stable parallel processing around GPU resources. VLLM adopts the PagedAttention technology to store the key-value cache required for attention calculation in pages, effectively reducing GPU video memory fragmentation and greatly improving video memory utilization. Relying on this feature of VLLM, the module can process more inference tasks on a single GPU, thereby enhancing the overall parallel processing ability.
[0062] In terms of task scheduling, the module constructs a dynamic thread pool scheduling algorithm based on the asynchronous inference interface of VLLM. Each thread in the thread pool corresponds to an independent inference task, and these tasks are submitted to the inference engine of VLLM for processing. The module monitors the video memory occupancy of the GPU in real time. When the video memory occupancy exceeds a preset threshold (e.g., 90%), the system automatically reduces the number of concurrent threads and puts new inference tasks into the task queue for waiting. At the same time, to avoid the overhead caused by frequent thread creation and destruction, the module maintains a certain number of idle threads so that they can be quickly reused after the video memory resources are released.
[0063] To ensure the stability of the system under high load, the module designs a perfect exception handling mechanism. During the inference process, if the inference task of a certain clause times out or fails, the system will make a judgment based on the error feedback information of VLLM, automatically reduce the number of concurrent threads, and then re-add the task to the task queue. If the retry fails multiple times, the system will record detailed error logs, including relevant information of the task, error type, and occurrence time, etc., and skip this clause to continue processing other tasks to avoid a single point of failure affecting the entire inference process.
[0064] Further, input the structured data of each clause into a language model, and obtain the semantic vector of each clause through model encoding.
[0065] S27. Retrieve the corresponding knowledge semantic vectors in the target domain knowledge base corresponding to the target text based on the semantic vectors of each clause.
[0066] S28. Compare the similarity between the semantic vectors of each clause and the knowledge semantic vectors, and determine the knowledge entries most relevant to each clause.
[0067] In the implicit thought chain construction module, when using a pre-trained language model for in-depth semantic parsing, the model will be fine-tuned according to different domain characteristics to better meet the needs of text semantic understanding in a specific domain. At the same time, when constructing a logical reasoning chain in combination with a specific domain knowledge base, the content in the knowledge base will be updated and maintained in real time according to industry regulations updates, social opinion fluctuations, etc., so as to ensure the accuracy and effectiveness of the logical reasoning chain.
[0068] Specifically, in the logical reasoning chain construction link, a retrieval enhancement architecture is adopted. The clause is encoded by the model to obtain a semantic vector , and it is input into a (Facebook AI Similarity Search, FAISS) vector database for retrieval. Assume there are domain knowledge entries in the database, and the vector corresponding to each entry is . By calculating the similarity between the semantic vector to determine the most relevant knowledge entry. The commonly used similarity calculation method is cosine similarity:
[0069] Sort according to the similarity, and select the top most relevant knowledge entries to form a knowledge set , which will be used as additional context to supplement the original clause. These knowledge entries are like adding more background information to the original clause, making the information of the original clause more abundant. The information of the original clause is limited. When relevant knowledge entries are added as context, the model can better understand the semantics of the clause. In the construction of the logical reasoning chain, these supplementary context knowledge can be used as the basis for reasoning. Abundant context knowledge can help the model avoid incorrect reasoning. If there is only the original clause, the model may have incorrect associations or reasoning due to insufficient information. And the additional context knowledge can provide correct guidance.
[0070] The domain knowledge base maintains real-time performance through an asynchronous update mechanism. For a specific domain knowledge base, an incremental update scheme is implemented. Taking the legal provisions database as an example, assume that the set of legal documents stored in the original database is , and the set of newly obtained legal texts is . Calculate the hash value of each text through the SimHash algorithm . Detect text changes by comparing the Hamming distance of the hash values . When the threshold is exceeded, it is considered that the text has changed, and only the changed part is synchronized to the vector database to reduce the update cost.
[0071] S29. Use the knowledge entry as a constraint condition to perform semantic parsing on the structured data of each clause, and identify the error type in the structured data.
[0072] S210. Adjust the error correction path based on the error type, and re-identify the error information in the structured data based on the adjusted error correction path.
[0073] S211. Recombine the multiple clauses based on the sequence ID, error content, error type, and error location to generate an error correction result containing the error location, error type, and correction suggestions.
[0074] The dynamic policy optimization module, as the core component of the text intelligent error correction system, consists of two key parts, namely, a fine-tuning model for accurate error category judgment, and path selection based on the preliminary judgment result to alleviate overcorrection.
[0075] In the model fine-tuning stage, the module uses the LoRA (Low - Rank Adaptation) technique to efficiently adapt the pre-trained large language model (LLM). The parameters of the pre-trained model are divided into frozen backbone parameters and trainable low-rank matrix parameters . Assume that the weight matrix of the model is , and after LoRA adjustment, it becomes , and its calculation formula is:
[0076] Among them, is a low-rank matrix with a shape of , is a low-rank matrix with a shape of , is the rank of the low-rank matrix, is the dimension of the original weight matrix. In this way, the number of trainable parameters is greatly reduced, significantly reducing the video memory requirements.
[0077] During the training process, in order to enable the model to better adapt to the text semantic understanding requirements of a specific knowledge domain, a special loss function was designed. Taking cross-sentence logical error recognition as an example, assume that the input text contains sub-clauses, and the true label corresponding to each sub-clause is , and the model prediction result is . The triplet loss function is used to optimize the model:
[0078] where is the anchor sample, is the positive sample, is the negative sample, is the distance metric function (such as Euclidean distance), is the margin parameter. This loss function prompts the model to learn the logical relationships between different sub-clauses and improves the recognition ability for problems such as time expression contradictions and knowledge expression errors. By continuously adjusting the parameters of the trainable low-rank matrix, the model reaches the optimal under this loss function, so as to accurately judge the possible error categories in the text.
[0079] The model after the above fine-tuning can make a preliminary judgment on the error category of the input text. For example, Figure 3 is the error correction type supported by the embodiments of the present invention. Another part of the dynamic policy optimization module will use this preliminary judgment result as the basis for path selection. If the preliminary judgment result shows that a certain type of error is not involved in the text, the system will adaptively adjust the error correction path and skip the inspection module for this type of error, so as to alleviate the problem of over-correction and improve the processing efficiency of the entire system. For example, if the model determines that the text does not involve professional term expressions or only involves conventional expressions, then the system will not start the inspection process for knowledge errors and invest more resources in aspects where errors may exist. In this way, while ensuring the accuracy of error correction, the dynamic policy optimization module realizes the efficient use of resources and the improvement of system performance.
[0080] S212. Perform a confidence score on each error message in the error correction result.
[0081] S213. Integrate the error messages for each type of error in the error correction result based on the confidence score.
[0082] As the final output hub of the text intelligent error correction system, the result integration module efficiently presents the structured error correction results through sequence reorganization, confidence evaluation, and dynamic filtering mechanisms. After receiving the error correction request results initiated for each clause during the parallel inference stage, the module constructs an ordered data structure based on the clause sequence ID and integrates the results of these parallel requests in the order of the original text.
[0083] In terms of confidence evaluation, the module uses a specially designed confidence scoring model to score each error item. The model adopts a hierarchical weighted algorithm, and its core formula is:
[0084] Among them, is the weight of the error hazard level. For example, in the case of time errors that seriously affect the accuracy of the content, the hazard level takes a relatively high value, such as 4. is the weight of the description matching degree, and the matching degree between the error description and the actual situation of the text is calculated by using a large language model (LLM) fine-tuned in a specific domain. is the weight of the context consistency. The logical coherence between the correction suggestion and the context before and after is verified through dependency syntactic analysis. Hazard(e), Match(d), and Contexd(c) are all set on a 5-point scale. Hazard(e) represents the current error level score, taking the error category as the input, with a maximum of 5 points indicating the most significant error, and a preset error list is used to rank the error categories; Match(d) represents the current description consistency score, taking the error category, error location, error correction content, and error suggestion as the input, and evaluating the consistency of the four. Because there is a certain degree of instability in model generation, there may be a mismatch among the four, and a maximum of 5 points indicates that the four descriptions are consistent; Contexd(c) represents the coherence of the correction result in the original text, taking the correction content and the context as the input, and evaluating whether the current correction result is abrupt in the original text, with a maximum of 5 points indicating compliance with the original text expression.
[0085] The confidence scoring range is set on a 1 - 5 point scale. Errors with a score of 3 or above are considered credible errors, and 5 points indicate a serious error with high confidence and must be corrected.
[0086] In the dynamic filtering phase, the system supports multi-dimensional filtering strategies. For repeatedly identified errors, a location-based clustering algorithm is adopted: First, errors are grouped according to the locations where they occur (clause ID and text offset), for example, multiple error descriptions within the same clause are grouped into one category; then, within each clustering group, they are sorted in descending order of confidence scores, and only the error description with the highest confidence is retained. This method ensures that among different error descriptions at the same location, the most credible suggestion is retained, avoiding redundant prompts. In specific scenarios, regular expressions are also used to automatically mask correction suggestions involving sensitive information. Errors with confidence scores lower than 3 are automatically ignored through a threshold filtering mechanism.
[0087] Finally, the result integration module will sort the filtered and organized error correction results in descending order of error severity, and attach correction suggestions, original text fragments, and confidence star ratings to each error, presenting them to the user in a clear and structured manner.
[0088] The core of the embodiment of the present invention lies in the collaborative architecture of implicit thought chain verification and dynamic inference path optimization. First, through multi-granularity slicing technology, long texts are decomposed into logically related clause units according to semantic boundaries, while retaining the strong correlation of adjacent clauses to avoid semantic fragmentation, forming structured processing units; then, a dynamic thread scheduling mechanism is used to perform parallel inference on the sliced clauses, adjust the concurrency by real-time monitoring of the system resource usage status, improve the processing efficiency while ensuring the depth of semantic parsing, and introduce a task retry mechanism to ensure the stability of inference. At the semantic parsing layer, the system uses a domain-adapted pre-trained language model to deeply encode the text, combines specific domain knowledge bases such as a regularly updated standard term library and a sensitive word library, and constructs a logical inference chain containing specific domain knowledge constraint conditions, so as to accurately capture deep problems such as non-standard professional expressions and cross-sentence reference errors.
[0089] The dynamic policy optimization module, as the decision-making center of the system, adaptively adjusts the error correction path according to the error type (such as grammar errors, knowledge errors) and context features. For example, when it is initially determined that a certain type of error is not involved in the content, the error correction path will be adaptively adjusted to alleviate over-correction and improve the overall efficiency at the same time. Finally, the result integration module reorganizes the clauses in the original sequence, generates a structured report containing error locations, types, and correction suggestions, and sorts the error priorities according to the confidence scores to ensure that high-risk errors are presented first.
[0090] The innovation of the embodiment of the present invention lies in the deep integration of "domain knowledge-driven dynamic reasoning" and "multi-dimensional semantic collaborative verification". By combining implicit thinking chain modeling with adaptive error correction strategies, the system not only solves the semantic fault problem of static models in long text processing, but also realizes the collaborative correction of grammatical, semantic, and logical errors. Its technical value lies in providing an error correction paradigm that takes into account both efficiency and accuracy for high-risk areas. Through a real-time updated knowledge base and a dynamic path optimization mechanism, it can effectively respond to dynamic contextual challenges such as industry specification updates and social public opinion fluctuations, and provides innovative solutions for the industrialization of text intelligent error correction technology.
[0091] The text error correction device provided by the present invention is described below. The text error correction device described below and the text error correction method described above can be referenced to each other.
[0092] Figure 5 : is a structural schematic diagram of a text error correction device provided by the present invention, which specifically comprises: The acquisition module 501 is used to acquire the target text to be corrected. For detailed description, please refer to the relevant description corresponding to the above method embodiment, which will not be repeated here.
[0093] The boundary recognition module 502 is used to perform semantic boundary recognition on the target text to obtain all semantic boundaries of the target text. For detailed description, please refer to the relevant description corresponding to the above method embodiment, which will not be repeated here.
[0094] The text segmentation module 503 is used to segment the target text into multiple clauses based on all the semantic boundaries, and assign an ID to each clause according to the order in which each clause appears in the target text, thereby generating structured data with a sequence ID. For detailed description, please refer to the relevant description corresponding to the above method embodiment, which will not be repeated here.
[0095] The error identification module 504 is used to input the structured data of each clause into the language model, and call the target domain knowledge base corresponding to the target text as a constraint to perform semantic analysis on the structured data, and identify the error type in the structured data. For detailed description, please refer to the relevant description corresponding to the above method embodiment, which will not be repeated here.
[0096] The error identification module 504 is further configured to adjust the error correction path based on the error type, and re-identify error information in the structured data based on the adjusted error correction path, wherein the error information includes error content, error type, and error location. For detailed description, please refer to the relevant description corresponding to the above method embodiment, which will not be repeated here.
[0097] An error correction module 505, configured to reorganize the plurality of clauses based on the sequence ID, the error content, the error type, and the error location, and generate an error correction result including the error location, the error type, and a correction suggestion. For detailed description, please refer to the relevant description corresponding to the foregoing method embodiment, which will not be elaborated herein.
[0098] Figure 6 An example of a schematic physical structure diagram of an electronic device is shown as Figure 6 shown. The electronic device may include: a processor 810, a communication interface 820, a memory 830, and a communication bus 840. Among them, the processor 810, the communication interface 820, and the memory 830 communicate with each other through the communication bus 840. The processor 810 can call the logical instructions in the memory 830 to execute a text error correction method, which includes: obtaining a target text to be corrected; performing semantic boundary recognition on the target text to obtain all semantic boundaries of the target text; based on all the semantic boundaries, splitting the target text into a plurality of clauses, and assigning an ID to each clause according to the order in which each clause appears in the target text, generating structured data with a sequence ID; inputting the structured data of each clause into a language model, and calling the target domain knowledge base corresponding to the target text as a constraint condition to perform semantic parsing on the structured data, and identifying the error type in the structured data; adjusting the error correction path based on the error type, and re-identifying the error information in the structured data based on the adjusted error correction path, where the error information includes error content, error type, and error location; reorganizing the plurality of clauses based on the sequence ID, the error content, the error type, and the error location, and generating an error correction result including the error location, the error type, and a correction suggestion.
[0099] In addition, when the logical instructions in the foregoing memory 830 are implemented in the form of a software functional unit and sold or used as an independent product, they may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc that can store program codes.
[0100] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the text error correction method provided by each of the above methods. The method includes: obtaining a target text to be corrected; performing semantic boundary recognition on the target text to obtain all semantic boundaries of the target text; based on all the semantic boundaries, splitting the target text into multiple clauses, and assigning an ID to each clause according to the order in which each clause appears in the target text to generate structured data with sequence IDs; inputting the structured data of each clause into a language model, and calling the target domain knowledge base corresponding to the target text as a constraint condition to perform semantic parsing on the structured data to identify the error types in the structured data; adjusting the error correction path based on the error types, and re-identifying the error information in the structured data based on the adjusted error correction path, where the error information includes error content, error type, and error location; reorganizing the multiple clauses based on the sequence IDs and the error content, error type, and error location to generate an error correction result including error location, error type, and correction suggestions.
[0101] On another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it realizes the execution of the text error correction method provided by each of the above methods. The method includes: obtaining a target text to be corrected; performing semantic boundary recognition on the target text to obtain all semantic boundaries of the target text; based on all the semantic boundaries, splitting the target text into multiple clauses, and assigning an ID to each clause according to the order in which each clause appears in the target text to generate structured data with sequence IDs; inputting the structured data of each clause into a language model, and calling the target domain knowledge base corresponding to the target text as a constraint condition to perform semantic parsing on the structured data to identify the error types in the structured data; adjusting the error correction path based on the error types, and re-identifying the error information in the structured data based on the adjusted error correction path, where the error information includes error content, error type, and error location; reorganizing the multiple clauses based on the sequence IDs and the error content, error type, and error location to generate an error correction result including error location, error type, and correction suggestions.
[0102] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative work.
[0103] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution, in essence, or the part that contributes to the prior art can be embodied in the form of a software product, and this computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0104] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or equivalently replace some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A text error correction method, characterized in that, Including: Obtain a target text to be error-corrected; Perform semantic boundary recognition on the target text to obtain all semantic boundaries of the target text; Based on all the semantic boundaries, segment the target text into multiple clauses, and assign an ID to each clause according to the order in which each clause appears in the target text, generating structured data with sequence IDs; Input the structured data of each clause into a language model, and call the target domain knowledge base corresponding to the target text as a constraint condition to perform semantic parsing on the structured data, and identify the error types in the structured data; Adjust the error correction path based on the error types, and re-identify the error information in the structured data based on the adjusted error correction path, where the error information includes error content, error type, and error location; Recombine the multiple clauses based on the sequence IDs and the error content, error type, and error location to generate an error correction result including error location, error type, and correction suggestions.
2. The method according to claim 1, wherein The method further includes: Perform confidence scoring on each error information in the error correction result; Integrate the error information for each type of error in the error correction result based on the confidence scoring, where the error information integration includes streamlining error prompts and filtering error information with a confidence scoring lower than a threshold.
3. The method according to claim 1, wherein The performing semantic boundary recognition on the target text to obtain all semantic boundaries of the target text includes: Identify the punctuation marks of the target text through a semantic boundary recognition algorithm; Segment the target text into multiple first short sentences based on the punctuation marks; Identify the semantic features of the multiple first short sentences, and identify semantic boundaries based on the semantic features to obtain all semantic boundaries of the target text.
4. The method according to claim 3, wherein The segmenting the target text into multiple clauses based on all the semantic boundaries includes: Segment the target text into multiple second short sentences based on all the semantic boundaries; Extract the semantic embedding vectors of each second short sentence, and calculate the cosine similarity of adjacent second short sentences based on the semantic embedding vectors; Determine whether there is a dependency relationship between pairs of second short sentences with a cosine similarity greater than a threshold; If it is determined that there is a dependency relationship between the pairs of second short sentences, then merge the pairs of second short sentences into independent clauses to obtain the multiple clauses into which the target text is segmented.
5. The method according to claim 1, wherein The method further includes: Generate an inference task for each clause through a preset model inference service engine, and each inference task corresponds to a thread; Perform multi-threaded parallel processing based on the occupancy of the processor.
6. The method according to claim 1, characterized in that The inputting the structured data of each clause into a language model, and calling the target domain knowledge base corresponding to the target text as a constraint condition to perform semantic parsing on the structured data, and identifying the error types in the structured data includes: Input the structured data of each clause into the language model and obtain the semantic vector of each clause through model encoding; Retrieve the corresponding knowledge semantic vector in the target domain knowledge base corresponding to the target text based on the semantic vector of each clause; Compare the similarity between the semantic vectors of each clause and the knowledge semantic vector to determine the knowledge entry most relevant to each clause; Use the knowledge entry as a constraint condition to perform semantic parsing on the structured data of each clause, and identify the error types in the structured data.
7. The method according to claim 2, characterized in that, The error information integration for each type of error in the error correction result based on the confidence score includes: If there are multiple error messages in each type of error, arrange the multiple error messages in descending order based on the confidence score from high to low; Retain the target error message with the highest confidence in each type of error; Generate a correction suggestion for the target error message retained in each type of error.
8. A text error correction device, characterized in that, Including: An acquisition module for acquiring the target text to be error-corrected; A boundary recognition module for performing semantic boundary recognition on the target text to obtain all semantic boundaries of the target text; A text segmentation module for segmenting the target text into multiple clauses based on the all semantic boundaries, and assigning an ID to each clause according to the order in which each clause appears in the target text, generating structured data with sequence IDs; An error recognition module for inputting the structured data of each clause into a language model, and calling the target domain knowledge base corresponding to the target text as a constraint condition to perform semantic parsing on the structured data, and identifying the error types in the structured data; The error recognition module is further configured to adjust the error correction path based on the error type, and re-identify the error information in the structured data based on the adjusted error correction path, where the error information includes error content, error type, and error location; An error correction module for reorganizing the multiple clauses based on the sequence ID and the error content, error type, and error location, and generating an error correction result including error location, error type, and correction suggestion.
9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the text error correction method according to any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the text error correction method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Remote consultation record text error correction method based on natural language processing
CN110110334A
Medical text error correction method and device, storage medium and electronic equipment
CN115048937A
Text error correction method and device, electronic equipment and computer storage medium
CN115455947A
Text error correction method and device, computer equipment and storage medium
CN116484843A
Text error correction method and device and electronic equipment
CN119047460A
Cited By
Error correction method and system based on large language model, and medium
CN120953018A
Text analysis method for default management based on artificial intelligence
CN121071115A
A text analysis method for default management based on artificial intelligence
CN121071115B
Data error correction method, data error correction device and storage medium
CN121434323A
Data error correction method, data error correction device, and storage medium
CN121434323B