Automatic text proofreading system and method based on natural language processing

By building an automatic text proofreading system, combining grammatical relationship tree and term consistency verification, the problem of inaccurate term proofing in the existing technology is solved, and the automatic proofreading effect of high-quality text processing is achieved.

CN120337909AInactive Publication Date: 2025-07-18GUANGZHOU BODUO ENG TECH CONSULTING CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510483526.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-07-18
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

When existing natural language processing technologies are comprehensively proofreading of multi-level grammatical structures and professional fields, they cannot accurately identify and associate multiple error types, resulting in the automatic proofreading results being inaccurate enough to meet the needs of high-quality text processing.

Method used

Build a text automatic proofreading system based on natural language processing, including text preprocessing module, grammatical topology analysis module, domain term dynamic adaptation module, error association map construction module and intelligent correction generation module. Through the construction of a grammatical relationship tree, term consistency check and error association map, accurate recognition and orderly correction of multiple types of errors can be achieved.

Benefits of technology

It realizes the full-link processing from text infrastructure analysis to error priority correction, improves the logic and pertinence of the proofreading process, ensures correct grammar and unified terms, and is suitable for professional scenarios such as education, publishing, medical care, and law.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120337909A_ABST
    Figure CN120337909A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of natural language processing, in particular to an automatic text proofreading system and method based on natural language processing, and the system comprises a text preprocessing module, a grammar topology analysis module, a domain term dynamic adaptation module, an error association graph construction module and an intelligent correction generation module. Wherein the text preprocessing module is used for receiving original text input and performing sentence segmentation processing; the grammar topology analysis module is used for constructing a grammar relation tree; the domain term dynamic adaptation module is used for checking term consistency; the error association graph construction module is used for generating an error association graph; and the intelligent correction generation module is used for generating a final proofreading text based on a correction priority algorithm. According to the method, a processing mechanism combining grammar topology analysis and term consistency verification is introduced, the error correlation graph with the weight is constructed, the proofreading text is generated based on the correction priority strategy, and accurate recognition and orderly correction of multiple types of errors are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and particularly to a text automatic proofreading system and method based on natural language processing. Background Art

[0002] In the environment where natural language processing and artificial intelligence technologies are becoming increasingly mature, text automatic proofreading has become an important requirement in the field of language applications; due to the fact that texts contain various grammatical structures and professional terms at the same time, manual proofreading is prone to problems such as omissions and low efficiency, and the demand for automated proofreading technologies has gradually increased; existing statistical models and simple rule engines can only perform basic misspelling detection at the grammatical level, and it is difficult to take into account the consistency of context and professional terms, resulting in inaccurate automatic proofreading results and being unable to fully meet the requirements of high-quality text processing.

[0003] Existing technologies have deficiencies in dealing with comprehensive proofreading of multi-level grammatical structures and professional field terms, are unable to accurately identify and associate multiple error types, and are difficult to effectively quantitatively analyze the relationship between grammatical errors and term errors, resulting in possible neglect of more highly correlated errors or repeated corrections during the automatic proofreading process. Summary of the Invention

[0004] Based on the above purposes, the present invention provides a text automatic proofreading system and method based on natural language processing.

[0005] A text automatic proofreading system based on natural language processing includes a text preprocessing module, a grammatical topology analysis module, a domain term dynamic adaptation module, an error correlation graph construction module, and an intelligent correction generation module; wherein:

[0006] The text preprocessing module: is used to receive the original text input and perform sentence splitting processing to generate a standardized text with position markers;

[0007] The grammatical topology analysis module: is used to receive the standardized text and construct a grammatical relationship tree through dependency syntactic analysis, and output a corrected indication text containing error markers;

[0008] The domain term dynamic adaptation module: is used to receive the corrected indication text, call an external domain term database for term consistency verification, and generate a term-unified text;

[0009] The error correlation graph construction module: is used to receive the term-unified text, analyze the correlation relationship between its grammatical errors and term errors, and generate an error correlation graph with weights;

[0010] The intelligent correction generation module: is used to receive the error correlation graph and generate a final proofread text based on a correction priority algorithm.

[0011] Optionally, the text preprocessing module includes an end-of-sentence recognition unit, a boundary location unit, and a position marking unit; where:

[0012] End-of-sentence recognition unit: It is used to perform character-by-character scanning on the input original text, and use a preset end-of-sentence punctuation library to identify the termination position of each sentence;

[0013] Boundary location unit: It is used to receive the sentence termination position recognized by the end-of-sentence recognition unit, trace back forward according to the sentence termination position to determine the starting position of the sentence, and record the position coordinates of the starting character and the termination character of the sentence in the text, generating coordinate data containing the boundary positions of each sentence;

[0014] Position marking unit: It is used to intercept the original text sentence by sentence according to the sentence coordinate data output by the boundary location unit, and assign a unique position identification number to each sentence in turn. This identification number is arranged in ascending order according to the sequence of sentences in the original text, so as to form a standardized text with position markings.

[0015] Optionally, the syntactic topology analysis module includes a dependency syntactic analysis unit, a syntactic relationship tree construction unit, and an error marking output unit; where:

[0016] Dependency syntactic analysis unit: It is used to receive the standardized text with position markings generated by the text preprocessing module, and use a preset dependency syntactic analysis model to perform dependency syntactic analysis on each word in the sentence one by one with the sentence as the unit, determine the subject-predicate, verb-object, attributive-middle or adverbial-middle syntactic dependency relationships between the words, and generate a syntactic analysis result with dependency relationships;

[0017] Syntactic relationship tree construction unit: It is used to receive the syntactic analysis result generated by the dependency syntactic analysis unit, use the core predicate in the sentence as the root node, and connect other words step by step according to the dependency relationship to form a syntactic relationship tree with a tree-like topological structure, and retain the hierarchical dependency paths between the nodes;

[0018] Error marking output unit: It is used to detect the structural integrity and dependency relationship correctness of the syntactic relationship tree, mark the nodes or paths that violate the grammar rules detected as error positions, and record the original position identification number corresponding to the error node and the specific error type, forming a corrected indication text containing error markings.

[0019] Optionally, the dependency syntactic analysis unit includes:

[0020] Word vector generation: Receive the sentences in the standardized text with position markings, and convert each word in the sentence into a word vector sequence;

[0021] Dependency relationship prediction: Input the word vector sequence into a preset deep dependency parsing model for dependency relationship prediction, and calculate the dependency relationship probability score R between each pair of words in the sentence ij ;

[0022] Dependency structure decoding: Based on the dependency relationship probability scores, use the maximum spanning tree algorithm to optimize and decode the probability scores of all word pairs, determine the dependency parent node of each word, and thus obtain the optimal dependency structure tree. The calculation formula is: In the formula, T * is the optimal dependency structure tree after decoding; is the set of all dependency trees; (w i , w j ) represents the dependency edge from node w i to node w j in the dependency tree.

[0023] Optionally, the domain term dynamic adaptation module includes a term extraction unit, a database matching unit, and a term replacement unit; where:

[0024] Term extraction unit: Used to receive the corrected instruction text output by the syntactic topology analysis module, and identify and extract suspected term error words sentence by sentence according to the position and type of error marks in the text, and generate a list of terms to be verified;

[0025] Database matching unit: Used to receive the list of terms to be verified, call the external domain term database one by one, compare the terms to be verified one by one through the term similarity calculation formula, determine the standard form of the terms, and generate a term verification matching mapping table according to the standard term with the highest similarity;

[0026] Term replacement unit: Used to receive the term verification matching mapping table output by the database matching unit, and uniformly replace the suspected term error words in the corrected instruction text with the standard terms according to the corresponding relationship in the mapping table, so as to generate a unified term text.

[0027] Optionally, the term similarity calculation formula is: In the formula, Sim(t a , t b ) represents the similarity between the term to be verified t a and the standard term t b ; C(t a ) and C(t b ) respectively represent the character sets of the terms t a , t b ; |C(t a ) ∩ C(t b )| represents the number of characters in the intersection of the character sets; |C(t a ) ∪ C(tb ) | Represents the number of characters in the union of character sets.

[0028] Optionally, the error correlation graph construction module includes an error feature extraction unit, a correlation relationship calculation unit, and a correlation graph generation unit; where:

[0029] Error feature extraction unit: Used to receive the term unified text output by the domain term dynamic adaptation module, and respectively extract the syntactic relationship path features of syntactic errors and the lexical semantic features of term errors according to the position identifiers of the syntactic error nodes and term error nodes marked therein, to form an error feature set;

[0030] Correlation relationship calculation unit: Used to receive the error feature set, and calculate the correlation weight values between each syntactic error node and term error node according to the distance factor and semantic correlation factor between error nodes;

[0031] Correlation graph generation unit: Based on the correlation weight values generated by the correlation relationship calculation unit, construct a weighted error correlation graph with term error nodes and syntactic error nodes as graph nodes and correlation weights as graph edge weights.

[0032] Optionally, the correlation graph generation unit includes:

[0033] Node generation sub-unit: Used to establish an error node set based on the error feature set generated by the error feature extraction unit and based on the position identifiers of each term error node and syntactic error node;

[0034] Edge construction sub-unit: Used to screen according to the correlation weight values output by the correlation relationship calculation unit with a weight threshold θ. When the correlation weight value between any two nodes is greater than the threshold, establish a weighted edge between the nodes to form an edge set;

[0035] Graph storage sub-unit: Used to store the attribute information of the node set and the edge set to form a complete error correlation graph.

[0036] Optionally, the intelligent correction generation module includes a correction priority calculation unit, a correction scheme generation unit, and a proofreading text output unit; where:

[0037] Correction priority calculation unit: Used to receive the weighted error correlation graph output by the error correlation graph construction module, and calculate the correction priority value of each node according to the number of connected edges, the total edge weight, and the node type of each error node in the graph;

[0038] Correction Plan Generation Unit: It is used to receive the node correction priority values generated by the correction priority calculation unit, determine the correction plans for each error node in order from high to low according to the priority, call the preset syntax correction rule library and term correction rule library one by one, and match the optimal correction plan according to the error node type and error characteristics;

[0039] Proofreading Text Output Unit: It is used to correct and replace the error nodes in the text one by one according to the correction plan output by the correction plan generation unit to form the final proofread text.

[0040] The automatic text proofreading method based on natural language processing is implemented by the above-mentioned automatic text proofreading system based on natural language processing, and includes the following steps:

[0041] S1: Receive the input of the original text, perform sentence splitting on the original text and assign position identifiers to obtain a standardized text with position markers;

[0042] S2: Input the standardized text with position markers into the dependency parsing model, construct a syntax relationship tree sentence by sentence, and output a correction instruction text for the identified syntax error positions;

[0043] S3: Receive the correction instruction text output by S2, call the external domain term database for term consistency verification and replacement, and generate a term-unified text;

[0044] S4: Receive the term-unified text, analyze the correlation relationship between syntax errors and term errors in the text, and construct a weighted error correlation graph based on the calculated weight values;

[0045] S5: Based on the error correlation graph, plan multi-node correction plans according to the correction priority algorithm, and perform final correction on the text to output the proofread text.

[0046] Advantages of the present invention:

[0047] In the present invention, by constructing an automatic proofreading process including modules such as text preprocessing, syntax topology analysis, dynamic adaptation of domain terms, construction of error correlation graphs, and intelligent correction generation, full-link processing from text basic structure parsing to error priority correction is realized, syntax errors and term errors can be accurately identified, and the correlation relationship between the two is established, making the proofreading process more logical and targeted.

[0048] In the present invention, by constructing an error correlation graph with weights and introducing a correction priority algorithm, it is ensured that in a text with multiple errors, errors with higher relevance and greater impact are preferentially processed, improving the rationality of the correction strategy and the accuracy of automatic processing; the finally generated proofread text has the characteristics of correct grammar and unified terms, and is applicable to professional scenarios with high requirements for text quality such as education, publishing, medical care, and law. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only those of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0050] Figure 1 It is a schematic diagram of the text automatic proofreading system according to an embodiment of the present invention;

[0051] Figure 2 It is a schematic diagram of the text automatic proofreading method according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0052] The present invention will be described in detail below in conjunction with the drawings and specific embodiments. At the same time, it should be noted here that in order to make the embodiments more detailed, the following embodiments are the best and preferred embodiments. For some well-known technologies, those skilled in the art can also adopt other alternative methods for implementation; moreover, the drawings are only for more specifically describing the embodiments and are not intended to specifically limit the present invention.

[0053] It should be noted that in the specification, when referring to "an embodiment", "embodiment", "exemplary embodiment", "some embodiments", etc., it indicates that the described embodiment may include specific features, structures, or characteristics, but not necessarily every embodiment includes such specific features, structures, or characteristics. Additionally, when combining an embodiment to describe a specific feature, structure, or characteristic, implementing such a feature, structure, or characteristic in combination with other embodiments (whether explicitly described or not) should be within the knowledge of those skilled in the relevant art.

[0054] Generally, terms can be understood at least in part from their use in the context. For example, at least in part depending on the context, the term "one or more" used herein can be used to describe any feature, structure, or characteristic in a singular sense, or can be used to describe a combination of features, structures, or characteristics in a plural sense. Additionally, the term "based on" can be understood as not necessarily intended to convey a set of exclusive factors, but rather, at least in part depending on the context, allowing for the existence of other factors that may not be explicitly described.

[0055] Such asFigure 1 As shown in Figure 1 , the text automatic proofreading system based on natural language processing includes a text preprocessing module, a syntactic topology analysis module, a domain term dynamic adaptation module, an error correlation graph construction module, and an intelligent correction generation module; where:

[0056] Text preprocessing module: It is used to receive the original text input and perform sentence segmentation processing to generate a standardized text with position markers;

[0057] Syntactic topology analysis module: It is used to receive the standardized text and construct a syntactic relationship tree through dependency syntactic analysis, and output a corrected instruction text containing error markers;

[0058] Domain term dynamic adaptation module: It is used to receive the corrected instruction text, call an external domain term database for term consistency verification, and generate a term-unified text;

[0059] Error correlation graph construction module: It is used to receive the term-unified text, analyze the correlation relationship between its syntactic errors and term errors, and generate an error correlation graph with weights;

[0060] Intelligent correction generation module: It is used to receive the error correlation graph and generate the final proofread text based on the correction priority algorithm.

[0061] The text preprocessing module includes an end-of-sentence recognition unit, a boundary positioning unit, and a position marking unit; where:

[0062] End-of-sentence recognition unit: It is used to perform character-by-character scanning on the input original text, and use a preset end-of-sentence punctuation library to identify the termination position of each sentence;

[0063] Table 1 End-of-sentence punctuation library

[0064]

[0065] The steps to identify the end position of a sentence are as follows:

[0066] Character sequence traversal: After receiving the original text, convert the text into a character sequence and traverse it character by character from left to right;

[0067] Match termination symbol: During the traversal process, compare one by one whether the current character belongs to any symbol in the end-of-sentence punctuation library;

[0068] Record end-of-sentence position: When it is detected that the character matches a symbol in the end-of-sentence punctuation library, immediately take the index position of the character in the original text as the termination position of the current sentence and record it;

[0069] Eliminate the interference of pseudo-terminators: Combine the context information of the preceding and following characters (for example, detect whether a punctuation mark is immediately followed by a quotation mark, parentheses, or an abbreviation form), exclude misrecognized punctuation marks in abbreviations or non-terminating structures, and ensure the accuracy of recognition;

[0070] Output the set of termination positions: After traversing the entire text, output the set of all recognized termination positions for use by the boundary localization unit;

[0071] The specific examples are as follows:

[0072] Original text example: The weather is nice today. Let's go for a walk in the park! What do you want to bring? Remember to wear a coat.

[0073] Results of the recognition process:

[0074] The position index of the first termination punctuation mark "!" is 16, corresponding to the sentence: "The weather is nice today. Let's go for a walk in the park!"

[0075] The position index of the second termination punctuation mark "?" is 26, corresponding to the sentence: "What do you want to bring?"

[0076] The position index of the third termination punctuation mark "." is 35, corresponding to the sentence: "Remember to wear a coat."

[0077] Boundary localization unit: Used to receive the sentence termination positions recognized by the end-of-sentence recognition unit, determine the starting position of the sentence by tracing back from the sentence termination position, and record the position coordinates of the starting character and the termination character of the sentence in the text, generating coordinate data containing the boundary positions of each sentence;

[0078] Position marking unit: Used to intercept the original text sentence by sentence according to the sentence coordinate data output by the boundary localization unit, and assign a unique position identification number to each sentence in turn. This identification number increases in sequence according to the order of the sentences in the original text, thus forming a standardized text with position markings; Through the design of each unit of the above text preprocessing module, it can be ensured that subsequent analysis and processing can be quickly and accurately located and called based on the clear position identification of each sentence, effectively guaranteeing the accuracy and efficiency of subsequent text proofreading and processing.

[0079] The syntactic topology analysis module includes a dependency parsing unit, a syntactic relation tree construction unit, and an error marking output unit; among them:

[0080] Dependency parsing unit: Used to receive the standardized text with position markings generated by the text preprocessing module, and use a pre-set dependency parsing model to perform dependency parsing on each word in the sentence one by one, determine the subject-predicate, verb-object, attributive-middle or adverbial-middle syntactic dependency relationships between the words, and generate a syntactic analysis result with dependency relationships;

[0081] Syntactic relation tree construction unit: used to receive the syntactic analysis results generated by the dependency parsing unit, take the core predicate in the sentence as the root node, connect other words step by step according to the dependency relationship, form a syntactic relation tree with a tree-like topological structure, and retain the hierarchical dependency paths between each node;

[0082] Error marker output unit: used to detect the structural integrity and the correctness of the dependency relationship of the syntactic relation tree, mark the nodes or paths that violate the grammar rules detected as error positions, and record the original position identification number and the specific error type corresponding to the error nodes, form a corrected indication text containing error markers for subsequent dynamic adaptation of terms.

[0083] The dependency parsing unit includes:

[0084] Word vector generation: receive the sentences in the standardized text with position markers, and convert each word in the sentence into a sequence of low-dimensional dense word vectors, expressed as: S = {w1, w2,..., w n}, where S represents the word vector sequence of the sentence, w i represents the word vector corresponding to the i-th word, and n represents the number of words included in the sentence;

[0085] Dependency relationship prediction: input the word vector sequence into a pre-set deep dependency parsing model for dependency relationship prediction, calculate the dependency relationship probability score R ij between each pair of words in the sentence, and the specific formula is expressed as: R ij = Softmax(U·tanh(W h w i + W d w j + b)), where R ij is the probability score of the dependency relationship between the word w i and the word w j ; W h and W d are the model weight matrices, corresponding to the head word and the dependent word respectively; b is the bias term; U is the output layer weight matrix; Softmax represents the normalization function;

[0086] Dependency structure decoding: based on the dependency relationship probability scores, use the maximum spanning tree algorithm (MST) to optimize and decode the probability scores of all word pairs, determine the dependency parent node of each word, and thus obtain the optimal dependency structure tree. The calculation formula is: where T * is the optimal dependency structure tree after decoding; is the set of all dependency trees; (w i , w j ) represents the node w in the dependency treei To node w j Dependency edges; through the above steps, the dependency structure of each word in the sentence can be efficiently and accurately determined, significantly improving the accuracy and stability of syntactic analysis.

[0087] The domain term dynamic adaptation module includes a term extraction unit, a database matching unit, and a term replacement unit; among them:

[0088] Term extraction unit: used to receive the corrected instruction text output by the syntax topology analysis module, and identify and extract suspected error terms of terms sentence by sentence according to the position and type of error marks in the text, generating a list of terms to be verified;

[0089] Database matching unit: used to receive the list of terms to be verified, call the external domain term database one by one, compare the terms to be verified one by one through the term similarity calculation formula, determine the standard form of the terms, and generate a term verification matching mapping table according to the standard term with the highest similarity;

[0090] Term replacement unit: used to receive the term verification matching mapping table output by the database matching unit, and uniformly replace the suspected error terms of terms in the corrected instruction text with the standard terms according to the corresponding relationship in the mapping table, thereby generating a unified term text for use by the subsequent error correlation graph construction module; through the setting of the above domain term dynamic adaptation module, the unified and standardized processing of terms can be realized, ensuring the consistency of term usage, and improving the accuracy and reliability of subsequent text analysis and correction.

[0091] The term similarity calculation formula is: In the formula, Sim(t a , t b ) represents the similarity between the term t to be verified a and the standard term t b ; C(t a ) and C(t b ) respectively represent the character sets of the terms t a , t b ; |C(t a ) ∩ C(t b )| represents the number of characters in the intersection of the character sets; |C(t a ) ∪ C(t b )| represents the number of characters in the union of the character sets.

[0092] Steps to generate a term verification matching mapping table

[0093] Obtain the list of terms to be verified: Receive the list of terms to be verified from the term extraction unit, and the list contains multiple words that need to be verified and replaced;

[0094] Query the external domain term database: Call the external domain term database to retrieve all registered standard term information to build a matching candidate set;

[0095] Calculate the similarity score: Calculate the similarity score for each term to be verified and each candidate standard term respectively;

[0096] Select the optimal matching standard term: Based on the similarity score, select the standard term with the highest similarity as the final matching result, and record the corresponding candidate standard term and similarity value;

[0097] Fill in the term verification matching mapping table: For each term to be verified, the database matching unit fills in the "term to be verified", "candidate standard term", "similarity", and "final standard term" in the term verification matching mapping table in sequence for subsequent unified replacement operations by the term replacement unit;

[0098] Output the mapping table: After completing the matching of all terms to be verified, the database matching unit generates the final term verification matching mapping table.

[0099] Table 2 Term Verification Matching Mapping

[0100] Serial number Term to be verified Candidate standard term Similarity Final standard term 1 "Air conditioner” "Air conditioner” 0.95 "Air conditioner” 2 "Statistical analysis” "Statistical analysis” 0.88 "Statistical analysis” 3 "datbase” "database” 0.93 "database” 4 "Respons” "Response” 0.9 "Response” ... ... ... ... ...

[0101] Through the multi-step processing of the above database matching unit, the optimal matching standard term can be efficiently retrieved in a large-scale external domain term database and presented in the form of a clear mapping table, providing a technical basis with high accuracy and operability for subsequent term replacement and text proofreading.

[0102] The error correlation graph construction module includes an error feature extraction unit, a correlation relationship calculation unit, and a correlation graph generation unit; among them:

[0103] Error feature extraction unit: Used to receive the unified term text output by the domain term dynamic adaptation module, and respectively extract the syntactic relationship path features of syntax errors and the lexical semantic features of term errors according to the position identifiers of the marked syntax error nodes and term error nodes in it to form an error feature set;

[0104] Correlation relationship calculation unit: Used to receive the error feature set and calculate the correlation weight value between each syntax error node and term error node according to the distance factor and semantic correlation factor between error nodes; The correlation weight calculation formula is as follows: In the formula, W ij represents the correlation weight value between the syntax error node g j and the term error node t i ; D ij represents the path distance between nodes, expressed by the number of edges of the nodes in the syntactic relationship tree; SemRel(ti , g j ) represents the semantic association strength between nodes; α and β are association weight adjustment coefficients, satisfying α + β = 1, and are used to adjust the weight ratio of the distance factor and the semantic factor;

[0105] Association graph generation unit: Based on the association weight values generated by the association relationship calculation unit, construct a weighted error association graph with term error nodes and syntax error nodes as graph nodes and association weights as graph edge weights, and store and output the attributes of the nodes and edges in the graph, and provide them for the intelligent correction generation module to use.

[0106] The association graph generation unit includes:

[0107] Node generation subunit: Used to establish an error node set based on the error feature set generated by the error feature extraction unit and based on the position identifiers of each term error node and syntax error node, denoted as: V = {v1, v2,..., v n}, where V is the error node set; v i represents the i-th error node, including the node type, position identifier, and error feature;

[0108] Edge construction subunit: Used to screen according to the association weight values output by the association relationship calculation unit with a weight threshold θ. When the association weight value between any two nodes is greater than the threshold, establish a weighted edge between the nodes to form an edge set, and the edge set is denoted as: E = {(v i , v j , W ij )|W ij > θ}, where E is the edge set in the graph; (v i , v j , W ij ) represents the edge connecting nodes v i and v j , and the weight of the edge is W ij ;

[0109] Graph storage subunit: Used to store the attribute information of the node set V and the edge set E to form a complete error association graph G, denoted as: G = (V, E); Through the above settings of the association graph generation unit, the association relationship between error nodes can be intuitively reflected in a quantitative manner, which is convenient for efficiently determining the correction strategy and correction order, and ensuring that the subsequent correction process has a clear logical basis.

[0110] The intelligent correction generation module includes a correction priority calculation unit, a correction scheme generation unit, and a proofreading text output unit; among them:

[0111] Correction Priority Calculation Unit: It is used to receive the weighted error correlation graph output by the error correlation graph construction module, and calculate the correction priority value of each node according to the number of connection edges, the total edge weight, and the node type in the graph. The correction priority calculation formula is as follows: In the formula, P(v i ) represents the correction priority value of node v i ; ∑ j W ij represents the sum of the weights of all edges connected to node v i ; M i represents the number of edges connected to node v i ; T(v i ) represents the node type coefficient, which takes the value of 1 when the node is a term error type and 0.8 when the node is a syntax error type; γ, δ are priority adjustment coefficients, satisfying γ + δ = 1;

[0112] Correction Scheme Generation Unit: It is used to receive the node correction priority values generated by the correction priority calculation unit, determine the correction scheme for each error node in order from high to low according to the priority, and call the preset syntax correction rule library and term correction rule library one by one to match the optimal correction scheme according to the error node type and error characteristics;

[0113] Table 3 Syntax Correction Rule Library

[0114]

[0115] Table 4 Term Correction Rule Library

[0116]

[0117] The steps to match the optimal correction scheme according to the error node type and error characteristics are as follows:

[0118] Step 1: Receive the set of error nodes to be corrected from the correction priority calculation unit, and each node has clear type information (term error or syntax error);

[0119] Step 2: According to the error characteristics recorded by the error nodes (such as subject-verb disagreement, term spelling error, etc.), clarify the specific error manifestation form of each node;

[0120] Step 3: According to the type of the error node, call the corresponding correction rule library: if the node type is a term error, call the term correction rule library; if the node type is a syntax error, call the syntax correction rule library;

[0121] Step 4: In the corresponding rule library, according to the specific error characteristics of the error node, match the error characteristic entries defined in the library one by one to determine the rule scheme matched by the current node;

[0122] Step 5: If the node matches multiple correction rule schemes simultaneously, preferentially select the rule scheme with the highest degree of match with the error characteristics, that is, the scheme with the error characteristics exactly the same as those in the rule library is the optimal scheme;

[0123] Step 6: Record the optimal correction scheme and output it to the proofreading text output unit for text correction.

[0124] Proofreading text output unit: used to perform correction and replacement on the error nodes in the text one by one according to the correction scheme generated by the correction scheme generation unit, form the final proofreading text, and output the result of the final proofreading text in the order of the original text; through the setting of the above intelligent correction generation module, the priority order of correcting each error node can be clarified, and accurate correction can be achieved according to the quantitative standard and the rule library, thereby improving the accuracy and rationality of the automatic proofreading result.

[0125] As Figure 2 shown, the text automatic proofreading method based on natural language processing is implemented by the above-mentioned text automatic proofreading system based on natural language processing, including the following steps:

[0126] S1: Receive the input of the original text, perform sentence splitting on the original text and assign position identifiers to obtain the standardized text with position marks;

[0127] S2: Input the standardized text with position marks into the dependency parsing model, construct the syntax relationship tree sentence by sentence, and output the correction instruction text for the identified syntax error positions;

[0128] S3: Receive the correction instruction text output by S2, call the external domain term database for term consistency verification and replacement, and generate the term-unified text;

[0129] S4: Receive the term-unified text, analyze the correlation relationship between the syntax errors and term errors in the text, and construct a weighted error correlation graph based on the calculated weight values;

[0130] S5: Based on the error correlation graph, plan the multi-node correction scheme according to the correction priority algorithm, and perform the final correction on the text to output the proofreading text; through the above steps, the syntax errors and term errors in the text can be detected and correlated in stages, and finally a highly accurate proofreading text is formed, improving the reliability and efficiency of the text automatic proofreading.

[0131] The present invention encompasses any alternatives, modifications, equivalent methods, and solutions that are made to the essence and scope of the present invention. For the public to have a thorough understanding of the present invention, specific details are described in detail in the following preferred embodiments of the present invention. However, those skilled in the art can fully understand the present invention without these detailed descriptions. Additionally, well-known methods, processes, procedures, components, and circuits, etc., are not described in detail to avoid unnecessary confusion to the essence of the present invention.

[0132] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. An automatic text proofreading system based on natural language processing, characterized in that, It includes a text preprocessing module, a syntactic topology analysis module, a domain term dynamic adaptation module, an error correlation graph construction module, and an intelligent correction generation module; among them: The text preprocessing module: is used to receive the original text input and perform sentence splitting processing to generate a standardized text with position markers; The syntactic topology analysis module: is used to receive the standardized text and construct a syntactic relationship tree through dependency syntactic analysis, and output a corrected instruction text containing error markers; The domain term dynamic adaptation module: is used to receive the corrected instruction text, call an external domain term database for term consistency verification, and generate a term-unified text; The error correlation graph construction module: is used to receive the term-unified text, analyze the correlation relationship between its syntactic errors and term errors, and generate an error correlation graph with weights; The intelligent correction generation module: is used to receive the error correlation graph and generate a final proofreading text based on the correction priority algorithm.

2. The text automatic proofreading system based on natural language processing according to claim 1, characterized in that, The text preprocessing module includes an end-of-sentence recognition unit, a boundary positioning unit, and a position marking unit; among them: The end-of-sentence recognition unit: is used to perform character-by-character scanning on the input original text, and use a preset end-of-sentence punctuation library to identify the termination position of each sentence; The boundary positioning unit: is used to receive the sentence termination position recognized by the end-of-sentence recognition unit, trace back forward according to the sentence termination position to determine the starting position of the sentence, and record the position coordinates of the starting character and the termination character of the sentence in the text, and generate coordinate data including the boundary positions of each sentence; The position marking unit: is used to intercept the original text sentence by sentence according to the sentence coordinate data output by the boundary positioning unit, and assign a unique position identification number to each sentence in turn. This identification number increases in the order of the sentences in the original text, so as to form a standardized text with position markers.

3. The text automatic proofreading system based on natural language processing according to claim 1, characterized in that, The syntactic topology analysis module includes a dependency syntactic analysis unit, a syntactic relationship tree construction unit, and an error marker output unit; among them: The dependency syntactic analysis unit: is used to receive the standardized text with position markers generated by the text preprocessing module, and use a preset dependency syntactic analysis model to perform dependency syntactic analysis on each word in the sentence unit by unit, determine the subject-predicate, verb-object, attributive-middle or adverbial-middle syntactic dependency relationships between the words, and generate a syntactic analysis result with dependency relationships; The syntactic relationship tree construction unit: is used to receive the syntactic analysis result generated by the dependency syntactic analysis unit, use the core predicate in the sentence as the root node, and connect other words step by step according to the dependency relationship to form a syntactic relationship tree with a tree-like topological structure, and retain the hierarchical dependency paths between the nodes; The error marker output unit: is used to detect the structural integrity and dependency relationship correctness of the syntactic relationship tree, mark the nodes or paths that violate the grammar rules detected as error positions, and record the original position identification number and specific error type corresponding to the error nodes, so as to form a corrected instruction text containing error markers.

4. The text automatic proofreading system based on natural language processing according to claim 3, characterized in that, The dependency syntactic analysis unit includes: Word vector generation: Receive the sentences in the standardized text with position markers, and convert each word in the sentence into a word vector sequence; Dependency relationship prediction: Input the sequence of word vectors into a pre-set deep dependency parsing model for dependency relationship prediction, and calculate the dependency relationship probability score R between each pair of words in the sentence ij ; Dependency structure decoding: Based on the dependency relationship probability scores, the maximum spanning tree algorithm is used to optimize the decoding of the probability scores of all word pairs, determine the dependency parent node of each word, and thus obtain the optimal dependency structure tree. The calculation formula is as follows: In the formula, T * is the optimal dependency structure tree after decoding; is the set of all dependency trees; (w i , w j ) represents the dependency edge from node w i to node w j in the dependency tree.

5. The text automatic proofreading system based on natural language processing according to claim 1, characterized in that, The domain term dynamic adaptation module includes a term extraction unit, a database matching unit, and a term replacement unit; where: The term extraction unit: is used to receive the corrected instruction text output by the syntactic topology analysis module, and identify and extract suspected term error words sentence by sentence according to the positions and types of error marks in the text, and generate a list of terms to be verified; The database matching unit: is used to receive the list of terms to be verified, call the external domain term database one by one, compare the terms to be verified one by one through the term similarity calculation formula, determine the standard form of the terms, and generate a term verification matching mapping table according to the standard term with the highest similarity; The term replacement unit: is used to receive the term verification matching mapping table output by the database matching unit, and uniformly replace the suspected term error words in the corrected instruction text with the standard terms according to the corresponding relationship in the mapping table, so as to generate a unified term text.

6. The text automatic proofreading system based on natural language processing according to claim 5, characterized in that, The formula for calculating the term similarity is as follows: In the formula, Sim(t a , t b ) represents the similarity between the term t a to be verified and the standard term t b ; C(t a ) and C(t b ) respectively represent the character sets of the terms t a , t b ; |C(t a ) ∩ C(t b )| represents the number of characters in the intersection of the character sets; |C(t a ) ∪ C(t b )| represents the number of characters in the union of the character sets.

7. The text automatic proofreading system based on natural language processing according to claim 1, characterized in that The error correlation graph construction module includes an error feature extraction unit, a correlation relationship calculation unit, and a correlation graph generation unit; where: The error feature extraction unit: is used to receive the unified term text output by the domain term dynamic adaptation module, and extract the syntactic relationship path features of the syntactic errors and the lexical semantic features of the term errors respectively according to the position identifiers of the syntactic error nodes and term error nodes marked therein, and form an error feature set; The correlation relationship calculation unit: is used to receive the error feature set, and calculate the correlation weight values between each syntactic error node and term error node according to the distance factor and semantic correlation factor between the error nodes; The correlation graph generation unit: based on the correlation weight values generated by the correlation relationship calculation unit, constructs a weighted error correlation graph with term error nodes and syntactic error nodes as graph nodes and correlation weights as graph edge weights.

8. The text automatic proofreading system based on natural language processing according to claim 7, characterized in that The correlation graph generation unit includes: The node generation subunit: is used to establish an error node set based on the position identifiers of each term error node and syntactic error node according to the error feature set generated by the error feature extraction unit; The edge construction subunit: is used to screen according to the correlation weight values output by the correlation relationship calculation unit with a weight threshold θ. When the correlation weight value between any two nodes is greater than the threshold, establish a weighted edge between the nodes to form an edge set; The graph storage subunit: is used to store the attribute information of the node set and the edge set to form a complete error correlation graph.

9. The text automatic proofreading system based on natural language processing according to claim 1, characterized in that, The intelligent correction generation module includes a correction priority calculation unit, a correction scheme generation unit, and a proofreading text output unit; where: The correction priority calculation unit: is used to receive the weighted error correlation graph output by the error correlation graph construction module, and calculate the correction priority value of each node according to the number of connection edges, the total edge weight, and the node type of each error node in the graph; The correction scheme generation unit: is used to receive the node correction priority values generated by the correction priority calculation unit, determine the correction scheme of each error node in turn according to the order from high to low priority, call the preset syntactic correction rule library and term correction rule library one by one, and match the optimal correction scheme according to the error node type and error features; Proofreading text output unit: used to correct and replace the error nodes in the text one by one according to the correction scheme output by the correction scheme generation unit, and form the final proofread text.

10. A text automatic proofreading method based on natural language processing, implemented by the text automatic proofreading system according to any one of claims 1-9, characterized in that, The steps are as follows: S1: Receive the original text input, perform sentence segmentation on the original text and assign position identifiers to obtain the standardized text with position markers. S2: Input the standardized text with position markers into the dependency syntax analysis model, construct a syntax relationship tree sentence by sentence, and output the correction instruction text for the identified syntax error positions. S3: Receive the correction instruction text output by S2, call the external domain term database for term consistency verification and replacement, and generate the text with unified terms. S4: Receive the text with unified terms, analyze the correlation relationship between syntax errors and term errors in the text, and construct a weighted error correlation graph based on the calculated weight values. S5: Based on the error correlation graph, plan the multi-node correction scheme according to the correction priority algorithm, and perform the final correction on the text to output the proofread text.

Citation Information

Cited By

  • Automatic speech recognition method and system based on large model

    CN120564696A

  • Automatic speech recognition method and system based on large model

    CN120564696B

  • Automatic error correction method and system based on knowledge graph

    CN121724031A