A computer processing method for Chinese text semantic error correction

CN122735656APending Publication Date: 2026-09-11LIAONING RENREN CHANGXIANG TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611211636.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-11
Publication Date
2026-09-11

AI Technical Summary

Technical Problem

[0004]本发明的目的是为了解决现有技术中存在的对字词合法、局部表达通顺但依存关系中支配词与被支配词语义角色不匹配的隐性语义错误难以准确识别、定位和纠正的缺点,而提出的一种面向中文文本语义纠错的计算机处理方法

Benefits of technology

[0039] 1. This invention transforms the Chinese text to be corrected from a single character or sentence fluency judgment into a candidate dependency relation jointly defined by the governing word, the governed word, and the dependency relation type by performing word segmentation, part-of-speech tagging, and dependency parsing on the Chinese text to be corrected. Furthermore, it extracts the corresponding local semantic fragments. Based on this, it calculates the surface fluency confidence value and semantic role fit of the candidate dependency relation, and determines the semantic mutual exclusion relationship of dependency roles through the divergence between the two. This invention can identify implicit semantic errors where the words themselves are legal, the syntactic structure is superficially stable, and the local expression is relatively fluent, but the semantic roles of the governing and governed words do not match. It avoids missing surface fluency errors when correcting solely based on misspellings, word co-occurrence, or overall language fluency, thereby improving the accuracy and interpretability of Chinese text semantic correction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122735656A_ABST
    Figure CN122735656A_ABST
Patent Text Reader

Abstract

This invention discloses a computer processing method for semantic error correction of Chinese text, belonging to the field of natural language processing technology. The method includes: acquiring the Chinese text to be corrected and performing sentence segmentation, word segmentation, part-of-speech tagging, and dependency parsing to obtain candidate dependency relations and local semantic fragments; determining surface fluency confidence values ​​based on the local semantic fragments; determining semantic role fit based on the governing word, governed word, and dependency relation type; and determining semantic mutual exclusion relationships of dependency roles based on their divergence; further determining the error-bearing word; generating structure-preserving error-correction candidate words; filtering target error-correction words; and replacing the output error-correction result. This invention can identify implicit semantic errors where the words are legal and the local text is fluent, but the dependency roles do not match, and reduces the risk of over-rewriting.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of natural language processing technology, and in particular to a computer processing method for semantic error correction of Chinese text. Background Technology

[0002] With the development of applications such as office automation, intelligent writing, text review, government document processing, and business report generation, the demand for automatic error correction in Chinese text is constantly increasing. During the input, editing, copying, speech-to-text, intelligent generation, and manual modification processes, Chinese text is prone to problems such as inappropriate word collocation, mismatched action objects, unreasonable modification relationships, and semantic deviations. Unlike typos or explicit grammatical errors, some Chinese semantic errors involve words that are themselves legal and whose local expressions are relatively grammatically fluent, but exhibit a mismatch in the semantic roles of the governing and governed words in specific syntactic relationships. These errors are not easily detected at the surface level of language but can cause deviations in text meaning, mismatches of technical objects, or unclear expression logic, especially in technical specifications, contracts, and business rules, easily affecting the accuracy and rigor of the text.

[0003] Existing Chinese text correction methods typically include dictionary-based misspelling detection, homophone or near-homophone substitution, language model-based fluency assessment, sequence labeling-based error location identification, and generative model-based sentence rewriting. These methods are effective in handling misspellings, omissions, repetitions, word order errors, and common grammatical errors, and can also use language models to determine whether a text segment conforms to common expression habits. However, most of these methods rely on character validity, word co-occurrence frequency, or overall fluency as the main criteria, easily confusing "fluency" with "semantic correctness." For errors where sentences appear fluent but semantic roles within dependency relationships do not match, they lack fine-grained judgment on the semantic fit between the governing word, the governed word, and dependency relationship types. Summary of the Invention

[0004] The purpose of this invention is to address the shortcomings of existing technologies in accurately identifying, locating, and correcting implicit semantic errors, such as the mismatch between the semantic roles of the governing and governed words in dependency relationships, even when the words are legal and the local expressions are fluent. Therefore, this invention proposes a computer processing method for semantic error correction of Chinese text.

[0005] To address the problems existing in the prior art, the present invention adopts the following technical solution:

[0006] A computer processing method for semantic error correction of Chinese text, comprising:

[0007] S1. Obtain the Chinese text to be corrected, and perform sentence segmentation, word segmentation, part-of-speech tagging and dependency parsing on the Chinese text to be corrected to obtain sentences, word sequences and dependency relation sets;

[0008] S2. Based on the dependency relation set, determine the candidate dependency relations containing the governing word, the governed word, and the dependency relation type, and extract the local semantic fragments corresponding to the candidate dependency relations;

[0009] S3. Determine the surface fluency confidence value of candidate dependencies based on local semantic fragments;

[0010] S4. Determine the semantic role fit of candidate dependency relations based on the governing word, the governed word, and the dependency relation type;

[0011] S5. Determine the semantic mutual exclusion relationship of dependent roles based on the divergence between surface fluency confidence value and semantic role fit.

[0012] S6. Determine the error-bearing word based on the semantic mutual exclusion relationship of dependency roles, generate structure-preserving error-correction candidate words, determine the target error-correction word from the structure-preserving error-correction candidate words, and output the error correction result after replacing the error-bearing word with the target error-correction word.

[0013] Preferably, the dependency relation set includes multiple dependency relations, each of which includes a governing word, a governed word, and a dependency relation type. The dependency relation type includes at least one of the following: subject-predicate relation, verb-object relation, attributive-head relation, adverbial-head relation, prepositional-object relation, and complementary relation.

[0014] Preferably, the local semantic segment consists of the governing word, the governed word, and the modifying element located between or adjacent to the governing word and the governed word in the candidate dependency relation.

[0015] Preferably, determining the surface fluency confidence value of candidate dependencies based on local semantic fragments includes:

[0016] Calculate the language fluency loss for local semantic segments;

[0017] Based on the language fluency loss of each candidate dependency relation in the same Chinese text to be corrected, the relative fluency between each candidate dependency relation is determined;

[0018] Candidate dependencies whose relative fluency is higher than the average level of candidate dependencies in the same Chinese text to be corrected are identified as surface fluency high-confidence candidate dependencies, and the relative fluency corresponding to the surface fluency high-confidence candidate dependencies is used as the surface fluency confidence value.

[0019] Preferably, the semantic role fit of candidate dependency relations is determined based on the governing word, the governed word, and the dependency relation type, including:

[0020] Based on the dependency relationship type of the candidate dependency relationship, determine the semantic slot of the dependency role of the governing word to the governed word;

[0021] Obtain the semantic category of the governed word;

[0022] The semantic role fit of candidate dependency relationships is determined based on the degree of matching between the semantic category of the governed word and the semantic slot of the dependency role.

[0023] Preferably, the dependency role semantic slots are obtained from at least one of a general Chinese semantic knowledge base, a domain terminology database, dependency collocation statistics in the correct corpus, or semantic vector clustering results, and are stored corresponding to the semantic category of the governing word, dependency relationship type, and governed word.

[0024] Preferably, determining the semantic mutual exclusion relationship of dependent roles based on the divergence between surface fluency confidence value and semantic role fit includes:

[0025] From the surface-level fluent high-confidence candidate dependency relations, select candidate dependency relations whose semantic role fit is lower than the average level of candidate dependency relations in the same Chinese text to be corrected;

[0026] The selected candidate dependencies are ranked according to the degree of deviation between their surface fluency confidence value and semantic role fit.

[0027] The candidate dependency relationship with the highest degree of deviation is identified as a semantic mutual exclusion relationship of dependency roles.

[0028] Preferably, the error-bearing word is determined based on the semantic mutual exclusion relationship of dependency roles, including:

[0029] When the dependency relationship type of the mutually exclusive relationship of dependency roles is verb-object relationship, prepositional object relationship or complementary relationship, the governed word is first identified as the error-bearing word to be verified, and the corresponding candidate word of the governed word is generated.

[0030] If the highest semantic role fit of the candidate word after replacement is higher than the semantic role fit before replacement, the word being replaced is identified as the incorrectly assigned word.

[0031] If the highest semantic role fit after replacing the candidate word corresponding to the governed word is not higher than the semantic role fit before replacement, the governing word will be identified as the erroneous bearer word.

[0032] When the dependency relationship type of the mutually exclusive semantic relationship is a nominative-head relation or an adverbial-head relation, the semantic role fit after replacing the dominant word and the subordinate word is judged respectively, and the side that increases the semantic role fit is identified as the erroneous bearer word.

[0033] Preferably, generating structure-preserving error correction candidate words, and determining the target error correction word from the structure-preserving error correction candidate words, includes:

[0034] Generate a set of candidate words that are in the same syntactic position as the error-bearing word;

[0035] From the candidate word set, retain candidate words that do not change the dependency relationship type after replacement, satisfy the corresponding dependency role semantic slot, and have a language fluency loss no greater than that of the original local semantic segment, to obtain structure-preserving error correction candidate words;

[0036] The candidate words for structure-preserving error correction are ranked in order of semantic role fit from high to low, language fluency loss from low to high, and edit distance from the error-bearing word from small to large. The candidate word with the highest ranking is then selected as the target error correction word.

[0037] Preferably, when outputting the error correction result, interpretable error correction information is output simultaneously. The interpretable error correction information includes the original dependency relationship, dependency relationship type, governing word, governed word, fault-bearing word, semantic category of fault-bearing word, semantic slot of dependency role, target error correction word, and whether the original dependency relationship type is maintained after replacement.

[0038] Compared with the prior art, the beneficial effects of the present invention are:

[0039] 1. This invention transforms the Chinese text to be corrected from a single character or sentence fluency judgment into a candidate dependency relation jointly defined by the governing word, the governed word, and the dependency relation type by performing word segmentation, part-of-speech tagging, and dependency parsing on the Chinese text to be corrected. Furthermore, it extracts the corresponding local semantic fragments. Based on this, it calculates the surface fluency confidence value and semantic role fit of the candidate dependency relation, and determines the semantic mutual exclusion relationship of dependency roles through the divergence between the two. This invention can identify implicit semantic errors where the words themselves are legal, the syntactic structure is superficially stable, and the local expression is relatively fluent, but the semantic roles of the governing and governed words do not match. It avoids missing surface fluency errors when correcting solely based on misspellings, word co-occurrence, or overall language fluency, thereby improving the accuracy and interpretability of Chinese text semantic correction.

[0040] 2. After determining the semantic mutual exclusion relationship of dependency roles, this invention determines the error-bearing word based on the dependency relationship type and the suitability of the replaced semantic role, and generates structure-preserving error-correction candidate words from the semantic slots of dependency roles. The target error-correction word is selected in the order of semantic role suitability, language fluency loss, and edit distance, so that the target error-correction word can replace the error-bearing word while maintaining the original dependency relationship type and local syntactic structure. It can correct implicit semantic errors while reducing whole-sentence rewriting and excessive polishing, maintaining the original text's expressive intent, technical object, and semantic structure, and simultaneously outputs interpretable error-correction information such as the original dependency relationship, error-bearing word, semantic slot, and target error-correction word, making it easy for users to confirm the basis for error correction. It is suitable for Chinese text processing scenarios with high requirements for semantic accuracy and text rigor, such as technical documents and business reports. Attached Figure Description

[0041] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this application, illustrate exemplary embodiments of the invention and, together with their description, serve to explain the invention and do not constitute an undue limitation thereof. In the drawings:

[0042] Figure 1 This is a flowchart illustrating a computer processing method for semantic error correction of Chinese text, provided as an embodiment of the present invention. Detailed Implementation

[0043] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.

[0044] Example: Figure 1 As shown, this embodiment provides a computer processing method for semantic error correction of Chinese text. The method is executed by a computer device with a processor and a memory. The memory stores a computer program. When the processor executes the computer program, it processes the Chinese text to be corrected according to the steps described in this embodiment. The Chinese text to be corrected in this embodiment includes Chinese sentences, Chinese paragraphs, Chinese technical texts, Chinese explanatory texts, Chinese business texts, or Chinese review texts. The Chinese text to be corrected may contain implicit semantic errors where the words themselves are legal, the local expressions are fluent, and the syntactic structure appears stable, but the governing word and the governed word in the dependency relationship do not match in semantic roles.

[0045] In this embodiment, the Chinese text to be corrected refers to the sequence of Chinese characters input into a computer device that needs to undergo semantic error correction processing. It can be a single sentence or a paragraph of text composed of multiple sentences. The presence or absence of typos is not a limiting condition for the Chinese text to be corrected. Even if all the words in the Chinese text to be corrected are legal Chinese words, as long as there are expressions with mismatched semantic roles, it is still subject to processing in this embodiment.

[0046] The method described in this embodiment identifies semantic mutual exclusion of dependency roles under surface fluency and high confidence. Semantic mutual exclusion of dependency roles under surface fluency and high confidence means that the local semantic segment containing the candidate dependency relationship is fluent or highly fluent at the character, part-of-speech, and local syntax levels, but the semantic category of the governed word in the candidate dependency relationship does not match the semantic slot of the dependency role required by the governing word under this dependency relationship type, or the semantic slot of the dependency role matching the governed word can only be formed after replacing the governing word. For example, in "repair risk", "repair" and "risk" are both legal Chinese words. "Repair risk" forms a verb-object relationship in form, but "repair" usually requires the object to be a loophole, defect, fault, error, or other object that can be repaired, while "risk" usually belongs to an object that can be reduced. Therefore, "repair-risk" constitutes semantic role mutual exclusion in the verb-object dependency relationship.

[0047] This embodiment includes the following steps;

[0048] S1. Obtain the Chinese text to be corrected, and perform sentence segmentation, word segmentation, part-of-speech tagging and dependency parsing on the Chinese text to be corrected to obtain sentences, word sequences and dependency relation sets;

[0049] Specifically, the computer device receives the Chinese text to be corrected, performs encoding cleaning on the Chinese text to be corrected, removes meaningless control characters, and retains Chinese characters, numbers, letters, punctuation marks and the original separation information in the text; then, according to Chinese punctuation marks, line breaks and text paragraph boundaries, the Chinese text to be corrected is segmented into sentences to obtain at least one sentence;

[0050] For any sentence, denoted as:

[0051] ;

[0052] in, To represent a sentence. This indicates the first [number] in the sentence. One word, This indicates the number of words in the sentence; computer devices display the sentence. Word segmentation and part-of-speech tagging are performed to obtain a word sequence and the part of speech of each word; the part of speech includes noun, verb, adjective, adverb, preposition, conjunction, auxiliary word, numeral, classifier, and punctuation category;

[0053] After obtaining the word sequence and part-of-speech sequence, the computer device processes the sentence. Dependency parsing is performed to obtain a set of dependency relations:

[0054] ;

[0055] in, Sentence The corresponding set of dependencies, This indicates the number of dependency relations in the sentence. Indicates the first The dominator in a dependency relation, Indicates the first The governed word in a dependency relationship, Indicates the first The types of dependency relations include at least one of the following: subject-predicate, verb-object, attributive-head, adverbial-head, prepositional-object, and complementary relations.

[0056] In this embodiment, the dependency relation set refers to the set of governing words, governed words, and dependency relation types obtained after performing dependency parsing on a sentence; the governing word is a word that forms syntactic control or semantic constraint on another word in the dependency relation, and the governed word is a word that is constrained by the governing word in the dependency relation; the dependency relation type is used to represent the syntactic relationship between the governing word and the governed word, including at least one of the following: subject-predicate relation, verb-object relation, attributive-head relation, adverbial-head relation, prepositional-object relation, and complement relation;

[0057] For example, when the Chinese text to be corrected contains the sentence "The system can automatically identify anomalies and repair risks;", word segmentation yields "the / system / can / automatically / identify / anomalies / and / repair / risks". Dependency parsing reveals that "repair—risk" is a verb-object relationship, where "repair" is the governing word and "risk" is the governed word, and "verb-object relationship" is the type of dependency relationship.

[0058] Through step S1, the Chinese text to be corrected is transformed from the original character sequence into a structured text with word boundaries, part-of-speech information, and dependency relationships, providing input for subsequent identification of implicit semantic errors at the dependency relationship level.

[0059] S2. Based on the dependency relation set, determine the candidate dependency relations containing the governing word, the governed word, and the dependency relation type, and extract the local semantic fragments corresponding to the candidate dependency relations;

[0060] Specifically, computer devices traverse dependency sets Each dependency in the table; for any dependency... If the dependency relationship type belongs to the dependency relationship type that requires semantic role adaptation judgment, then the dependency relationship is determined as a candidate dependency relationship; the dependency relationship types that require semantic role adaptation judgment include subject-predicate relationship, verb-object relationship, attributive-head relationship, adverbial-head relationship, prepositional-object relationship and supplementary relationship.

[0061] In this embodiment, candidate dependency relations refer to dependency relations selected from the dependency relation set that need to be further judged to determine whether there is a semantic role mismatch; a candidate dependency relation does not mean that the dependency relation has an error, but only that the dependency relation has the conditions to be further calculated for surface fluency confidence value and semantic role fit.

[0062] For candidate dependency relations, the computer device extracts the corresponding local semantic fragments; in this embodiment, a local semantic fragment refers to the smallest semantic judgment fragment extracted around the candidate dependency relation, which is used to express the semantic relationship formed by the governing word, the governed word and its adjacent modifying components within a local range; the local semantic fragment is not a whole sentence, nor a single governing word or governed word, but is used to judge whether the candidate dependency relation is superficially fluent and whether there are text fragments with mismatched semantic roles.

[0063] For the same Chinese text to be corrected, the computer device obtains a set of candidate dependency relations:

[0064] ;

[0065] in, Represents the set of candidate dependencies. This indicates the number of candidate dependency relations involved in the relativization process within the same Chinese text to be corrected. Indicates the first The dominant word in the candidate dependency relation, Indicates the first The subordinate word in the candidate dependency relation, Indicates the first The dependency types of the candidate dependencies, Indicates the first Local semantic fragments corresponding to candidate dependency relations;

[0066] Local semantic fragments Dominant word in candidate dependency relation Subordinate words and located in the governing word and the governed word The local semantic segment is composed of modifiers in positions between or adjacent to the dominant word and the dominated word; if there are no other words between the dominant word and the dominated word, the local semantic segment is composed of the dominant word and the dominated word; if there are adverbs, attributives, complements or degree adverbs in the dominant word or the dominated word, the modifiers that are directly adjacent to the candidate dependency relationship and affect the semantic judgment will be incorporated into the local semantic segment.

[0067] For example, for "repair risk", the local semantic fragment is "repair risk"; for "significantly increase difficulty of use", the candidate dependency relation can be "increase - difficulty", and the local semantic fragment is "significantly increase difficulty of use"; for "convene a new test plan", the candidate dependency relation can be "convene - plan", and the local semantic fragment is "convene a new test plan".

[0068] The output of this step includes a set of candidate dependency relations and a local semantic fragment corresponding to each candidate dependency relation.

[0069] S3. Determine the surface fluency confidence value of candidate dependencies based on local semantic fragments;

[0070] In this embodiment, the surface fluency confidence value refers to the fluency of a local semantic segment at the linguistic level relative to other local semantic segments in the same Chinese text to be corrected. The surface fluency confidence value is used to characterize the naturalness of a local semantic segment in terms of characters, word order, and conventional language expression, but does not directly indicate that the local semantic segment is necessarily correct in terms of semantic role. In other words, a high surface fluency confidence value only indicates that the local semantic segment is relatively fluent on the surface, and does not exclude the possibility that there is a mismatch between the semantic roles of the governing word and the governed word.

[0071] Specifically, for each candidate dependency Corresponding local semantic fragments Computer equipment calculates the language fluency loss of local semantic segments based on Chinese language models;

[0072] In this embodiment, the Chinese language model is a pre-trained Chinese autoregressive language model based on correct Chinese corpus. The input of the Chinese autoregressive language model is the word sequence in a local semantic segment, and the output is the conditional probability of each word in the local semantic segment under the condition of its preceding words. The correct Chinese corpus is Chinese text corpus that has been manually reviewed, screened by publicly available standard corpus, or confirmed through domain text review processes. The correct Chinese corpus does not contain text segments marked as having typos, grammatical errors, or semantic collocation errors. For any local semantic segment, the computer device, according to the order of words in the local semantic segment, takes the preceding words as conditional input and the current word as the prediction object to obtain the conditional probability corresponding to the current word, and calculates the language fluency loss of the local semantic segment accordingly. Thus, the input of the language fluency loss is the word sequence of the local semantic segment, and the output is the numerical fluency level of the local semantic segment.

[0073] In another implementation, the Chinese language model is an N-gram language model obtained based on the statistics of correct Chinese corpus. The computer device calculates the conditional probability of the current word based on one or more words preceding the current word, and obtains the language fluency loss of the local semantic segment according to the same language fluency loss calculation method. Regardless of whether a Chinese autoregressive language model or an N-gram language model is used, its output is the conditional probability of each word in the local semantic segment. The conditional probability is used to subsequently determine the surface fluency confidence value.

[0074] If local semantic fragment Depend on Composed of several words, denoted as:

[0075] ;

[0076] in, Represents the first in a local semantic segment One word, Indicates the number of words in a local semantic segment; local semantic segment Language fluency loss for:

[0077] ;

[0078] in, Indicates the preceding words Generate words under known conditions The conditional probability; the language fluency loss is the negative log-likelihood form commonly used in natural language processing. The smaller the value, the more fluent the local semantic fragment is in the language model;

[0079] This embodiment performs relativization processing based on candidate dependency relations in the same Chinese text to be corrected; the computer device counts the language fluency loss corresponding to all candidate dependency relations in the same Chinese text to be corrected, and obtains the mean language fluency loss. and standard deviation of language fluency loss ;in:

[0080] ;

[0081] ;

[0082] in, This indicates the number of candidate dependency relations involved in relativization processing within the same Chinese text to be corrected; when When the relative fluency is zero, it indicates that the language fluency loss of each candidate dependency relation in the same Chinese text to be corrected is the same. In this case, the relative fluency of each candidate dependency relation is set to zero, and all candidate dependencies are treated as surface fluency-consistent candidate dependencies and entered into the subsequent semantic role adaptation judgment; when When the value is not zero, the relative fluency of candidate dependencies for:

[0083] ;

[0084] in, Indicates candidate dependency relationship Surface fluency confidence value; due to language fluency loss A smaller value indicates a smoother flow of the local semantic segment, therefore The larger the value, the smoother the semantic segment is relative to other semantic segments in the same Chinese text to be corrected.

[0085] When the surface fluency confidence value of a candidate dependency relation is higher than the average level of surface fluency confidence values ​​of candidate dependencies in the same Chinese text to be corrected, the computer device identifies the candidate dependency relation as a high-confidence surface fluency candidate dependency relation. When the language fluency loss of all candidate dependencies in the same Chinese text to be corrected is the same, the computer device identifies all candidate dependencies as consistent surface fluency candidate dependencies and proceeds to the subsequent semantic role adaptation judgment. In this embodiment, a high-confidence surface fluency candidate dependency relation refers to a candidate dependency relation whose local semantic segments are at a relatively high level in the language fluency judgment, and a consistent surface fluency candidate dependency relation refers to a candidate dependency relation that cannot be further distinguished by surface fluency confidence value when the language fluency loss is the same.

[0086] For example, in the sentence "The system can automatically identify anomalies and repair risks", although "repair risks" has a semantic role mismatch, as a local phrase, its language model fluency loss may not be high, so it is identified as a surface fluency high-confidence candidate dependency. This embodiment does not regard low fluency segments as the only source of error, but specifically retains surface fluency high-confidence segments so as to identify semantic role mutual exclusion hidden in fluent expressions in the future.

[0087] S4. Determine the semantic role fit of candidate dependency relations based on the governing word, the governed word, and the dependency relation type;

[0088] In this embodiment, semantic category refers to the category to which a word belongs in terms of semantic function. It is used to characterize the type of object, action, attribute, or event that the word can assume in a sentence. For example, "vulnerability, defect, malfunction, error" can be classified into the category of repairable objects; "risk, cost, error, loss, complexity" can be classified into the category of objects that can be reduced; "meeting, symposium, hearing, review meeting" can be classified into the category of events that can be held; and "plan, scheme, system, process" can be classified into the category of text plan objects. Semantic categories can be obtained from statistical results or semantic vector clustering results in a general Chinese semantic knowledge base, a domain terminology database, or a correct Chinese corpus.

[0089] Specifically, regarding candidate dependencies Computer devices are based on the dependency type of candidate dependencies. Determine the governing word For the governed word The semantic slots of dependency roles; the semantic slots of dependency roles represent the semantic category of the dependent word required by the governing word under a certain dependency relationship type, or the matching requirements between the semantic category of the governing word and the semantic category of the dependent word formed after replacing the governing word.

[0090] In this embodiment, the dependency role semantic slot is not a single word, nor is it a simple word co-occurrence relationship. Instead, it is a semantic constraint structure jointly defined by the governing word, the dependency relationship type, and the allowed semantic category of the governed word. For verb-object relationships, the dependency role semantic slot is used to indicate what semantic category of object a certain verb can accept when it is a governing word. For attributive-head relationships, the dependency role semantic slot is used to indicate the attribute matching relationship between the modifier and the head word. For adverbial-head relationships, the dependency role semantic slot is used to indicate the semantic matching relationship between the adverbial component and the action or state.

[0091] In this embodiment, the dependency role semantic slots are constructed as follows: The computer device pre-acquires correct Chinese corpus and performs word segmentation, part-of-speech tagging, and dependency parsing on the correct Chinese corpus to obtain a set of correct corpus dependency relations; for each dependency relation in the set of correct corpus dependency relations, the computer device extracts the governing word, dependency relation type, and governed word to form a correct dependency triple; the computer device maps the governed word in the correct dependency triple to the governed word semantic category, and establishes a correspondence between the governing word, dependency relation type, and governed word semantic category in the correct dependency triple;

[0092] For the same governing word and the same dependency relation type, the computer device counts the occurrence frequency of each corresponding governed word semantic category and forms a governed word semantic category distribution based on the occurrence frequency; this governed word semantic category distribution is the dependency role semantic slot of the governing word under the dependency relation type; in other words, the dependency role semantic slot is stored in the form of "governing word - dependency relation type - governed word semantic category distribution", which is used to represent the range of governed word semantic categories that a specific governing word can accept under a specific dependency relation type;

[0093] For example, regarding the governing word "repair" and the dependency relationship type "verb-object relationship," the governed words that form a verb-object relationship with "repair" in correct Chinese corpora include "loophole," "defect," "fault," "error," "file," and "system," etc. Computer equipment maps these governed words to categories of repairable or recoverable objects, and determines the semantic category distribution as the dependency role semantic slot for "repair" under the verb-object relationship. Similarly, regarding the governing word "reduce" and the dependency relationship type "verb-object relationship," the governed words that form a verb-object relationship with "reduce" in correct Chinese corpora include "risk," "cost," etc. The computer device maps the aforementioned governed words such as "cost, error, loss, and complexity" to categories that can be reduced, and determines the distribution of this semantic category as the semantic slot of the dependency role under the verb-object relationship. For the governed word "hold a meeting" and the dependency relationship type "verb-object relationship", the governed words that form a verb-object relationship with "hold a meeting" in the correct Chinese corpus include "meeting", "symposium", "hearing", "review meeting", etc. The computer device maps the aforementioned governed words to categories of events that can be held, and determines the distribution of this semantic category as the semantic slot of the dependency role under the verb-object relationship.

[0094] When there is no slot record in the dependency role semantic slot table indexed by the governing word and the dependency relation type, or when the number of correct dependency triples corresponding to the governing word under the dependency relation type is zero, the computer device maps the governing word to a governing word semantic category and performs backtracking statistics in the form of "governing word semantic category - dependency relation type - governed word semantic category distribution". For example, when there are insufficient samples of a specific verb, the computer device maps it to one of the following: disposal action, repair action, reduction action, generation action, or convening action, and determines the dependency role semantic slot based on the correct dependency triples of other words under the same governing word semantic category.

[0095] Dependency role semantic slots can be obtained in at least one of the following ways and stored accordingly in computer devices;

[0096] The first approach is based on a general Chinese semantic knowledge base. The computer device maps words to semantic categories. For example, "vulnerability, defect, malfunction, error" are classified as repairable objects, "risk, cost, loss, error, complexity" are classified as objects that can be reduced, "meeting, symposium, hearing, review meeting" are classified as events that can be held, and "plan, scheme, document, report" are classified as text plan objects. The computer device further establishes the correspondence between dominant words, dependency relationship types, and allowed semantic categories.

[0097] The second method is based on a domain terminology database. When the Chinese text to be corrected belongs to a technical document, patent text, legal text, business report, or other professional text, the computer device obtains domain terms and their semantic categories from the domain terminology database. For example, in the patent text domain, "claims, description, abstract, and embodiments" are classified as patent document objects; "limitation, recording, disclosure, and support" are classified as text processing actions. This allows for the determination of semantic adaptation differences between expressions such as "sufficient disclosure effect" and "sufficient disclosure content."

[0098] The third method is based on the statistical results of dependency collocations in the correct corpus; computer equipment pre-processes the correct Chinese corpus for word segmentation, part-of-speech tagging, and dependency parsing, extracts the dependency collocation relationships in the correct corpus, and statistically analyzes them in the governing words. Dependency Relationship Types Given the semantic category of the governed word The conditional probability of occurrence; for candidate dependencies If the governed word The semantic category is Then semantic role fit It can be represented as:

[0099] ;

[0100] in, Indicates candidate dependency relationship Semantic role fit Indicating the subject of a word semantic categories, This indicates that the governing word in the correct corpus is And the dependency relationship type is When the semantic category of the governed word is The conditional probability;

[0101] Specifically, The data was obtained from the dependency collocation statistics table in the correct Chinese corpus; its denominator is the number of governing words in the correct Chinese corpus. And the dependency relationship type is The total number of all correct dependency triples, whose numerator is the semantic category of the governed word in all the above correct dependency triples. The number of correct dependency triples; if the denominator is not zero, the computer determines the semantic role fit based on the ratio of the numerator to the denominator. If the denominator is zero, it indicates that the governing word in the correct Chinese corpus is... In dependency relationship types There is no corresponding correct dependency triple; the computer device uses the semantic category of the dominator word. Perform backtracking statistics; if the denominator is not zero but the numerator is zero, it indicates that the semantic category of the governed word was not observed in the correct Chinese corpus. With governing words In dependency relationship types Once a correct match is formed, the computer device records the semantic role fit of the candidate dependency relationship as zero, and determines whether it constitutes a mutually exclusive dependency role semantic relationship through subsequent relativization processing within the same Chinese text to be corrected.

[0102] If the governed word It can correspond to multiple semantic categories. The computer device calculates the semantic role fit for each semantic category and takes the highest semantic role fit as the semantic role fit for the candidate dependency relationship; if the dominant word It can correspond to multiple semantic categories of governing words. The computer device performs backtracking statistics for each semantic category of governing words and takes the semantic category of governing words that yields the highest semantic role fit as the backtracking statistics result. This processing method is used to avoid misjudgment caused by polysemous words.

[0103] When using the semantic category of the dominant word for backtracking statistics, the computer device calculates the semantic role fit based on the semantic category of the dominant word:

[0104] ;

[0105] in, Indicating governing words semantic categories, This indicates that the semantic category of the governing word in the correct corpus is And the dependency relationship type is When the semantic category of the governed word is The conditional probability.

[0106] The fourth method is based on semantic vector clustering results; the computer device converts words or phrases in the correct corpus into semantic vectors, and clusters the semantic vectors, using the clustering results as the semantic categories of words; then, the computer device calculates the matching relationship between the semantic categories of the dominant word, the type of dependency relationship, and the semantic categories of the governed word based on the dependency relationships in the correct corpus.

[0107] After obtaining the dependent role semantic slot, the computer device acquires the dominated word. semantic categories Furthermore, based on the degree of matching between the semantic category of the governed word and the semantic slot of the dependency role, the semantic role fit of the candidate dependency relationship is determined. In this embodiment, semantic role fit refers to the degree of matching between the semantic category of the governed word in the candidate dependency relation and the semantic slot of the dependency role of the governing word under the corresponding dependency relation type. The higher the semantic role fit, the more the governing word and the governed word conform to the semantic role requirements of correct Chinese expression under the dependency relation type. The lower the semantic role fit, the more likely there is a semantic role mismatch in the candidate dependency relation. Semantic role fit is not the same as word frequency or sentence fluency. Its evaluation object is the semantic role fit relation within the candidate dependency relation.

[0108] For example, in the candidate dependency relation "repair-risk", the governing word is "repair" and the governed word is "risk", and the dependency relation type is verb-object relation. The semantic slots of the dependency role of "repair" under the verb-object relation usually include repairable objects such as vulnerabilities, defects, faults, and errors, while "risk" belongs to objects that can be reduced. Therefore, the semantic category of "risk" does not match the semantic slots of the dependency role of "repair", and the semantic role fit of this candidate dependency relation is low.

[0109] For example, in the candidate dependency relationship “hold a meeting-plan”, the governing word is “hold a meeting” and the governed word is “plan”, and the dependency relationship type is verb-object relationship. The semantic slots of the object of “hold a meeting” usually include event-type objects such as meetings, seminars, and hearings, while “plan” belongs to text plan objects. Therefore, the semantic role fit of this candidate dependency relationship is low.

[0110] To facilitate execution by computer devices, the dependency role semantic slots in this embodiment can be stored in a table structure. This table structure includes at least a governing word field, a governing word semantic category field, a dependency relationship type field, a allowed governed word semantic category field, a slot word field, and a frequency field. Specifically, the governing word field records the specific governing word; the governing word semantic category field is used for backtracking statistics when there are insufficient samples of the specific governing word; the dependency relationship type field records subject-verb, verb-object, attributive-head, adverbial-head, prepositional-object, or supplementary relationships; the allowed governed word semantic category field records the semantic categories of governed words that the governing word can accept under the dependency relationship type; the slot word field records the specific words belonging to the allowed governed word semantic category; and the frequency field records the number of times the semantic category or the slot word appears in the correct Chinese corpus.

[0111] For example, the dependency role semantic slot table may include the following records: "Repair - verb-object relationship - repairable object category - vulnerability, defect, fault, error"; "Reduce - verb-object relationship - object category that can be reduced - risk, cost, error, loss, complexity"; "Hold - verb-object relationship - event category that can be held - meeting, seminar, hearing, review meeting"; "Formulate - verb-object relationship - text plan object category - scheme, plan, system, process"; When judging candidate dependencies, the computer device queries the table structure based on the governing words and dependency relationship type in the candidate dependencies, and matches the semantic categories of the allowed governed words found in the query with the semantic categories of the governed words in the candidate dependencies.

[0112] S5. Determine the semantic mutual exclusion relationship of dependent roles based on the divergence between surface fluency confidence value and semantic role fit.

[0113] In this embodiment, the degree of divergence refers to the degree of inconsistency between the surface fluency confidence value and the semantic role fit of the candidate dependency relationship; the degree of divergence is used to describe a special phenomenon: local semantic segments are relatively fluent in linguistic form, but the dominating word and the dominated word in the local semantic segment do not match in the semantic slot of the dependency role; the higher the degree of divergence, the closer the candidate dependency relationship is to the semantic mutual exclusion feature of dependency roles under high confidence of surface fluency.

[0114] Specifically, the computer device selects candidate dependencies from surface fluent high-confidence candidate dependencies or surface fluent consistent candidate dependencies whose semantic role fit is lower than the average level of candidate dependencies in the same Chinese text to be corrected.

[0115] To perform relativistic judgment, the computer device statistically analyzes the semantic role fit of all candidate dependency relations in the same Chinese text to be corrected, and obtains the mean semantic role fit. Standard deviation of semantic role fit ;in:

[0116] ;

[0117] ;

[0118] in, This indicates the number of candidate dependency relations involved in relativization processing within the same Chinese text to be corrected; when When the value is zero, it indicates that the semantic role fit of each candidate dependency relation in the same Chinese text to be corrected is the same. In this case, the semantic role fit standard value of each candidate dependency relation is set to zero; when When the value is not zero, the semantic role adaptation standard value of the candidate dependency relation for:

[0119] ;

[0120] in, The larger the value, the higher the semantic role fit of the candidate dependency relation relative to other candidate dependencies in the same Chinese text to be corrected. The smaller the value, the lower the semantic role fit of the candidate dependency relationship;

[0121] The computer device further calculates the degree of deviation between the surface fluency confidence value and the semantic role fit. :

[0122] ;

[0123] in, Indicates candidate dependency relationship The degree of deviation; The larger the value, the more the candidate dependency relationship conforms to the characteristics of high surface fluency confidence and low semantic role fit;

[0124] In this embodiment, a dependency role semantic mutual exclusion relationship refers to a candidate dependency relationship that satisfies the surface fluency high confidence condition or the surface fluency consistency condition, and whose semantic role fit is lower than the average level of candidate dependency relationships in the same Chinese text to be corrected. The determination of dependency role semantic mutual exclusion relationship is not based on the overall incoherence of the sentence, but on the deviation between the local expression fluency and the dependency role misfit.

[0125] A computer device determines a candidate dependency relationship as a mutually exclusive dependency role relationship when the following conditions are met:

[0126] First, the candidate dependency relationship belongs to the surface fluency high-confidence candidate dependency relationship, or it belongs to the surface fluency consistent candidate dependency relationship when all candidate dependencies in the same Chinese text to be corrected have the same language fluency loss.

[0127] Second, the semantic role fit of this candidate dependency relation is lower than the average level of candidate dependencies in the same Chinese text to be corrected;

[0128] Third, the deviation of this candidate dependency relation is the highest among the candidate dependency relations in the current sentence;

[0129] If there are two or more candidate dependency relations with the same degree of divergence in the same sentence, and both are the highest, the computer selects the candidate dependency relations in order of semantic role fit from low to high. If the semantic role fit is still the same, the candidate dependency relations are selected in order of local semantic segment length from short to long. If the local semantic segment length is still the same, the candidate dependency relations are selected according to the order in which they appear in the sentence. The above processing is used to obtain a unique mutually exclusive dependency role semantic relation when there are two parallel highest degrees of divergence, so as to avoid the problem of non-unique input in the subsequent error-bearing word determination step.

[0130] If all candidate dependency relations in the same Chinese text to be corrected have the same semantic role fit, the computer will not identify any candidate dependency relation as a mutually exclusive dependency relation and will output that there is no verifiable mutually exclusive dependency relation in the sentence. This processing method is used to avoid forced error correction when there is a lack of relative differences in semantic role fit.

[0131] The above judgment method is based on relative judgment of candidate dependency relationships in the same Chinese text to be corrected, and can be adapted to Chinese texts of different lengths, fields and styles of expression.

[0132] For example, in the sentence "The system can automatically identify anomalies and repair risks", the surface fluency confidence value of the local semantic fragment "repair risks" is high, but the semantic role fit of the verb-object relationship "repair-risk" is low. Therefore, its degree of deviation is high, and it is identified as a semantically exclusive relationship of dependent roles. This result shows that the fragment is not identified because it is not fluent, but because there is a significant deviation between its surface fluency and semantic role misfit.

[0133] S6. Determine the error-bearing word based on the semantic mutual exclusion relationship of dependency roles, generate structure-preserving error-correction candidate words, determine the target error-correction word from the structure-preserving error-correction candidate words, and output the error correction result after replacing the error-bearing word with the target error-correction word.

[0134] In this embodiment, an erroneous term refers to a word that is determined to need to be replaced in a semantically exclusive dependency relationship. An erroneous term can be either a subordinate term or a dominant term, and its specific nature is determined by whether the semantic role fit is improved after replacement. An erroneous term does not mean that the term itself is not a legitimate Chinese term, but rather that the semantic role that the term assumes in the current candidate dependency relationship does not match the corresponding dependency role semantic slot.

[0135] Specifically, for the mutually exclusive semantic relationships of dependency roles determined by S5 The computer device determines the error-bearing word based on the dependency relationship type; the error-bearing word indicates the word that needs to be replaced in the semantic mutual exclusion relationship of the dependency role;

[0136] When the dependency relationship type of the mutually exclusive dependency role is a verb-object relationship, a prepositional phrase relationship, or a complementary relationship, the computer device will first select the governed word. The original term is identified as the erroneous term to be verified, and candidate terms corresponding to the erroneous term are generated. These candidate terms are generated from the dependency role semantic slot table. The computer device first queries the dependency role semantic slot table based on the original erroneous term and the original dependency relationship type to obtain the allowed semantic categories of the erroneous term under the original dependency relationship type. Then, specific words are extracted from the slot word field corresponding to the allowed semantic categories of the erroneous term. Subsequently, words that are consistent with the part of speech of the original erroneous term, can be in the syntactic position of the original erroneous term, and do not change the original dependency relationship type after replacement are retained as candidate terms corresponding to the erroneous term. If there are similar words in the domain terminology library in the slot word field, words that are consistent with the domain of the Chinese text to be corrected are retained first. If there is no corresponding slot record in the dependency role semantic slot table for the original erroneous term, the computer device queries the corresponding fallback slot based on the semantic category of the erroneous term and generates candidate terms corresponding to the erroneous term in the same way.

[0137] The computer device replaces the original governed word with the candidate word corresponding to the governed word, and calculates the semantic role fit after replacement. If the highest semantic role fit after replacement of the candidate word corresponding to the governed word is higher than the semantic role fit before replacement, the governed word is identified as the wrongly assigned word. If the highest semantic role fit after replacement of the candidate word corresponding to the governed word is not higher than the semantic role fit before replacement, the assigned word is identified as the wrongly assigned word.

[0138] For example, regarding "repairing risks," if the context includes phrases like "vulnerability, defect, security gap," then when "risk" is used as the error-bearing word to be verified, candidate words such as "vulnerability, defect, fault" can be generated. Replacing it with "repair vulnerabilities" improves the semantic role fit, thus "risk" is identified as the error-bearing word. However, if the context emphasizes "risk management, risk reduction, and risk avoidance," then replacing "risk" does not yield a higher fit. In this case, the computer device identifies "repair" as the error-bearing word and generates candidate words such as "reduce, control, and avoid."

[0139] In the context of replacing dominant words, the computer device no longer uses the semantic slot of the dependency role corresponding to the original dominant word as the target slot. Instead, it uses the semantic category of the original dominated word as the retrieval condition and searches the dependency role semantic slot table for dominant words that can accept the semantic category of the dominated word. For example, if the semantic category of the original dominated word "risk" is a category of objects that can be reduced, the computer device searches the dependency role semantic slot table for dominant words that allow the semantic category of the dominated word to be a category of objects that can be reduced under the verb-object relationship, and obtains candidate dominant words such as "reduce," "control," "avoid," and "reduc." The computer device constructs candidate dependency relationships after replacement, such as "reduce risk," "control risk," "avoid risk," and "reduce risk," and recalculates the semantic role fit and language fluency loss of each candidate dependency relationship to determine the target correction word.

[0140] When the dependency relationship type of the mutually exclusive dependency role is a modifier-head relation or an adverbial-head relation, the computer device judges the semantic role fit after replacing the governing word and the governed word, and identifies the side that improves the semantic role fit as the incorrect bearing word. For example, in a modifier-head relation such as "heavy efficiency", if "heavy" as a modifier cannot form a reasonable attribute relationship with "efficiency", and replacing the modifier with "higher" can improve the fit, then "heavy" is identified as the incorrect bearing word. In an adverbial-head relation such as "complete quickly and difficult", if there is a semantic mismatch between the adverbial and the action or state, the incorrect bearing word is determined according to the fit after replacement.

[0141] In this embodiment, structure-preserving error correction candidate words refer to candidate words that can replace the error-bearing word and maintain the original dependency relationship type after replacement; structure-preserving error correction candidate words need to simultaneously meet the requirements of semantic role adaptation, local language fluency and syntactic position preservation; they are used to distinguish them from whole-sentence rewriting candidate results and avoid extending the semantic error correction process into free polishing or rewriting of the Chinese text to be corrected;

[0142] After identifying the erroneous bearer, the computer device generates a set of candidate words that occupy the same syntactic position as the erroneous bearer. Specifically, the computer device determines the target semantic slot based on the mutually exclusive dependency role semantic relationship of the erroneous bearer. If the erroneous bearer is a subordinate word, the target semantic slot is the dependency role semantic slot corresponding to the original subordinate word under the original dependency relationship type. The computer device extracts candidate words from the slot word field corresponding to the target semantic slot and retains candidate words with the same part of speech as the erroneous bearer, or retains candidate words that remain in the original subordinate word position after dependency parsing and whose dependency relationship type remains unchanged. If the erroneous bearer is a dominant word, the target semantic slot is the dominant word semantic category or set of dominant words that can form a matching relationship with the semantic category of the original subordinate word. The computer device retrieves dominant words that can accept the semantic category of the original subordinate word from the dependency role semantic slot table and retains candidate words with the same part of speech as the erroneous bearer, or retains candidate words that remain in the original dominant word position after dependency parsing and whose dependency relationship type remains unchanged.

[0143] After generating the candidate word set, the computer device replaces the erroneous word with the candidate word and re-executes dependency parsing or local dependency relation verification. If the replaced word, the governed word, and the dependency relation type can still form a candidate dependency relation consistent with the original dependency relation type, the candidate word is retained. If the dependency relation type changes after replacement, or the replaced local semantic fragment cannot form a stable candidate dependency relation, the candidate word is deleted. Thus, the generation process of structure-preserving error correction candidate words does not rely on subjective human selection, but is completed jointly based on the dependency role semantic slot table, part-of-speech consistency, dependency relation type preservation, and semantic role fit recalculation.

[0144] Each candidate word in the candidate word set must be able to replace the erroneous word and maintain the basic structure of the original candidate dependency relationship in the sentence; the computer device filters the candidate word set and retains the candidate words that meet the following conditions to obtain structure-preserving error-correcting candidate words;

[0145] First, the type of dependency relationship remains unchanged after replacement; that is, if the original dependency relationship is a verb-object relationship, the verb-object relationship is retained after replacement; if the original dependency relationship is a noun-head relationship, the noun-head relationship is retained after replacement.

[0146] Second, the replaced candidate dependency relations satisfy the corresponding dependency role semantic slots; that is, regardless of whether the replaced word is the dominant word or the dominated word, the replaced dominant word, dominated word and dependency relation type can form a semantic role adaptation relationship.

[0147] Third, the language fluency loss of the replaced local semantic fragment is no higher than that of the original local semantic fragment; this condition is used to prevent the sentence from becoming awkward even though the semantic role fit is improved.

[0148] Fourth, the replacement scope is limited to the word bearing the error, and the entire sentence is not freely rewritten; this condition is used to keep the modifications to a minimum and avoid turning semantic error correction into sentence polishing or rewriting.

[0149] After obtaining the structure-preserving error correction candidate words, the computer device sorts the candidate words according to the following rules: first, sorting them from high to low semantic role fit; if the semantic role fit is the same, sorting them from low to high language fluency loss; and if both semantic role fit and language fluency loss are the same, sorting them from small to large edit distance from the error-bearing word. Edit distance is used to measure the difference between the candidate words and the error-bearing word in terms of character-level substitution, insertion, and deletion. The smaller the edit distance, the smaller the change to the original text.

[0150] After sorting, the computer device determines the structure-preserving error correction candidate word with the highest ranking result as the target error correction word, and replaces the error-bearing word with the target error correction word to obtain the error correction result. In this embodiment, the target error correction word refers to the final replacement word determined by sorting from the structure-preserving error correction candidate words. The target error correction word is used to directly replace the error-bearing word and form the error correction result. The determination order of the target error correction word takes into account semantic role fit, language fluency loss and edit distance in turn, so as to ensure that the local semantic fragment after replacement satisfies the semantic slot of the dependency role, maintains the fluency of local expression, and minimizes the modification of the original text.

[0151] For example, the original sentence was: "The system can automatically identify anomalies and repair risks;"

[0152] After processing through S1 to S5, the semantic relationship between "repair" and "risk" is determined to be mutually exclusive. If the context contains information such as "system vulnerability," "security flaw," or "abnormal patch," the candidate word set includes "vulnerability," "flaw," and "fault." After replacement, the semantic role fit of "repair vulnerability" is higher than that of "repair risk," and the language fluency loss of the local semantic fragment is no greater than that of the original local semantic fragment. Therefore, "vulnerability" can be used as a structure-preserving error correction candidate word. If its ranking result is the highest, the target error correction word is "vulnerability," and the error correction result is: "The system can automatically identify anomalies and repair vulnerabilities."

[0153] For example, the original sentence was: "This method can increase the difficulty for users in complex scenarios;"

[0154] After processing, the semantic relationship between "improvement" and "difficulty" was determined to be mutually exclusive. The verb-object semantic slot of "improvement" typically corresponds to objects that can be improved, such as efficiency, accuracy, ability, and quality, while "difficulty" typically corresponds to objects that can be reduced or overcome. If the context emphasizes reducing usage barriers, the computer device can identify the governing word "improvement" as the error-bearing word and generate candidate words such as "reduce," "reduce," and "alleviate." After replacement, "reduce usage difficulty" satisfies the corresponding dependent role semantic slot and maintains the verb-object relationship, thus the error correction result can be output: "This method can reduce the difficulty of use for users in complex scenarios;"

[0155] For example, the original sentence was: "The project team convened a meeting to discuss the new testing plan;"

[0156] After processing, it was determined that "convenient-plan" has a mutually exclusive semantic relationship of dependent roles; the object semantic slot of "convenient" is usually a meeting-type event, and "plan" belongs to the text plan object; if there is no "meeting" object in the context but the emphasis is on the formation of the plan, the computer device will identify the governing word "convenient" as the erroneous word and generate candidate words such as "formulate," "formulate," and "review"; if the semantic role fit is the highest after replacing "formulate" and the loss of language fluency is not higher than that of the original local semantic fragment, the error correction result is output: "The project team has formulated a new test plan;".

[0157] When S6 outputs the error correction results, the computer device also outputs interpretable error correction information simultaneously; the interpretable error correction information includes the original dependency relation, dependency relation type, governing word, governed word, fault-bearing word, semantic category of fault-bearing word, semantic slot of dependency role, target error correction word, and whether the original dependency relation type is maintained after replacement;

[0158] In this embodiment, explainable error correction information refers to structured explanatory data that is output synchronously with the error correction result, which is used to explain the basis for the generation of the error correction result. Explainable error correction information does not change the error correction result itself. Its role is to indicate why the original candidate dependency relationship was determined to be a mutually exclusive dependency role relationship, and why the target error correction word can replace the error-bearing word. Through explainable error correction information, users can confirm that the error correction basis comes from the dependency role semantic slot and semantic role fit, rather than just from word frequency or language model fluency.

[0159] Taking "remediation of risks" as an example, the corrective information that can be explained includes:

[0160] Original dependency relationship: "Repair-Risk";

[0161] Dependency relationship type: verb-object relationship;

[0162] Governing word: “repair”;

[0163] Subjective word: "risk";

[0164] Incorrectly assigned word: "risk";

[0165] Semantic categories of erroneous words: can reduce object categories;

[0166] Dependency role semantic slots: "Repair" in a verb-object relationship requires the category of the object that can be repaired;

[0167] Target error correction keyword: "vulnerability";

[0168] Does the original dependency relationship type remain after replacement? Yes, the relationship remains verb-object after replacement.

[0169] Through the aforementioned interpretable error correction information, users can clearly understand that the basis for error correction is not simply sentence fluency, but rather the discrepancy between the surface fluency confidence value and the semantic role fit.

[0170] To further illustrate the feasibility of this embodiment, a set of exemplary data is provided below; this data is used to illustrate how a computer device determines the semantic mutual exclusion relationship of dependent roles based on surface fluency confidence value, semantic role fit, and degree of deviation, and is not intended to limit the scope of protection of this invention;

[0171] Repair risks Repair - Risk Verb-object relationship 0.62 0.08 -0.83 1.45 Dependent role semantic mutual exclusion Fix vulnerabilities Fix - Vulnerability Verb-object relationship 0.25 0.76 0.84 -0.59 Semantic adaptation Reduce risk Reduce risk Verb-object relationship 0.38 0.81 1.02 -0.64 Semantic adaptation Meeting Plan Convening—Plan Verb-object relationship 0.46 0.05 -0.76 1.22 Dependent role semantic mutual exclusion Hold a meeting Meeting Verb-object relationship 0.18 0.83 0.83 -0.65 Semantic adaptation Increase difficulty Improvement - Difficulty Verb-object relationship 0.58 0.07 -0.78 1.36 Dependent role semantic mutual exclusion Improve efficiency Improve efficiency Verb-object relationship 0.31 0.79 0.91 -0.60 Semantic adaptation

[0172] In the examples above, the surface fluency confidence values ​​of "fix risks," "convene a plan," and "improve difficulties" are all positive, indicating that their local expressions are not obviously awkward under the language model; however, their semantic role fit is significantly lower than the average level in the same Chinese text to be corrected, resulting in a high degree of deviation, and therefore they are identified as mutually exclusive dependent role semantic relationships. Conversely, the surface fluency confidence values ​​of "fix vulnerabilities," "reduce risks," "convene a meeting," and "improve efficiency" are all high in terms of semantic role fit and low in degree of deviation, and are therefore identified as semantic fit relationships. This example shows that this embodiment can identify implicit semantic errors that are fluent on the surface but do not match dependent roles, rather than simply judging all low-frequency expressions or all high-fluency expressions as errors.

[0173] Taking the Chinese text to be corrected, “The system can automatically identify anomalies and repair risks,” as an example, the computer device obtains the correction result according to the following process;

[0174] The computer device performs word segmentation, part-of-speech tagging, and dependency parsing on the sentence, and obtains the candidate dependency relation "repair-risk", where "repair" is the governing word and "risk" is the governed word, and the dependency relation type is verb-object relation;

[0175] The computer device extracts a local semantic segment called "repair risk" and calculates the language fluency loss of the local semantic segment using a Chinese language model. Then, it obtains the surface fluency confidence value based on the language fluency loss of each candidate dependency in the same Chinese text to be corrected. If the surface fluency confidence value of "repair risk" is higher than the average level of candidate dependencies in the same Chinese text to be corrected, it is regarded as a high-confidence candidate dependency in surface fluency.

[0176] The computer device queries the semantic slot of the dependency role under the verb-object relationship according to the dependency role semantic slot table, and obtains that its allowed subject semantic category is the category of repairable object. The slot words include "vulnerability, defect, fault, error", etc. The computer device maps "risk" to the category of object that can be reduced. It determines that the semantic category of "risk" does not match the semantic slot of the dependency role under the verb-object relationship of "repair", and obtains a low semantic role fit.

[0177] The computer device identifies "repair-risk" as a mutually exclusive dependency relationship based on the degree of deviation between the surface fluency confidence value and the semantic role fit. Since this dependency relationship is a verb-object relationship, the computer device first identifies the governed word "risk" as the error-bearing word to be verified, and generates candidate words such as "vulnerability, defect, fault, error" from the dependent role semantic slots of "repair" under the verb-object relationship. The computer device then constructs candidate dependency relationships with replacements such as "repair vulnerability", "repair defect", "repair fault", and "repair error", and recalculates the semantic role fit and language fluency loss.

[0178] If "fix vulnerabilities" has the highest semantic role fit, its language fluency loss is no greater than that of "fix risks," and its edit distance to "risk" is the smallest or highest-ranked among the candidate words that meet the aforementioned conditions, then the computer device will identify "vulnerabilities" as the target correction word and replace the erroneous word "risk" with "vulnerabilities," outputting the correction result "The system can automatically identify anomalies and fix vulnerabilities." At the same time, the computer device will output interpretable correction information, explaining that "fix-risk" has a verb-object relationship, the semantic category of "risk" is a category of objects that can be reduced, and it does not belong to the category of objects that can be patched under the verb-object relationship required by "fix," while the replaced "fix-vulnerabilities" maintains the verb-object relationship and satisfies the semantic slot of the dependent role.

[0179] In this embodiment, the input of the method is the Chinese text to be corrected, and the output is the correction result and interpretable correction information. The correction result is obtained by replacing the erroneous word with the target correction word, and the interpretable correction information is used to explain the original dependency relationship, the reason for semantic mutual exclusion, and the structural preservation after replacement. Thus, this embodiment realizes the identification and correction of implicit semantic errors that are legal in themselves, fluent in local expression, and have a stable syntactic structure but semantically mutually exclusive dependency roles. It can avoid missed detections caused by simply relying on misspelling detection, grammatical error detection, or overall language fluency judgment, and avoids excessive rewriting of the original text by screening structurally preserved candidate words.

[0180] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A computer processing method for semantic error correction of Chinese text, characterized in that, Includes the following steps: S1. Obtain the Chinese text to be corrected, and perform sentence segmentation, word segmentation, part-of-speech tagging and dependency parsing on the Chinese text to be corrected to obtain sentences, word sequences and dependency relation sets; S2. Based on the dependency relation set, determine the candidate dependency relations containing the governing word, the governed word, and the dependency relation type, and extract the local semantic fragments corresponding to the candidate dependency relations; S3. Determine the surface fluency confidence value of candidate dependencies based on local semantic fragments; S4. Determine the semantic role fit of candidate dependency relations based on the governing word, the governed word, and the dependency relation type; S5. Determine the semantic mutual exclusion relationship of dependent roles based on the divergence between surface fluency confidence value and semantic role fit. S6. Determine the error-bearing word based on the semantic mutual exclusion relationship of dependency roles, generate structure-preserving error-correction candidate words, determine the target error-correction word from the structure-preserving error-correction candidate words, and output the error correction result after replacing the error-bearing word with the target error-correction word.

2. The computer processing method for semantic error correction of Chinese text according to claim 1, characterized in that, The dependency relation set includes multiple dependency relations, each of which includes a governing word, a governed word, and a dependency relation type. The dependency relation type includes at least one of the following: subject-predicate relation, verb-object relation, attributive-head relation, adverbial-head relation, prepositional-object relation, and supplementary relation.

3. The computer processing method for semantic error correction of Chinese text according to claim 1, characterized in that, The local semantic fragment consists of the governing word, the governed word, and the modifying elements located between or adjacent to the governing word and the governed word in the candidate dependency relation.

4. The computer processing method for semantic error correction of Chinese text according to claim 1, characterized in that, The surface fluency confidence value for determining candidate dependencies based on local semantic fragments includes: Calculate the language fluency loss for local semantic segments; Based on the language fluency loss of each candidate dependency relation in the same Chinese text to be corrected, the relative fluency between each candidate dependency relation is determined; Candidate dependencies whose relative fluency is higher than the average level of candidate dependencies in the same Chinese text to be corrected are identified as surface fluency high-confidence candidate dependencies, and the relative fluency corresponding to the surface fluency high-confidence candidate dependencies is used as the surface fluency confidence value.

5. A computer processing method for semantic error correction of Chinese text according to claim 1, characterized in that, The semantic role fit of candidate dependency relations is determined based on the governing word, the governed word, and the dependency relation type, including: Based on the dependency relationship type of the candidate dependency relationship, determine the semantic slot of the dependency role of the governing word to the governed word; Obtain the semantic category of the governed word; The semantic role fit of candidate dependency relationships is determined based on the degree of matching between the semantic category of the governed word and the semantic slot of the dependency role.

6. The computer processing method for semantic error correction of Chinese text according to claim 5, characterized in that, The dependency role semantic slots are obtained from at least one of the following: a general Chinese semantic knowledge base, a domain terminology database, dependency collocation statistics from correct corpora, or semantic vector clustering results, and are stored corresponding to the semantic category of the governing word, dependency relationship type, and governed word.

7. A computer processing method for semantic error correction of Chinese text according to claim 4, characterized in that, Determining the semantic mutual exclusion relationship of dependency roles based on the divergence between surface fluency confidence value and semantic role fit includes: From the surface-level fluent high-confidence candidate dependency relations, select candidate dependency relations whose semantic role fit is lower than the average level of candidate dependency relations in the same Chinese text to be corrected; The selected candidate dependencies are ranked according to the degree of deviation between their surface fluency confidence value and semantic role fit. The candidate dependency relationship with the highest degree of deviation is identified as a semantic mutual exclusion relationship of dependency roles.

8. A computer processing method for semantic error correction of Chinese text according to claim 7, characterized in that, Error-bearing words are determined based on the semantic mutual exclusion relationship of dependency roles, including: When the dependency relationship type of the mutually exclusive relationship of dependency roles is verb-object relationship, prepositional object relationship or complementary relationship, the governed word is first identified as the error-bearing word to be verified, and the corresponding candidate word of the governed word is generated. If the highest semantic role fit of the candidate word after replacement is higher than the semantic role fit before replacement, the word being replaced is identified as the incorrectly assigned word. If the highest semantic role fit after replacing the candidate word corresponding to the governed word is not higher than the semantic role fit before replacement, the governing word will be identified as the erroneous bearer word. When the dependency relationship type of the mutually exclusive semantic relationship is a nominative-head relation or an adverbial-head relation, the semantic role fit after replacing the dominant word and the subordinate word is judged respectively, and the side that increases the semantic role fit is identified as the erroneous bearer word.

9. A computer processing method for semantic error correction of Chinese text according to claim 8, characterized in that, Generate structure-preserving error correction candidate words, and determine the target error correction word from the structure-preserving error correction candidate words, including: Generate a set of candidate words that are in the same syntactic position as the error-bearing word; From the candidate word set, retain candidate words that do not change the dependency relationship type after replacement, satisfy the corresponding dependency role semantic slot, and have a language fluency loss no greater than that of the original local semantic segment, to obtain structure-preserving error correction candidate words; The candidate words for structure-preserving error correction are ranked in order of semantic role fit from high to low, language fluency loss from low to high, and edit distance from the error-bearing word from small to large. The candidate word with the highest ranking is then selected as the target error correction word.

10. A computer processing method for semantic error correction of Chinese text according to claim 9, characterized in that, When outputting the error correction results, explainable error correction information is output simultaneously. Explainable error correction information includes the original dependency relation, dependency relation type, governing word, governed word, fault-bearing word, semantic category of fault-bearing word, semantic slot of dependency role, target error correction word, and whether the original dependency relation type is maintained after replacement.