A term proofreading method based on fuzzy matching and large model semantic discrimination

CN121724012BActive Publication Date: 2026-09-18TRS INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610202514.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-02-12
Publication Date
2026-09-18
Estimated Expiration
2046-02-12

AI Technical Summary

Technical Problem

[0003]为了解决现有技术中专业术语存在多样性、拼写错误、表述差异以及相似正确术语的干扰,导致对文本中的专业术语进行校对时误判率高和准确率低的问题,提出了一种基于模糊匹配与大模型语义判别的术语校对方法,解决了上述问题

Benefits of technology

[0061] This invention belongs to the field of natural language processing and proposes a terminology proofreading method based on fuzzy matching and large-scale model semantic discrimination. This method reduces the misjudgment rate of professional terminology proofreading and simultaneously improves the accuracy of the proofreading results. Specifically, in the coarse terminology query process, an n-gram retrieval engine and a prefix tree fuzzy retrieval engine are used to retrieve candidate words. These candidate words can effectively identify all possible terms in the proofreading text, achieving rapid screening and comprehensive coverage of terms in the proofreading text. In the fine terminology query process, by calculating the five-dimensional similarity between candidate words and matching targets extracted from the proofreading text, the terms and their positions in the proofreading text can be accurately determined. Complexity judgment rules are set based on five-dimensional similarity and boundary consistency, which can effectively and quickly determine the complexity of fine query terms. Different proofreading methods are selected according to the complexity of the fine query terms, improving the efficiency of terminology proofreading of the input proofreading text. When the complexity of the fine query terms is complex, terminology proofreading is performed using a large-scale model and the matched correct professional terms, i.e., candidate words, reducing the misjudgment rate of correct terms in the proofreading text that are not included in the professional terminology database. This invention combines coarse term query, fine term query, and large-scale model term verification, which not only achieves rapid term matching and accurate term location in the text being verified, but also reduces the misjudgment rate caused by similar correct terms, thereby improving the accuracy of term verification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121724012B_ABST
    Figure CN121724012B_ABST
Patent Text Reader

Abstract

The present application belongs to the field of natural language processing, and proposes a term proofreading method based on fuzzy matching and large model semantic discrimination. The n-gram search engine and the prefix tree fuzzy search engine are used for term retrieval, which can complement each other, improve the matching efficiency and recall rate; in the fine query stage, the five-dimensional similarity is obtained by weighting and summing the similarity of five dimensions, which can comprehensively quantify the matching degree of the candidate words in the candidate word set obtained by coarse query and the matching target in the proofreading text, effectively improve the term correction accuracy in complex error scenarios; using a large model and the correct professional terms matched to proofread the terms, which can reduce the misjudgment rate caused by similar correct terms, and improve the accuracy of term proofreading of the proofreading text.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of natural language processing technology, and in particular to a terminology proofreading method based on fuzzy matching and large model semantic discrimination. Background Technology

[0002] With the rapid development of information technology, the scale of text data is growing exponentially. In professional fields such as law, medicine, and academic research, the standardization and accuracy of terminology are crucial, directly affecting the reliability of information transmission and the scientific nature of decision-making. However, due to the diversity of terminology, spelling errors, differences in expression (such as confusion between homophones, misuse of similar-looking characters, and reversed order), as well as interference from similar correct terms, traditional rule-based or simple statistical terminology proofreading methods face significant challenges, making it difficult to efficiently and accurately identify and correct errors. Summary of the Invention

[0003] To address the issues of high misjudgment rates and low accuracy in text proofreading caused by the diversity, spelling errors, and discrepancies in terminology, as well as interference from similar correct terms, existing technologies propose a terminology proofreading method based on fuzzy matching and large-scale model semantic discrimination, which solves the aforementioned problems.

[0004] The technical solution is as follows:

[0005] A terminology collation method based on fuzzy matching and large-model semantic discrimination, the specific steps of which include:

[0006] S1: Construct a terminology database and search engine: Construct a terminology database, and based on the terminology database, construct an n-gram syntax search engine and a prefix tree fuzzy search engine;

[0007] S2: Coarse terminology search: The proofreading text is searched using an n-gram syntax search engine and a prefix tree fuzzy search engine to obtain a set of candidate words, which is the coarse terminology search result.

[0008] S3: Detailed Terminology Query: The proofreading text is segmented to obtain the matching target and term positions; the five-dimensional similarity between the matching target and the candidate words in the coarse term query results is calculated; the detailed query terms in the proofreading text are determined based on the five-dimensional similarity and the matching target; a term proofreading set is constructed based on the detailed query terms, candidate words, and five-dimensional similarity; the five-dimensional similarity is obtained by weighted summation of the pinyin similarity, edit distance similarity, boundary consistency, character order similarity, and n-gram similarity between the matching target and candidate words.

[0009] S4: Terminology proofreading based on a large model: When the complexity of the detailed query term is complex, terminology proofreading prompts are constructed using the proofreading text, candidate words of the detailed query term, and matching examples; the terminology proofreading prompts are input into the large model to perform terminology verification on the proofreading text, and the proofreading results are obtained; the matching examples are generated based on the professional terminology database, term location, and candidate words of the detailed query term; the candidate words of the detailed query term are the candidate words corresponding to the detailed query term in the terminology proofreading set.

[0010] Preferably, the n-gram retrieval engine and the prefix tree fuzzy retrieval engine are constructed based on the character length of the professional terms in the professional terminology database.

[0011] Preferably, the method for coarse term lookup in step S2 is as follows:

[0012] S21: Use the n-gram grammar retrieval engine to perform long term retrieval on the proofreading text to obtain the first candidate word set;

[0013] S22: Perform short term retrieval on the proofreading text using a prefix tree fuzzy retrieval engine to obtain a second candidate word set;

[0014] S23: Merge the first candidate word set and the second candidate word set into a candidate word set; the candidate word set serves as the coarse query result for terms.

[0015] Preferably, the method for performing term retrieval on the proofreading text using an n-gram grammar retrieval engine in step S2 is as follows:

[0016] A1: Extract professional terms with a character length greater than or equal to a preset length from the professional terminology database to obtain long professional terms; segment the long professional terms using a sliding window algorithm to obtain a long professional term sequence; the long professional term sequence serves as the index of the long professional terms.

[0017] A2: The proofreading text is segmented using a sliding window algorithm to obtain a proofreading text sequence;

[0018] A3: Calculate the degree of matching between the proofreading text sequence and the index of the long technical terms; extract technical terms from the technical term database based on the degree of matching to obtain a first candidate word set.

[0019] Preferably, the method for performing term retrieval on the proofreading text using the prefix tree fuzzy search engine in step S2 is as follows:

[0020] A terminology prefix tree is constructed based on professional terms in the terminology database whose character length is less than a preset length; the proofreading text is then subjected to fuzzy term matching with the terminology prefix tree to obtain a second set of candidate words.

[0021] Preferably, the method for detailed term lookup is as follows:

[0022] S31: The proofreading text is segmented using a sliding window algorithm to obtain a set of matching targets; the set of matching targets includes: matching targets and position markers; the position markers include the start and end positions of the matching targets in the proofreading text.

[0023] S32: Calculate the five-dimensional similarity between each matching target in the matching target set and each candidate word in the candidate word set; the five-dimensional similarity is:

[0024] (1)

[0025] in, Represents five-dimensional similarity; Indicates the similarity of pinyin; Indicates edit distance similarity; Indicates boundary consistency; Indicates the similarity of character order; Indicates n-gram similarity; This represents the weight of the similarity in each of the five dimensions;

[0026] S33: Compare the five-dimensional similarity between the candidate words in the candidate word set and each matching target to obtain the highest five-dimensional similarity of the candidate words;

[0027] S34: The matching target corresponding to the highest five-dimensional similarity of the candidate words is taken as the fine query term; the position mark of the fine query term is taken as the term position; the fine query term is the term in the determined proofreading text;

[0028] S35: Based on the candidate words, construct a term collation set with the highest five-dimensional similarity between the detailed query terms and the candidate words.

[0029] Preferably, the method for calculating the pinyin similarity is as follows:

[0030] a1: Segment the candidate words to obtain candidate word segments; obtain the corresponding pinyin with tones, i.e., candidate pinyin, by querying the pinyin database;

[0031] a2: Segment the target word to obtain the target word segmentation; by querying the pinyin database, obtain the pinyin with tone marks corresponding to the target word segmentation, i.e., the target pinyin;

[0032] a3: Calculate the similarity between the candidate pinyin and the target pinyin; the similarity is the pinyin similarity.

[0033] Preferably, the edit distance similarity The calculation method is as follows:

[0034] (2)

[0035] Where d represents the edit distance; The length of the longer of the two characters being compared is indicated by the length of the longer character; the edit distance refers to the minimum number of single-character edit operations required to convert the target word into a candidate word; the single-character edit operations include insertion, deletion, and replacement; the two characters refer to the target word and the candidate word, respectively.

[0036] Preferably, the method for calculating the character order similarity is as follows:

[0037] b1: Determine the index positions of the same characters in the matching target and candidate words, respectively, within the matching target and candidate words;

[0038] b2: Calculate character order similarity based on the index position. ,for:

[0039] (3)

[0040] Where c represents the number of characters in the same position; N represents the total number of characters in the target and candidate words; the number of characters in the same position refers to the number of characters with the same character at the same index position in the target and candidate words.

[0041] Preferably, the method for calculating the n-gram similarity is as follows:

[0042] c1: Segment the matching target and candidate words according to preset values ​​to obtain a set of matching target text sequences and a set of candidate word text sequences;

[0043] c2: Match the text sequences in the target text sequence set and the candidate word text sequence set to obtain the number of identical text sequences in the target text sequence set and the candidate word text sequence set, that is, the number of text sequences that are matched by the target.

[0044] c3: Calculate the n-gram similarity based on the number of text sequences matched by the target, using the following formula:

[0045] (4)

[0046] in, This represents the n-gram similarity, where M represents the number of text sequences matched by the target. This indicates the number of text sequences in the candidate word text sequence set.

[0047] Preferably, the method for determining the weights corresponding to the similarities in each of the five dimensions of similarity is as follows:

[0048] B1: Construct a terminology test set based on a professional terminology database; the terminology test set includes a similar terminology sample set and a dissimilar terminology sample set;

[0049] B2: Set the weight values ​​for the five dimensions: Pinyin similarity, edit distance similarity, boundary consistency, character order similarity, and n-gram similarity.

[0050] B3: Select weight values ​​from the weight value set of the five dimensions respectively and arrange them to obtain the five-dimensional weight set;

[0051] B4: Calculate the judgment accuracy corresponding to each five-dimensional weight combination in the five-dimensional weight set; select the five-dimensional weight combination with the highest accuracy from the five-dimensional weight set according to the judgment accuracy to obtain the optimal five-dimensional weight combination; the judgment accuracy refers to the accuracy of calculating the five-dimensional similarity of samples in the terminology test set according to the five-dimensional weight combination, and judging whether the samples are similar according to the five-dimensional similarity.

[0052] B5: Determine whether the accuracy rate of similar samples corresponding to the optimal five-dimensional weight combination reaches 1, and whether the accuracy rate of dissimilar samples reaches a preset threshold; if so, the optimal five-dimensional weight combination is taken as the final five-dimensional weight combination; otherwise, based on the optimal five-dimensional weight combination, reset the weight value set of the five dimensions of pinyin similarity, edit distance similarity, boundary consistency, character order similarity and n-gram similarity, and repeat steps B3-B5.

[0053] Preferably, the method for determining the complexity of the detailed query terms in step S4 is as follows:

[0054] If the five-dimensional similarity of the detailed query term in the term collation set is greater than or equal to the preset similarity, and the boundary consistency score between the detailed query term and the corresponding candidate word is 1, then the complexity of the detailed query term is simple; otherwise, the complexity of the detailed query term is complex. If the candidate word corresponding to the detailed query term is a non-entity term, then the complexity of the detailed query term is complex. The corresponding candidate word refers to the candidate word corresponding to the detailed query term in the term collation set.

[0055] Preferably, when the complexity of the detailed query term is simple, the term verification method is as follows: the detailed query term and the candidate words and five-dimensional similarity corresponding to the detailed query term in the term verification set are used as the verification result.

[0056] Preferably, the term verification method when the detailed query term is determined to be complex in step S4 is as follows:

[0057] C1: Generate Matching Examples: Determine the specific statement of the detailed query term in the proofreading text based on the term position, and use the specific statement as the input context; calculate the similarity between the input context and the standard usage examples in the terminology database; determine the standard usage examples corresponding to the input context based on the similarity, and obtain matching examples; the standard usage examples are all standard usage examples corresponding to the candidate words of the detailed query term in the terminology database; the candidate words of the detailed query term are the candidate words corresponding to the detailed query term in the terminology proofreading set;

[0058] C2: Construct terminology proofreading prompts: Construct terminology proofreading prompts based on the proofreading text, the candidate words of the detailed query terms, and the matching examples;

[0059] C3: Terminology Validation: Input the terminology proofreading prompts into the large model, perform terminology validation on the proofreading text, and obtain the validation results.

[0060] Beneficial effects:

[0061] This invention belongs to the field of natural language processing and proposes a terminology proofreading method based on fuzzy matching and large-scale model semantic discrimination. This method reduces the misjudgment rate of professional terminology proofreading and simultaneously improves the accuracy of the proofreading results. Specifically, in the coarse terminology query process, an n-gram retrieval engine and a prefix tree fuzzy retrieval engine are used to retrieve candidate words. These candidate words can effectively identify all possible terms in the proofreading text, achieving rapid screening and comprehensive coverage of terms in the proofreading text. In the fine terminology query process, by calculating the five-dimensional similarity between candidate words and matching targets extracted from the proofreading text, the terms and their positions in the proofreading text can be accurately determined. Complexity judgment rules are set based on five-dimensional similarity and boundary consistency, which can effectively and quickly determine the complexity of fine query terms. Different proofreading methods are selected according to the complexity of the fine query terms, improving the efficiency of terminology proofreading of the input proofreading text. When the complexity of the fine query terms is complex, terminology proofreading is performed using a large-scale model and the matched correct professional terms, i.e., candidate words, reducing the misjudgment rate of correct terms in the proofreading text that are not included in the professional terminology database. This invention combines coarse term query, fine term query, and large-scale model term verification, which not only achieves rapid term matching and accurate term location in the text being verified, but also reduces the misjudgment rate caused by similar correct terms, thereby improving the accuracy of term verification.

[0062] Phased retrieval improves efficiency and coverage: In the coarse term query stage, a metasyntax retrieval engine is used for term retrieval of long terms, and a term prefix tree fuzzy retrieval engine is used for fuzzy term matching of short terms. By taking the advantages of each other, the matching efficiency and recall rate are improved, which solves the problem of incomplete term matching in the proofreading text, thereby improving the overall coverage of term proofreading of the input text.

[0063] Enhanced matching accuracy through five-dimensional similarity fusion: In the fine query stage, the similarity of five dimensions—phonetic similarity (capturing near-phonetic errors), edit distance similarity (quantifying character differences), boundary consistency (precise segmentation and positioning), character order similarity (recognizing reversed expressions), and n-gram similarity (evaluating structural overlap)—can comprehensively quantify the matching degree between candidate words in the candidate word set obtained from the coarse query and the matching target in the proofreading text. This enables precise determination of terms and their positions in the proofreading text, effectively improving the accuracy of terminology correction in complex error scenarios.

[0064] Large-scale model semantic discrimination reduces the false positive rate: To address the problem of similar correct terms being easily misjudged, the semantic understanding and knowledge reasoning capabilities of the Large Language Model (LLM) are utilized to perform contextual association analysis and semantic legality judgment on the results of coarse and fine queries, distinguishing correct terms from real erroneous words, breaking through the limitations of traditional rules and significantly reducing the false positive rate. Attached Figure Description

[0065] Figure 1 This is a flowchart of a terminology collation method based on fuzzy matching and large model semantic discrimination.

[0066] Figure 2 This is a flowchart illustrating the terminology proofreading of input text using a terminology proofreading method based on fuzzy matching and large model semantic discrimination. Detailed Implementation

[0067] The following detailed description is provided in conjunction with the accompanying drawings and specific embodiments.

[0068] Example 1:

[0069] like Figure 1 As shown, a terminology collation method based on fuzzy matching and large-model semantic discrimination includes the following steps:

[0070] S1: Construct a professional terminology database and search engine: Construct a professional terminology database, and construct an n-gram syntax search engine and a prefix tree fuzzy search engine based on the character length of the professional terms in the database;

[0071] S2: Coarse terminology search: The proofreading text is searched using an n-gram retrieval engine and a prefix tree fuzzy retrieval engine to obtain a first candidate word set and a second candidate word set; the first candidate word set and the second candidate word set are merged into a candidate word set as the coarse terminology search result;

[0072] S3: Detailed Terminology Query: The proofreading text is segmented using a sliding window algorithm to obtain the matching target and term positions; the five-dimensional similarity between the matching target and the candidate words in the candidate word set is calculated; the terms in the proofreading text are determined based on the five-dimensional similarity and the matching target corresponding to the five-dimensional similarity, resulting in detailed query terms; a term proofreading set is constructed based on the detailed query terms, candidate words, and five-dimensional similarity; the five-dimensional similarity is obtained by weighted summation of the similarity of the matching target and candidate words in five dimensions: pinyin similarity, edit distance similarity, boundary consistency, character order similarity, and n-gram similarity, along with the weights of each dimension.

[0073] S4: Terminology proofreading based on a large model: Determine the complexity of the detailed query terms; if determined to be complex, determine the input context of the detailed query terms based on the term location; filter in a professional terminology database based on the input context of the detailed query terms and the candidate words corresponding to the detailed query terms in the terminology proofreading set to obtain matching examples; construct terminology proofreading prompts using the proofreading text, the candidate words corresponding to the detailed query terms in the terminology proofreading set, and the matching examples; input the terminology proofreading prompts into the large model to perform terminology verification on the proofreading text to obtain the proofreading result.

[0074] Example 2:

[0075] like Figure 2 The process of proofreading terminology in the text is shown below, including:

[0076] 1. Construct a professional terminology database

[0077] The terminology database consists of two main parts: first, the terminology text, used for fuzzy matching of the terminology; and second, the parsing of the terminology, including its definition, standard usage examples, etc., to supplement the terminology information and enable the large model to understand the meaning and usage of the terminology.

[0078] The terminology database is sourced from domain-standard terminology databases and professional dictionaries, extracting terms and definitions. Terms and contextual sentences are also extracted from manually reviewed text. Additionally, example sentences corresponding to relevant terms are obtained from media articles, allowing the model to learn the latest usage scenarios of these terms.

[0079] Specifically, the format for saving technical terms in the technical terminology database is as follows:

[0080]

[0081] Specifically, the professional terminology library includes: entry, that is, professional term, and professional term analysis; the professional term analysis includes a standard reference example corresponding to the professional term.

[0082] 2. Rough query, that is, rough terminology query

[0083] Quickly filter data from the professional terminology library, return a large number of possibly relevant but not sufficiently accurate professional term matching results, reduce the number of candidate words in the fine query stage, and speed up the matching. In the rough query stage, retrieval is performed by combining an n-gram retrieval engine and a fuzzy term prefix tree retrieval engine based on term length. The rough query can match all possible correct versions of terms in the proofreading text, which improves the comprehensiveness of term verification in the term proofreading text, but cannot confirm the specific position of the term in the term proofreading text.

[0084] Specifically, in the rough query process, term matching is performed on the proofreading text respectively by the n-gram retrieval engine and the fuzzy term prefix tree retrieval engine in the professional terminology library.

[0085] Specifically, the steps of retrieving potentially existing terms in the term proofreading text by combining the n-gram retrieval engine and the fuzzy prefix tree retrieval engine based on term length are as follows:

[0086] First step: perform term retrieval on the proofreading text through the n-gram retrieval engine:

[0087] For terms with a character length greater than or equal to n (n=5), the n-gram retrieval engine is used for term retrieval.

[0088] Specifically, in order to more accurately match various incorrectly written terms in the proofreading text, an n-gram retrieval engine is constructed.

[0089] Specifically, the steps of performing term retrieval on the proofreading text through the n-gram retrieval engine are as follows:

[0090] 1) Segment long professional terms, that is, professional terms with character length greater than or equal to 5, in the professional terminology library, set the window size to 2, each binary group advances character by character through the sliding window, and adjacent binary groups share one character.

[0091] For example: construct an index for the professional term "da po sha guo wen dao di" (an idiom meaning to never stop digging until one gets to the bottom of a matter), the index includes:

[0092] "da po", "po sha", "sha guo", "guo wen", "wen dao", "dao di".

[0093] 2) The proofreading text is segmented using the sliding window algorithm to obtain a proofreading text sequence; the similarity between the proofreading text sequence and all indices in the terminology database is calculated; specifically, when segmenting the proofreading text and long terminology, it is necessary to ensure that the segmentation results of the proofreading text and long terminology are of the same length, such as dividing the long terminology and proofreading text into binary sequences.

[0094] Specifically, the n-gram algorithm is used for segmentation, and the n-gram algorithm is a specific implementation of the sliding window algorithm.

[0095] Specifically, the proofreading text is 2-gram segmented based on a preset value of n (n=2) to obtain multiple proofreading text binary sequences; the matching degree between each proofreading text binary sequence and the index of long technical terms constructed in step 1) is calculated; based on the matching degree, the matched technical terms are extracted from the technical terminology database to obtain a first candidate word set; the matching degree is the similarity between the binary sequence of the proofreading text and each index in the proofreading terminology database. The matched technical terms refer to the technical terms corresponding to a matching degree greater than or equal to a preset value.

[0096] Step 2: Perform term retrieval on the proofreading text using a prefix tree fuzzy search engine.

[0097] For terms with a character length less than n (n=5), terms with wildcards are automatically constructed, and a prefix tree fuzzy search engine is built.

[0098] (1) Constructing a term prefix tree

[0099] Specifically, wildcards are added to professional terms in the terminology database that have a character length of less than 5.

[0100] For example, after adding wildcards to the technical term "killing two birds with one stone," the following fuzzy matching terms are obtained:

[0101] "Double the coin with one arrow", "Double the coin with one arrow", "Double the coin with one arrow", etc.

[0102] Then, a term prefix tree is constructed based on the fuzzy matching words.

[0103] Specifically, a terminology prefix tree can be directly constructed based on professional terms with a character length of less than 5, without the need to add wildcards.

[0104] (2) Fuzzy matching of terms

[0105] Based on the term prefix tree constructed in (1), perform fuzzy term matching on the proofreading text. The specific steps are as follows:

[0106] 1) Start from the first character of the text being checked, and use it as the starting point of the current matching window;

[0107] 2) Match the characters in the child nodes of the term prefix tree root node with the current character; if the match fails, the matching window moves one position to the right and determines whether the current character is the last character of the proofreading text. If it is, the fuzzy matching ends; if not, the second step is repeated; if the match is successful, proceed to the next step.

[0108] 3) Proceed to the next level node of the term prefix tree, while simultaneously shifting the matching window one position to the right; match the current character with the child nodes of the current node. If the match is successful, determine whether the next level node of the term prefix tree is an end marker. If it is an end marker, add the term corresponding to the matching path to the second candidate word set, simultaneously shift the matching window one position to the right, and proceed to the second step; if it is not an end marker, repeat the third step; if the match fails, proceed to the second step.

[0109] Specifically, the term prefix tree is constructed based on fuzzy matching words. During the fuzzy matching process, when a wildcard is encountered, all possible child nodes are traversed. If the child node can continue to match characters, the matching continues.

[0110] Fuzzy matching using term prefix trees can effectively address the relatively low recall rate of n-gram retrieval when matching short terms. Combining term prefix trees with fuzzy matching can effectively identify shorter terms in the proofreading text.

[0111] Using n-gram retrieval and term prefix tree fuzzy matching, terms similar to those in the proofreading text can be retrieved quickly and comprehensively from the terminology database.

[0112] Specifically, a coarse search primarily involves fuzzy matching. It retrieves a set of candidate terms ["candidate term 1", "candidate term 2", ...] from a terminology database, allowing for a rough assessment that the input text is likely related to these candidate terms. Since the terminology database contains a large amount of data, the coarse search step aims to initially filter the input text for possible terms based on quantity, thereby improving retrieval speed. However, a coarse search cannot precisely identify the matching targets (terms in the input text) and their locations within the corresponding text that match the candidate terms.

[0113] 3. Detailed search, i.e., detailed terminology search.

[0114] The fine-grained query can accurately locate the corresponding positions of the candidate words given by the coarse query in the proofreading text. That is, from the perspective of five-dimensional similarity, it marks which texts in the text to be proofread are professional terms that differ from the standard terms in the professional terminology database, and locates the texts that are suspected of having errors.

[0115] During the detailed query process, the position of the matching target corresponding to the candidate word given by the coarse query in the proofreading text is accurately located by the sliding window, and the five-dimensional similarity between the candidate word and the matching target is calculated. The high matching result is obtained by using the five-dimensional similarity to determine the matching target corresponding to the candidate word in the proofreading text, that is, to determine the professional terms that may be incorrect in the proofreading text, so as to avoid the problem of using other content in the proofreading text as terms and making incorrect proofreading.

[0116] Specifically, the candidate words refer to the correct terms in the terminology database that are related to the proofreading text; the candidate words are the candidate words in the candidate word set obtained in the coarse query step.

[0117] Specifically, the matching target refers to technical terms similar to the candidate words, and is a text segment in the proofreading text.

[0118] Specifically, the five-dimensional similarity is obtained by weighted summation of the similarities in five dimensions: pinyin similarity, edit distance similarity, boundary consistency, character order similarity, and n-gram similarity. The calculation steps are as follows:

[0119] The first step is to calculate the pinyin similarity: convert the characters into their corresponding pinyin, and use the ratio of similarity between the two words as the pinyin similarity to improve the matching accuracy in cases of similar-sounding errors.

[0120] Specifically, the method for converting characters into corresponding pinyin is as follows: after segmenting the matching target and candidate words, the pinyin corresponding to each word is queried in the pinyin library; the pinyin is pinyin with tones.

[0121] The second step is to calculate the edit distance similarity: the minimum number of single-character editing operations required to convert one string to another; the single-character editing operations include insertion, deletion, and replacement; the smaller the edit distance, the greater the similarity between the matching target and the candidate words.

[0122] Specifically, the edit distance similarity for:

[0123] (2)

[0124] Where d represents the edit distance; This indicates the length of the longer character among the two characters being compared.

[0125] Specifically, the edit distance value can be determined using algorithms such as dynamic programming.

[0126] The third step is to calculate character order similarity: check whether the two strings differ only in their order of representation, and assign a similarity score based on the magnitude of this difference. The specific steps are as follows:

[0127] (1) determining the index positions of identical characters in the matching target and the candidate word within the matching target and the candidate word respectively; for example, taking "Lv Shui Jin Shan (green water and golden mountain)" as the reference, the recorded index positions are 0, 1, 2, 3. For the contrasted "Lv Shan Jin Shui (green mountain and golden water)", the recorded index positions are 0, 3, 2, 1.

[0128] (2) comparing the difference between the two sets of index positions to obtain the character order similarity. The character order similarity is:

[0129] (3)

[0130] wherein, c represents the number of characters at the same positions; N represents the total number of characters in the matching target and the candidate word;

[0131] The fourth step: calculating boundary consistency: checking whether the first character and the last character of the two character strings are consistent, so as to position and segment the matching target more accurately.

[0132] The fifth step: calculating n-gram similarity: in the n-gram retrieval process, calculating the similarity through the proportion of hit text sequences.

[0133] Through statistical experimental analysis, setting the preset value as 2 achieves the best effect, so segmentation is performed according to binary groups.

[0134] The specific steps are:

[0135] segmenting the candidate word "da po sha guo wen dao di (never stop digging into a subject)" into 6 binary groups, that is, segmenting into 6 text sequences, to obtain a candidate word text sequence set ["da po (break)", "po sha (break casserole)", "sha guo (casserole)", "guo wen (casserole ask)", "wen dao (ask to)", "dao di (to the end)"].

[0136] segmenting the matching target "gai you da po sha guo wen dao di de jin tou (should have the perseverance of never stopping digging into a subject)" into 11 binary groups, that is, segmenting into 11 text sequences, to obtain a matching target text sequence set ["gai you (should have)", "you da (have break)", "da po (break)", "po sha (break casserole)", "sha guo (casserole)", "guo wen (casserole ask)", "wen dao (ask to)", "dao di (to the end)", "di de (end's)", "de jin (perseverance's)", "jin tou (perseverance)"].

[0137] matching the candidate word text sequence set with the matching target text sequence set to obtain identical text sequences in the two sets, that is, to obtain the text sequences hit in the matching target: "da po (break)", "guo wen (casserole ask)", "wen dao (ask to)", "dao di (to the end)".

[0138] then calculating the n-gram similarity according to formula (3), the n-gram similarity is:

[0139] (4)

[0140] Where M represents the number of text sequences matched by the target; This indicates the number of text sequences in the candidate word text sequence set.

[0141] Step 6: Weighted summation of the similarity scores across five dimensions—pinyin similarity, edit distance similarity, boundary consistency, character order similarity, and n-gram similarity—results in the final similarity score:

[0142] (1)

[0143] in, Represents five-dimensional similarity. Indicates the similarity of pinyin. Indicates boundary consistency. This indicates the weights corresponding to the similarity in each of the five dimensions.

[0144] Specifically, The corresponding values ​​were obtained using the following statistical analysis methods:

[0145] a. Compile a batch of professional terms, and construct a terminology test set based on the professional terms, which is divided into two sample sets: similar and dissimilar.

[0146] b. Set the set of values ​​for the similarity weight of each dimension, such as [0.1, 0.2, 0.3, ...].

[0147] c. Arrange and combine each similarity weight, and calculate the five-dimensional similarity of each weight combination (five-dimensional weight combination) for each sample on the test set.

[0148] d. If the calculated five-dimensional similarity can determine whether a sample is similar, mark it as correct, and calculate the accuracy rate of judging similar terms and dissimilar terms for each five-dimensional weight combination on the test set; calculate the term test set judgment accuracy rate corresponding to each five-dimensional weight combination based on the similarity judgment accuracy rate and the dissimilarity judgment accuracy rate.

[0149] e. Select the five-dimensional weight combination with the highest accuracy in judging the terminology test set, and reset the precision of the value set of the five weights to further refine the value range, such as [0.51, 0.52, 0.53, ...].

[0150] f. Repeat the above process until, on the terminology test set, the accuracy rate of identifying similar samples as similar is 100%, and the accuracy rate of identifying dissimilar samples as dissimilar is relatively high (i.e., the accuracy rate reaches 94%).

[0151] Specifically, the determination principle for the five-dimensional weight combination is: when making a judgment based on the calculated five-dimensional similarity, it is allowable to judge dissimilar ones as similar, because any errors in terms determined to be similar can be accurately identified in the subsequent proofreading step, but it is not allowed to judge similar ones as dissimilar, which will result in missing data.

[0152] For example, the values of five-dimensional similarity calculated for matching targets in the matching target set and candidate words are as follows:

[0153] Proofread text: When encountering difficult problems, one should have the perseverance to keep digging until getting to the bottom of the matter, so as to master knowledge firmly.

[0154] Candidate word: da po sha guo wen dao di (to get to the bottom of something)

[0155] In the sliding window matching process, the different matched matching targets (the matching target set is constructed through these matching targets) and the five-dimensional similarity between some matching targets and candidate words are as follows:

[0156]

[0157] After the five-dimensional similarity between all matching targets and candidate words is calculated, all five-dimensional similarities corresponding to the candidate words are compared to determine the maximum five-dimensional similarity; according to the candidate word, the matching target corresponding to the maximum five-dimensional similarity of the candidate word is the fine query term, and a term proofreading set is constructed from the candidate word, the matching target and the maximum five-dimensional similarity of the candidate word.

[0158] The fine query is described below with reference to specific text content:

[0159] The input text is "dui shi yao da po sha guo wen dao di", and the candidate word set is: ["da po sha guo wen dao di"].

[0160] The input text is segmented by the sliding window algorithm to obtain a matching target set: [{dui shi yao da po sha guo wen dao di, 0, 9}, {shi yao da po sha guo wen dao di, 1, 9}, {yao da po sha guo wen dao di, 2, 9}, {da po sha guo wen dao di, 3, 9}], where the first column is the matching target, the second column is the starting position of the matching target in the proofread text, and the third column is the ending position of the matching target in the proofread text.

[0161] Each matching target in the matching target set is calculated separately with "da po sha guo wen dao di" in the candidate word set to obtain the five-dimensional similarity, and the results are as follows:

[0162]

[0163] After comparing all the five-dimensional similarities corresponding to the candidate words, the highest five-dimensional similarity of the candidate words is 0.95. The matching target corresponding to 0.95 is the fine query term, which is the term in the proofreading text. The position marker {3, 9} corresponding to the fine query term is the specific position of the fine query term in the proofreading text, that is, the position of the term in the proofreading text is {3, 9}.

[0164] The term collation set constructed based on the detailed query terms, candidate words, and the highest five-dimensional similarity of the candidate words is ["Breaking through the pot to the bottom, breaking through the pot to the bottom, 0.95"].

[0165] Specifically, when the highest five-dimensional similarity of a candidate word is lower than a preset threshold (e.g., 0.6), it indicates that the candidate word is not a term present in the proofreading text, and therefore the candidate word is deleted.

[0166] Specifically, the highest five-dimensional similarity of the candidate words in the constructed terminology collation set must all be greater than a preset threshold. Furthermore, the detailed query terms and candidate words in the constructed terminology collation set must correspond one-to-one to ensure that the highest five-dimensional similarity of the candidate words can be calculated between the detailed query terms and the candidate words.

[0167] Specifically, the structure of the term collation set is: ["detailed query term, candidate word, highest five-dimensional similarity of the candidate word"], wherein the highest five-dimensional similarity of the candidate word must be greater than a preset threshold.

[0168] Specifically, the fine query finds the fine query terms in the proofreading text by calculating five-dimensional similarity and sliding window positioning, which correspond to each candidate word in the candidate word set obtained from the coarse query. In other words, it finds the specific location and text content of the terms in the proofreading text.

[0169] 4. Terminology proofreading

[0170] In proofreading applications, fuzzy matching results generally cannot be directly used as proofreading results because detailed query terms that are similar to correct terms are not necessarily incorrect. For example, phrases like "City A Science and Technology Innovation Guidance Center," "City B Science and Technology Innovation Guidance Center," and "City C Science and Technology Innovation Guidance Center" are all similar words and terms related to correct institution names. In practical applications, it is difficult to include all correct and similar professional terms completely. Therefore, relying solely on fuzzy matching results can easily lead to correctly identified terms in the proofreading text being incorrect. Large-scale models, however, learn a wealth of knowledge during pre-training. Using large-scale models, combined with a professional terminology database, can more intelligently determine whether detailed query terms are incorrect.

[0171] Specifically, for example, if the candidate words are only "City A Science and Technology Innovation Guidance Center" and the detailed query term is "City B Science and Technology Innovation Guidance Center," the candidate words and detailed query term are similar from a string perspective, but not completely identical. For instance, both contain "City Science and Technology Innovation Guidance Center" and constitute a high proportion of the entire term, and these words are correct. However, if the judgment is based solely on the fuzzy matching results, the detailed query term would be considered incorrect, which is a misjudgment. Similarly, the detailed query term "City B Science and Technology Innovation Guidance Center" is also similar from a character perspective, containing both "City Science and Technology Innovation" and "Guidance Center" and constituting a high proportion of the entire term. However, this term is incorrect because "innovation" describes improvement, iteration, or breakthrough thinking, while "innovation" emphasizes the process of creation from scratch and subsequent development, and cannot be used as the name of an official institution.

[0172] Specifically, the primary goal of proofreading is to identify erroneous text in the text to be proofread. During the proofreading phase, from a semantic perspective, it is determined whether there are indeed problems with the detailed query terms identified in the detailed query steps within the text to be proofread.

[0173] Specifically, the steps for terminology proofreading are as follows:

[0174] The first step is to determine the complexity of the terms (i.e., the matching targets in the terminology check set) according to the rules. The specific rules are as follows:

[0175] 1. If the five-dimensional similarity is sufficiently high (above 0.9) and the boundaries of the matching target and candidate words are consistent, the judgment is simple. Consistent boundaries between the matching target and candidate words refer to a boundary consistency score of 1.

[0176] 2. Except for point 1, all non-entity terms are judged as complex.

[0177] The purpose of judging the complexity of the matching target is mainly to reduce the complexity of implementing term validation and increase the system's operating efficiency.

[0178] The second step is to perform terminology correction using different methods based on the terminology's complexity value:

[0179] When a term is judged to be simple, the detailed query result, i.e., the detailed query term corresponding to the five-dimensional similarity in the term correction set, the candidate word, and the five-dimensional similarity are directly output as the correction result. If the five-dimensional similarity in the term correction set is 1, it means that the matching target and the candidate word are completely consistent and there is no error. Therefore, the correction result does not include the term correction result corresponding to the five-dimensional similarity of 1.

[0180] When a detailed query term is deemed complex, terminology correction is performed using a larger model. When using a larger model for terminology correction, the larger model uses detailed information and references related to that term.

[0181] Specifically, the steps for terminology correction using a large model are as follows:

[0182] 1) Search for reference materials: After finding the detailed query term through detailed search, search for reference materials of candidate words with similar contexts to the detailed query term in the professional terminology database. The reference materials are standard usage examples of the candidate words.

[0183] Specifically, during the query, the specific sentence or paragraph in the terminology proofreading text is determined based on the terminology location, i.e., the location marker of the detailed query term. The specific sentence or paragraph is used as the input context. The similarity between all standard usage examples of the candidate words corresponding to the detailed query term in the professional terminology database and the input context is calculated. The standard usage examples with similarity reaching a preset threshold are used as part of the terminology proofreading prompts.

[0184] Specifically, the terminology proofreading prompts are:

[0185]

[0186] The terminology proofreading prompts mentioned above include: Goal: the role of the large model and the function of that role; Rules: how the large model specifically performs the terminology proofreading task; input: the text to be proofread, as well as candidate words corresponding to the detailed query terms and standard usage examples similar to the input context, and what kind of results are expected from the large model.

[0187] 2) Large Model Proofreading: Construct terminology proofreading prompts using the proofreading text, candidate words, and standard usage examples; input the terminology proofreading prompts into the large model, enabling the large model to gain a deeper understanding of the meaning and usage of professional terms and improve the accuracy of the model in judging similar terms.

[0188] It should be noted that the specific embodiments described above enable those skilled in the art to more fully understand the present invention, but do not limit the present invention in any way. Therefore, although the present invention has been described in detail with reference to the accompanying drawings and embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the present invention. In short, all technical solutions and improvements that do not depart from the spirit and scope of the present invention should be covered within the protection scope of the present invention patent.

Claims

1. A terminology proofreading method based on fuzzy matching and large-model semantic discrimination, characterized in that, S1: Construct a terminology database and search engine: Construct a terminology database, and based on the terminology database, construct an n-gram syntax search engine and a prefix tree fuzzy search engine; S2: Coarse terminology search: The proofreading text is searched in the professional terminology database by the n-gram retrieval engine and the prefix tree fuzzy retrieval engine respectively to obtain the candidate word set, which is the result of the coarse terminology search. S3: Detailed terminology query: Segment the proofreading text to obtain the matching target and term positions; calculate the five-dimensional similarity between the matching target and the candidate words in the results of the coarse terminology query; determine the detailed query terms in the proofreading text based on the five-dimensional similarity and the matching target; A term collation set is constructed based on the detailed query terms, candidate words, and five-dimensional similarity; the five-dimensional similarity is obtained by weighted summation of the pinyin similarity, edit distance similarity, boundary consistency, character order similarity, and n-gram similarity between the matching target and candidate words. S4: Term proofreading based on large model: When the complexity of the detailed query term is complex, construct term proofreading prompts by using the proofread text, candidate words of the detailed query term and matching examples; The terminology proofreading prompts are input into the large model to perform terminology verification on the proofreading text, and the proofreading results are obtained; the matching examples are generated based on the professional terminology database, terminology location, and candidate words of detailed query terms.

2. A terminology proofreading method based on fuzzy matching and large-model semantic discrimination as described in claim 1, characterized in that, The n-gram syntax retrieval engine and the prefix tree fuzzy retrieval engine are constructed based on the character length of the professional terms in the professional terminology database.

3. A terminology proofreading method based on fuzzy matching and large-model semantic discrimination as described in claim 1, characterized in that, The method for coarse term lookup in step S2 is as follows: S21: Use the n-gram grammar retrieval engine to perform long term retrieval on the proofreading text to obtain the first candidate word set; S22: Perform short term retrieval on the proofreading text using a prefix tree fuzzy retrieval engine to obtain a second candidate word set; S23: Merge the first candidate word set and the second candidate word set into a candidate word set; the candidate word set serves as the result of the coarse term query.

4. A terminology proofreading method based on fuzzy matching and large-model semantic discrimination as described in claim 1 or 3, characterized in that, The method for performing term retrieval on the proofreading text using an n-gram grammar retrieval engine in step S2 is as follows: A1: Extract professional terms with a character length greater than or equal to a preset length from the professional terminology database to obtain long professional terms; segment the long professional terms using a sliding window algorithm to obtain a long professional term sequence; the long professional term sequence serves as the index of the long professional terms. A2: The proofreading text is segmented using a sliding window algorithm to obtain a proofreading text sequence; A3: Calculate the degree of matching between the proofreading text sequence and the index of the long technical terms; extract technical terms from the technical term database based on the degree of matching to obtain a first candidate word set.

5. A terminology proofreading method based on fuzzy matching and large-model semantic discrimination as described in claim 1 or 3, characterized in that, Step S2 uses a prefix tree fuzzy search engine to perform term retrieval on the proofreading text as follows: A terminology prefix tree is constructed based on professional terms in the terminology database whose character length is less than a preset length; the proofreading text is then subjected to fuzzy term matching with the terminology prefix tree to obtain a second set of candidate words.

6. A terminology proofreading method based on fuzzy matching and large-model semantic discrimination as described in claim 1, characterized in that, The method for detailed term lookup is as follows: S31: The proofreading text is segmented using a sliding window algorithm to obtain a set of matching targets; the set of matching targets includes: matching targets and position markers; the position markers include the start and end positions of the matching targets in the proofreading text. S32: Calculate the five-dimensional similarity between each matching target in the matching target set and each candidate word in the candidate word set; the five-dimensional similarity is: (1) in, Represents five-dimensional similarity; Indicates the similarity of pinyin; Indicates edit distance similarity; Indicates boundary consistency; Indicates the similarity of character order; Indicates n-gram similarity; This represents the weight of the similarity in each of the five dimensions; S33: Compare the five-dimensional similarity between the candidate words in the candidate word set and each matching target to obtain the highest five-dimensional similarity of the candidate words; S34: The matching target corresponding to the highest five-dimensional similarity of the candidate words is taken as the fine query term; the position mark of the fine query term is taken as the term position; the fine query term is the term in the determined proofreading text; S35: Based on the candidate words, construct a term collation set with the highest five-dimensional similarity between the detailed query terms and the candidate words.

7. A terminology proofreading method based on fuzzy matching and large-model semantic discrimination as described in claim 1, characterized in that, The method for calculating the pinyin similarity is as follows: a1: Segment the candidate words to obtain candidate word segments; obtain the corresponding pinyin with tones, i.e., candidate pinyin, by querying the pinyin database; a2: Segment the target word to obtain the target word segmentation; by querying the pinyin database, obtain the pinyin with tone marks corresponding to the target word segmentation, i.e., the target pinyin; a3: Calculate the similarity between the candidate pinyin and the target pinyin; the similarity is the pinyin similarity.

8. A terminology proofreading method based on fuzzy matching and large-model semantic discrimination as described in claim 1, characterized in that, Edit distance similarity The calculation method is as follows: (2) Where d represents the edit distance; The length of the longer of the two characters being compared is indicated by the length of the longer character; the edit distance refers to the minimum number of single-character edit operations required to convert the target word into a candidate word; the single-character edit operations include insertion, deletion, and replacement; the two characters refer to the target word and the candidate word, respectively.

9. A terminology proofreading method based on fuzzy matching and large-model semantic discrimination as described in claim 1, characterized in that, The method for calculating the character order similarity is as follows: b1: Determine the index positions of the same characters in the matching target and candidate words, respectively, within the matching target and candidate words; b2: Calculate character order similarity based on the index position. ,for: (3) Where c represents the number of characters in the same position; N represents the total number of characters in the target and candidate words; the number of characters in the same position refers to the number of characters with the same character at the same index position in the target and candidate words.

10. A terminology proofreading method based on fuzzy matching and large-model semantic discrimination as described in claim 1, characterized in that, The method for calculating the n-gram similarity is as follows: c1: Segment the matching target and candidate words according to preset values ​​to obtain a set of matching target text sequences and a set of candidate word text sequences; c2: Match the text sequences in the target text sequence set and the candidate word text sequence set to obtain the number of identical text sequences in the target text sequence set and the candidate word text sequence set, that is, the number of text sequences that are matched by the target. c3: Calculate the n-gram similarity based on the number of text sequences matched by the target, using the following formula: (4) in, This represents the n-gram similarity, where M represents the number of text sequences matched by the target. This indicates the number of text sequences in the candidate word text sequence set.

11. A terminology proofreading method based on fuzzy matching and large-model semantic discrimination as described in claim 6, characterized in that, The method for determining the weights corresponding to the similarity in each of the five dimensions is as follows: B1: Construct a terminology test set based on a professional terminology database; the terminology test set includes a similar terminology sample set and a dissimilar terminology sample set; B2: Set the weight values ​​for the five dimensions: Pinyin similarity, edit distance similarity, boundary consistency, character order similarity, and n-gram similarity. B3: Select weight values ​​from the weight value set of the five dimensions respectively and arrange them to obtain the five-dimensional weight set; B4: Calculate the judgment accuracy corresponding to each five-dimensional weight combination in the five-dimensional weight set; select the five-dimensional weight combination with the highest accuracy from the five-dimensional weight set according to the judgment accuracy to obtain the optimal five-dimensional weight combination; the judgment accuracy refers to the accuracy of calculating the five-dimensional similarity of samples in the terminology test set according to the five-dimensional weight combination, and judging whether the samples are similar according to the five-dimensional similarity. B5: Determine whether the similarity judgment accuracy corresponding to the optimal five-dimensional weight combination reaches 1, and whether the dissimilar sample judgment accuracy reaches a preset threshold. If the optimal five-dimensional weight combination is achieved, then the optimal five-dimensional weight combination is taken as the final five-dimensional weight combination; otherwise, based on the optimal five-dimensional weight combination, the weight value set of the five dimensions of pinyin similarity, edit distance similarity, boundary consistency, character order similarity and n-gram similarity is reset, and steps B3-B5 are repeated.

12. A terminology proofreading method based on fuzzy matching and large-model semantic discrimination as described in claim 1, characterized in that, The method for determining the complexity of the detailed query terms in step S4 is as follows: If the five-dimensional similarity of the detailed query term in the term collation set is greater than or equal to the preset similarity, and the boundary consistency score between the detailed query term and the corresponding candidate word is 1, then the complexity of the detailed query term is simple; otherwise, the complexity of the detailed query term is complex. If the candidate term corresponding to the detailed query term is a non-entity term, then the complexity level of the detailed query term is complex. The corresponding candidate words refer to the candidate words that correspond to the detailed query terms in the term collation set.

13. A terminology proofreading method based on fuzzy matching and large-model semantic discrimination as described in claim 1 or 12, characterized in that, When the complexity of the detailed query term is simple, the term verification method is as follows: take the detailed query term and the candidate words and five-dimensional similarity corresponding to the detailed query term in the term verification set as the verification result.

14. A terminology proofreading method based on fuzzy matching and large-model semantic discrimination as described in claim 1, characterized in that, The term verification method when the detailed query terms are determined to be complex in step S4 is as follows: C1: Generate Matching Examples: Determine the specific statement of the detailed query term in the proofreading text based on the term position, and use the specific statement as the input context; calculate the similarity between the input context and the standard usage examples in the terminology database; determine the standard usage examples corresponding to the input context based on the similarity, and obtain matching examples; the standard usage examples are all standard usage examples corresponding to the candidate words of the detailed query term in the terminology database; the candidate words of the detailed query term are the candidate words corresponding to the detailed query term in the terminology proofreading set; C2: Construct terminology proofreading prompts: Construct terminology proofreading prompts based on the proofreading text, the candidate words of the detailed query terms, and the matching examples; C3: Terminology Validation: Input the terminology proofreading prompts into the large model, perform terminology validation on the proofreading text, and obtain the validation results.

Citation Information

Patent Citations

  • Document proofreading processing method based on large model

    CN119599003A

  • Artificial intelligence assisted cross-language automatic note taking and term marking system

    CN120183408A