Chinese error correction method and device and storage medium
By constructing a recall set containing candidate characters with similar pronunciations and shapes, and determining and correcting spelling errors in Chinese text, the problem of insufficient accuracy and generalization ability of Chinese spelling correction in the prior art is solved, and more efficient and accurate error correction effects are achieved.
Patent Information
- Application Number
- CN202510286181.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-06-27
AI Technical Summary
The existing Chinese spelling correction methods have problems with insufficient accuracy and generalization capabilities, especially the marking method memorizes error patterns, while the retelling method lacks sufficient information supplement, resulting in inaccurate and flexible error correction.
By determining the recall set of the target to be corrected characters in the Chinese text to be corrected, the error correction result is determined from the recall set based on the target to be corrected characters and the recall set, and the error correction Chinese text to be corrected based on the error correction result. The recall set includes target to be corrected characters and/or candidate characters, and the candidate characters and the target to be corrected characters are similar in pronunciation and/or glyph.
It improves the accuracy and generalization ability of Chinese error correction, enables the Chinese error correction process to have knowledge retrieval capabilities, can efficiently integrate external knowledge, enrich the amount of knowledge injection, and provides sufficient candidates for subsequent accurate error correction.
Smart Images

Figure CN120218053A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of natural language processing, and in particular, to a Chinese error correction method, apparatus, and storage medium. Background Art
[0002] The Chinese Spelling Correction (CSC) task is mainly used to detect and correct spelling mistakes (such as typos) in Chinese texts, and has always been a crucial basic task in Natural Language Processing (NLP). It is widely applied in fields such as web search, speech recognition, machine translation, text generation and editing (such as script generation and script editing).
[0003] In recent years, the state-of-the-art (SOTA) CSC methods tend to be more inclined to rephrasing methods rather than tagging methods. Research shows that tagging methods have inherent limitations - models often memorize fixed patterns of errors and corrections, rather than truly understanding sentence semantics and making reasonable corrections. In contrast, although rephrasing methods are more flexible, due to the lack of sufficient information supplementation, there are still certain expression bottlenecks. For example, ReLM (Rephrasing Language Model) attempts to simulate the semantic understanding process of humans, but lacks knowledge retrieval ability, which limits its error correction accuracy and generalization ability. How to more efficiently improve the accuracy and generalization ability of Chinese error correction remains an important challenge in current CSC research. Summary of the Invention
[0004] In view of this, the present disclosure proposes a Chinese error correction method, apparatus, and storage medium.
[0005] According to one aspect of the present disclosure, a Chinese error correction method is provided. The method includes:
[0006] Determine a recall set corresponding to a target character to be corrected in the Chinese text to be corrected, where the recall set includes the target character to be corrected and / or candidate characters, and the candidate characters are similar to the target character to be corrected in terms of pronunciation and / or glyph;
[0007] Based on the target character to be corrected and the recall set, determine the corrected result from the recall set;
[0008] Based on the corrected result, correct the Chinese text to be corrected.
[0009] In a possible implementation, the recall set further includes candidate words, and each character in the candidate words corresponds to the target character to be corrected and one or more characters after the target character to be corrected in the Chinese text to be corrected.
[0010] In a possible implementation, determining a recall set for a target character to be corrected in a Chinese text to be corrected includes:
[0011] Using any one or more of pinyin matching, four-corner code matching, radical matching, and shape matching to determine candidate characters and / or candidate words in the recall set;
[0012] Among them, pinyin matching determines candidate characters and / or candidate words based on the initials and finals of the character, four-corner code matching determines candidate characters and / or candidate words based on the encoding of the Chinese character structure components of the character, radical matching determines candidate characters and / or candidate words based on the radical of the character, and shape matching determines candidate characters and / or candidate words based on the shape structure of the character.
[0013] In a possible implementation, using any one or more of pinyin matching, four-corner code matching, radical matching, and shape matching to determine candidate characters and / or candidate words in the recall set includes:
[0014] Searching respectively on one or more search trees based on the target character to be corrected, and in response to the existence of a matching node in the search tree, using the character corresponding to the matching node as a candidate character; wherein each search tree corresponds to one matching method respectively.
[0015] In a possible implementation, using any one or more of pinyin matching, four-corner code matching, radical matching, and shape matching to determine candidate characters and / or candidate words in the recall set includes:
[0016] Searching respectively on one or more search trees based on the target character to be corrected and a preset number of characters after the target character to be corrected;
[0017] In response to the existence of a matching node in the search tree, using the word corresponding to the matching node as a candidate word, and after increasing the preset number, re-executing the steps of searching respectively on one or more search trees based on the target character to be corrected and a preset number of characters after the target character to be corrected and subsequent steps until there are no matching nodes in each search tree, to obtain the recall set.
[0018] In a possible implementation, based on the target character to be corrected and the recall set, determining the corrected result from the recall set includes:
[0019] Determining a first vector representation corresponding to the target character to be corrected;
[0020] Determining second vector representations corresponding to each character and / or word in the recall set;
[0021] Based on the first vector representation and the second vector representation, use the knowledge selection model to determine the corrected result of the target character to be corrected. The knowledge selection model is constructed based on the attention mechanism.
[0022] In a possible implementation, the knowledge selection model is a trained knowledge selection model, and the method further includes:
[0023] Obtain a training set, where the training set includes training texts and labels, and the labels are used to indicate the actual corrected texts corresponding to the training texts;
[0024] Determine the third vector representation corresponding to the training character to be corrected in the training text;
[0025] Determine the fourth vector representations corresponding to each character and / or word in the recall set corresponding to the training character to be corrected;
[0026] Based on the third vector representation and the fourth vector representation, use the initial knowledge selection model to calculate the attention weights corresponding to each fourth vector representation;
[0027] Weight each fourth vector representation based on the attention weights corresponding to each fourth vector representation to obtain a third fused vector representation;
[0028] Weight the third vector representation and the third fused vector representation to obtain a fourth fused vector representation;
[0029] Based on the attention weights corresponding to each fourth vector representation, the fourth fused vector representation, and the labels, optimize the parameters of the initial knowledge selection model to obtain a trained knowledge selection model.
[0030] In a possible implementation, based on the first vector representation and the second vector representation, using the knowledge selection model to determine the corrected result of the target character to be corrected includes:
[0031] Based on the first vector representation and the second vector representation, use the knowledge selection model to calculate the attention weights corresponding to each second vector representation;
[0032] Weight each second vector representation based on the attention weights corresponding to each second vector representation to obtain a first fused vector representation;
[0033] Weight the first vector representation and the first fused vector representation to obtain a second fused vector representation;
[0034] Based on the second fused vector representation, determine the corrected result from the recall set.
[0035] In a possible implementation, the corrected result is the target candidate character or target candidate word in the recall set. Based on the corrected result, correct the Chinese text to be corrected, including:
[0036] Replace the target character to be corrected in the Chinese text to be corrected with the target candidate character; or,
[0037] Replace the target character to be corrected and one or more characters after the target character to be corrected in the Chinese text to be corrected with the target candidate word.
[0038] According to another aspect of the present disclosure, a Chinese error correction device is provided. The device includes:
[0039] A first determination module, configured to determine a recall set corresponding to the target character to be corrected in the Chinese text to be corrected, the recall set including the target character to be corrected and / or candidate characters, and the candidate characters having similarity in pronunciation and / or glyph with the target character to be corrected;
[0040] A second determination module, configured to determine the corrected result from the recall set based on the target character to be corrected and the recall set;
[0041] A correction module, configured to correct the Chinese text to be corrected based on the corrected result.
[0042] In a possible implementation, the recall set further includes candidate words, and each character in the candidate words corresponds to the target character to be corrected and one or more characters after the target character to be corrected in the Chinese text to be corrected, respectively.
[0043] In a possible implementation, the first determination module is configured to:
[0044] Use any one or more of the matching methods of pinyin matching, four-corner code matching, radical matching, and shape matching to determine the candidate characters and / or candidate words in the recall set;
[0045] Among them, pinyin matching determines candidate characters and / or candidate words based on the initials and finals of the characters, four-corner code matching determines candidate characters and / or candidate words based on the encoding of the Chinese character structure components of the characters, radical matching determines candidate characters and / or candidate words based on the radicals of the characters, and shape matching determines candidate characters and / or candidate words based on the shape structure of the characters.
[0046] In a possible implementation, using any one or more of the matching methods of pinyin matching, four-corner code matching, radical matching, and shape matching to determine the candidate characters and / or candidate words in the recall set includes:
[0047] Search based on the target character to be corrected on one or more search trees respectively, and in response to the existence of a matching node in the search tree, use the character corresponding to the matching node as the candidate character; wherein, each search tree corresponds to a matching method respectively.
[0048] In a possible implementation, any one or more of pinyin matching, four-corner code matching, radical matching, and shape matching are used to determine candidate characters and / or candidate words in the recall set, including:
[0049] Searching on one or more search trees respectively based on the target character to be corrected and a preset number of characters after the target character to be corrected;
[0050] In response to the existence of a matching node in the search tree, using the word corresponding to the matching node as a candidate word, and after increasing the preset number, re-executing the steps of searching on one or more search trees respectively based on the target character to be corrected and a preset number of characters after the target character to be corrected and subsequent steps until there are no matching nodes in each search tree, to obtain the recall set.
[0051] In a possible implementation, a second determination module is used for:
[0052] Determining a first vector representation corresponding to the target character to be corrected;
[0053] Determining second vector representations corresponding to each character and / or word in the recall set;
[0054] Based on the vector representation and the second vector representations, using a knowledge selection model to determine the corrected result of the target character to be corrected, and the knowledge selection model is constructed based on an attention mechanism.
[0055] In a possible implementation, the knowledge selection model is a trained knowledge selection model, and the device further includes:
[0056] An acquisition module for acquiring a training set, where the training set includes training texts and labels, and the labels are used to indicate the actual corrected texts corresponding to the training texts;
[0057] A third determination module for determining a third vector representation corresponding to the training character to be corrected in the training text;
[0058] A fourth determination module for determining fourth vector representations corresponding to each character and / or word in the recall set corresponding to the training character to be corrected;
[0059] A calculation module for calculating the attention weights corresponding to each of the fourth vector representations respectively based on the third vector representation and the fourth vector representations by using an initial knowledge selection model;
[0060] A first weighting module for weighting each of the fourth vector representations based on the attention weights corresponding to each of the fourth vector representations respectively to obtain a third fused vector representation;
[0061] A second weighting module, configured to perform weighting based on the third vector representation and the third fusion vector representation to obtain a fourth fusion vector representation;
[0062] An optimization module, configured to optimize the parameters of the initial knowledge selection model based on the attention weights, the fourth fusion vector representation, and the labels respectively corresponding to each fourth vector representation, to obtain a trained knowledge selection model.
[0063] In a possible implementation manner, determining a corrected result of a target character to be corrected by using a knowledge selection model based on a first vector representation and a second vector representation includes:
[0064] Calculating the attention weights respectively corresponding to each second vector representation by using a knowledge selection model based on the first vector representation and the second vector representation;
[0065] Performing weighting on each second vector representation based on the attention weights respectively corresponding to each second vector representation to obtain a first fusion vector representation;
[0066] Performing weighting based on the first vector representation and the first fusion vector representation to obtain a second fusion vector representation;
[0067] Determining the corrected result from the recall set based on the second fusion vector representation.
[0068] In a possible implementation manner, the corrected result is a target candidate character or a target candidate word in the recall set, and a correction module is configured to:
[0069] Replacing the target character to be corrected in the Chinese text to be corrected with the target candidate character; or,
[0070] Replacing the target character to be corrected and one or more characters after the target character to be corrected in the Chinese text to be corrected with the target candidate word.
[0071] According to another aspect of the present disclosure, there is provided a Chinese error correction device, including a memory, a processor, and a computer program stored on the memory, where the processor executes the computer program to implement the steps of the above method.
[0072] According to another aspect of the present disclosure, there is provided a non-volatile computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above method are implemented.
[0073] According to another aspect of the present disclosure, there is provided a computer program product, including a computer program, or a non-volatile computer-readable storage medium carrying the computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.
[0074] According to an embodiment of the present disclosure, by determining a recall set corresponding to a target character to be corrected in a Chinese text to be corrected, based on the target character to be corrected and the recall set, determining a corrected result from the recall set, and based on the corrected result, correcting the Chinese text to be corrected, wherein the recall set includes the target character to be corrected and / or candidate characters, and the candidate characters have similarity in pronunciation and / or glyph with the target character to be corrected, which can enable the Chinese error correction process to have a knowledge retrieval ability, be able to efficiently integrate external knowledge, greatly enrich the knowledge injection amount, provide sufficient candidates for subsequent accurate error correction, and thus improve the accuracy and generalization ability of Chinese error correction.
[0075] Other features and aspects of the present disclosure will become clear from the following detailed description of exemplary embodiments with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0076] The drawings included in and constituting a part of the specification, together with the specification, illustrate exemplary embodiments, features, and aspects of the present disclosure and are used to explain the principles of the present disclosure.
[0077] Figure 1 A schematic diagram showing an application scenario according to an embodiment of the present disclosure.
[0078] Figure 2 A schematic diagram showing the structure of a knowledge recall and selection network according to an embodiment of the present disclosure.
[0079] Figure 3 A flowchart showing a Chinese error correction method according to an embodiment of the present disclosure.
[0080] Figure 4 A schematic diagram showing the process of using a knowledge recall and selection network according to an embodiment of the present disclosure.
[0081] Figure 5 A schematic diagram showing a knowledge representation model according to an embodiment of the present disclosure.
[0082] Figure 6 A schematic diagram showing a knowledge selection model according to an embodiment of the present disclosure.
[0083] Figure 7 A structural diagram showing a Chinese error correction device according to an embodiment of the present disclosure.
[0084] Figure 8 A block diagram showing an apparatus 1900 for Chinese error correction according to an exemplary embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0085] Various exemplary embodiments, features, and aspects of the present disclosure will be described in detail below with reference to the accompanying drawings. Like reference numerals in the drawings denote functionally identical or similar elements. Although various aspects of the embodiments are shown in the drawings, the drawings are not necessarily drawn to scale unless otherwise specified.
[0086] As used herein, the terms "comprising," "including," "having," or variations thereof are open-ended and include one or more stated features, integers, elements, steps, components, or functions, but do not preclude the presence or addition of one or more other features, integers, elements, steps, components, functions, or groups thereof.
[0087] When an element is referred to as being "connected," "coupled," "responsive," or variations thereof to another element, it can be directly connected, coupled, or responsive to the other element, or intervening elements may be present.
[0088] Although the terms first, second, third, etc. may be used herein to describe various elements / operations, these elements / operations should not be limited by these terms. These terms are only used to distinguish one element / operation from another. Thus, a first element / operation in some embodiments may be referred to as a second element / operation in other embodiments without departing from the teachings of the inventive concept.
[0089] The term "exemplary" as used herein means "serving as an example, embodiment, or illustration." Any embodiment illustrated herein as "exemplary" should not necessarily be construed as superior to or better than other embodiments.
[0090] In addition, for a better illustration of the present disclosure, numerous specific details are given in the following detailed description. Those skilled in the art should understand that the present disclosure can be implemented without some of these specific details. In some instances, methods, means, elements, and circuits well known to those skilled in the art are not described in detail so as to highlight the gist of the present disclosure.
[0091] The Chinese Spelling Correction (CSC) task is mainly used to detect and correct spelling mistakes (such as typos) in Chinese texts. It has always been a crucial basic task in Natural Language Processing (NLP) and is widely applied in fields such as web search, speech recognition, machine translation, text generation, and editing.
[0092] In recent years, the most advanced (SOTA) CSC methods have tended to favor rephrasing methods rather than tagging methods. Research has shown that tagging methods have inherent limitations - models often memorize fixed patterns of errors and corrections, rather than truly understanding sentence semantics and making reasonable corrections. In contrast, although rephrasing methods are more flexible, there are still certain expression bottlenecks due to the lack of sufficient information supplementation. For example, ReLM (Rephrasing Language Model) attempts to simulate the human semantic understanding process, but lacks knowledge retrieval capabilities, limiting its error correction accuracy and generalization ability.
[0093] The inventors recognize that currently, the main bottleneck in improving the performance of the CSC task lies in the effective injection of knowledge. Therefore, how to more efficiently integrate external knowledge and improve the accuracy and generalization ability of Chinese error correction remains an important challenge in current CSC research.
[0094] In view of this, the embodiments of the present disclosure provide a Chinese error correction method, apparatus, and storage medium. The method of the embodiments of the present disclosure determines a recall set corresponding to a target error-prone character in the Chinese text to be error-corrected, determines the error-corrected result from the recall set based on the target error-prone character and the recall set, and corrects the Chinese text to be error-corrected based on the error-corrected result. Among them, the recall set includes the target error-prone character and / or candidate characters, and the candidate characters are similar to the target error-prone character in terms of pronunciation and / or glyph. This can enable the Chinese error correction process to have knowledge retrieval capabilities, efficiently integrate external knowledge, greatly enrich the amount of knowledge injection, provide sufficient candidates for subsequent accurate error correction, and thus improve the accuracy and generalization ability of Chinese error correction.
[0095] Figure 1 A schematic diagram showing an application scenario according to an embodiment of the present disclosure. The Chinese error correction method of the embodiments of the present disclosure can be used in any application scenario related to natural language processing of Chinese, such as web search, speech recognition, machine translation, text generation, and editing, as Figure 1 shown, the method of the embodiments of the present disclosure can obtain the Chinese text to be error-corrected (such as "legal far body hope"), and correct the spelling errors (such as "far far", "hope") in the Chinese text to be error-corrected through the Knowledge Recall and Selection Network (ReSC) of the embodiments of the present disclosure, and output the corrected Chinese text (such as "legal origin system").
[0096] The method of the embodiments of the present disclosure can be applied to a terminal device or a server. The terminal device can be any one or more of a mobile phone, a foldable electronic device, a tablet computer, a desktop computer, a laptop computer, a handheld computer, a notebook computer, an ultra-mobile personal computer (UMPC), a netbook, a cellular phone, a personal digital assistant (PDA), and an in-vehicle device. The embodiments of the present disclosure do not impose any special restrictions on the specific type of the terminal device, and it can have wired or wireless communication functions. The server can be located locally or in the cloud, and can be a physical device or a virtual device, such as a virtual machine or a container, with wireless communication functions. Among them, the wireless communication functions can be set in the chip (system) or other components or assemblies of the server. The wireless communication functions can be implemented, for example, through mobile communication technologies such as 2G / 3G / 4G / 5G, as well as Wi-Fi, Bluetooth, frequency modulation (FM), data radio, satellite communication, etc. Communication can also be carried out through wired connections to achieve interaction with other devices.
[0097] Figure 2 FIG. shows a schematic structural diagram of a knowledge recall and selection network according to an embodiment of the present disclosure. As Figure 2 shown, the knowledge recall and selection network of the embodiments of the present disclosure can include a language model (LM), a knowledge recall model, a knowledge representation and selection network (including a knowledge representation model and a knowledge selection model), and an output layer. Among them, the language model can be used to determine the vector representation corresponding to each character in the Chinese text to be corrected, the knowledge recall model can be used to determine the recall set based on the characters in the Chinese text to be corrected, the knowledge representation model can be used to determine the vector representation corresponding to each character and / or word in the recall set, the knowledge selection model can be used to determine the fused knowledge representation based on the outputs of the language model and the knowledge representation model, and the output layer can be used to output the corrected results of each character in the Chinese text to be corrected based on the fused knowledge representation.
[0098] The following introduces the Chinese error correction method of the embodiments of the present disclosure based on Figure 2 this. Figure 3 FIG. shows a flowchart of a Chinese error correction method according to an embodiment of the present disclosure. This method can be applied to the above-mentioned terminal device or server. As Figure 3 shown, this method can include:
[0099] Step S301, determining a recall set corresponding to a target character to be corrected in a Chinese text to be corrected.
[0100] The Chinese text to be corrected is, for example, "法律远度希", which may contain spelling errors (such as "远" and "希"). In the embodiment of the present disclosure, the Chinese text to be corrected can be input into the above-mentioned knowledge recall and selection network. First, the first character ("法") in the Chinese text to be corrected is used as the target character to be corrected. After executing steps S301-S303, the next character ("律") is used as the target character to be corrected, until the last character is executed, and the corrected Chinese text corresponding to the Chinese text to be corrected is obtained.
[0101] You can use Figure 2 The knowledge recall model in performs step S301, wherein the recall set (also referred to as the confusion set) may refer to a set of characters or words that are often confused because of their similar shapes and / or pronunciations. In the CSC task, the recall set can be used to provide correction candidates. For example, "已" and "巳" are easy to confuse because of their similar appearances, and they can constitute a simple recall set. The recall set may include target characters to be corrected and / or candidate characters, and the candidate characters have similarities with the target characters to be corrected in pronunciation and / or shape. For example, the candidate character has the same pinyin as the target character to be corrected (the tone may be the same or different), or the candidate character has the same radical, similar shape or structure as the target character to be corrected, etc.
[0102] In order to further enhance the model's expressive power and improve the accuracy of error correction, the recall set of the embodiment of the present disclosure may also include candidate words, and each character in the candidate word corresponds to the target character to be corrected and one or more characters after the target character to be corrected in the Chinese text to be corrected. For example, for the target character to be corrected "法", when constructing the recall set corresponding to "法", the present disclosure can not only recall candidate characters for the single character "法", but also recall candidate words of two characters for the word "法律" composed of "法" and the character "律" after it, and recall candidate words of three characters for the word "法律远" composed of "法" and the two characters "律远" after it, etc., until the word composed of "法" and one or more characters after it cannot be recalled as a candidate word. By introducing multi-granular candidate words, the model can not only correct errors for single characters, but also recall more candidate results in combination with words composed of multiple characters, avoiding new semantic errors caused by correcting characters in isolation. This recall mechanism helps to improve the overall error correction effect.
[0103] The recall set in the related art has problems such as small scale and ineffective screening. It is only an additional feature, which limits the recall rate. To solve this problem, the embodiment of the present disclosure introduces multiple matching methods to construct the recall set. In step S301, it can:
[0104] Determine candidate characters and / or candidate words in the recall set by using any one or more matching methods including pinyin matching, four-corner code matching, radical matching, and shape matching.
[0105] Among them, pinyin matching can determine candidate characters and / or candidate words based on common pinyin errors using the initial consonants and finals of the characters. Through the pinyin matching method, characters with the same pinyin (i.e., initial consonants and finals) as the target character to be corrected can be used as candidate characters. For example, for the target character to be corrected "远", the determined candidate characters may include "渊". For words composed of multiple characters, words with the same pinyin can be used as candidate words. For example, for the word "远远" composed of the target character to be corrected "远" and the subsequent character "远", the determined candidate words may include "渊源".
[0106] Among them, since the most common incorrect characters in the CSC task are incorrect tones, for example, "epilepsy" (dian3xian2) and "dot line" (dian3xian4), the difference between the two is the different tones. The embodiment of the present disclosure may not limit the tones when using pinyin matching to recall more candidate characters or candidate words.
[0107] In order to enhance the recall capability for visual and character structure features, a four-corner code matching method may also be introduced in the disclosed embodiment. The four-corner code matching may utilize visual and character structure features for recall, wherein candidate characters and / or candidate words may be determined based on the Chinese character structure component encoding of the characters. Through the four-corner code matching method, characters having the same Chinese character structure component encoding (such as the four-corner code) as the target character to be corrected may be used as candidate characters. For example, for the target character to be corrected "远", its four-corner code is 31306, and the determined candidate characters may include "逼" (whose four-corner code is also 31306). For words composed of multiple characters, words having the same four-corner code may be used as candidate words.
[0108] The encoding of Chinese character structural components can be to decompose Chinese characters into smaller components or structural features, and then represent them in a specific encoding method so as to more efficiently process, identify and match the encoding method of Chinese characters. The encoding of Chinese character structural components can be a four-corner code, which can be used to encode Chinese characters. By decomposing Chinese characters into different structural components and assigning corresponding numbers (between 0 and 9) to each corner according to the characteristics of the upper left corner, upper right corner, lower left corner and lower right corner of the character, a five-digit code is formed (the fifth digit is the supplementary distinguishing code). For example, although the characters "訇、茐、句、旬、甸" have different shapes, they correspond to the same four-corner code: 27620. By using the four-corner code matching method, the structural information of the characters can be effectively captured, so that the model can more accurately recall the correct characters corresponding to possible spelling errors.
[0109] Radical matching can determine candidate characters and / or candidate words based on the radicals of the characters. For example, the characters "椅" and "桌" are similar in structure, both have the radical "木", and both have the meaning of furniture. Therefore, characters with the same radical can be recalled based on the connection between the radical and the meaning or pronunciation of the Chinese character. Through the radical matching method, characters with the same radical as the target character to be corrected can be used as candidate characters. For example, for the target character to be corrected "远", which includes the radical "辶", the determined candidate characters can include "逼" (also including the radical "辶"). For words composed of multiple characters, words with the same radical (determined according to the radical of each character in the word) can be used as candidate words. For example, for the word "远" composed of the target character to be corrected "远" and the next character "远", its radical is "辶辶", and the determined candidate words can include "邂逅" (also including the radical "辶辶").
[0110] Shape matching can recall visually similar "similar characters", thereby improving the accuracy of error correction by leveraging the structural similarities between characters. For example, the characters "句" and "甸" both contain the same part "勹", but their meanings and usages are different. By identifying and leveraging these shared structural features, shape matching can effectively improve the accuracy of spelling correction.
[0111] Among them, shape matching can determine candidate characters and / or candidate words based on the shape structure of the characters. By means of radical matching, characters with a shape structure similar or identical to the target character to be corrected can be used as candidate characters. For example, for the target character to be corrected "远", the determined candidate characters may include "虑". For words composed of multiple characters, words with a similar or identical shape structure (determined according to the shape structure of each character in the word) can be used as candidate words. Among them, based on a pre-constructed dictionary of similar-shaped characters, characters / words with similar or identical shape structures are classified and stored, and characters / words with similar or identical shape structures are found through the dictionary of similar-shaped characters to determine candidate characters and / or candidate words.
[0112] In the embodiment of the present disclosure, character-based retrieval can be performed to find candidate characters related to the target character to be corrected. In the process of determining the candidate characters and / or candidate words in the recall set by using any one or more matching methods of pinyin matching, four-corner code matching, radical matching, and shape matching, the following can be performed:
[0113] Based on the target character to be corrected, a search is performed on one or more search trees respectively, and in response to the existence of a matching node in the search tree, the character corresponding to the matching node is used as a candidate character.
[0114] Among them, each search tree can correspond to a matching method respectively. The search tree is, for example, a trie tree, which can be pre-established based on dictionaries corresponding to different matching methods, and the candidate characters can be one or more. The nodes of the search tree can represent information related to different matching methods. For example, for pinyin matching, the nodes of the search tree can represent initial consonants and / or finals, and finally point to candidate characters corresponding to a certain pinyin; for four-corner code matching, the nodes can represent digital sequences of four-corner codes, and each path represents the encoding of different four-corner codes, and finally points to candidate characters corresponding to a certain four-corner code; for radical matching, the nodes can represent different radicals, and finally point to candidate characters with the same radical; for shape matching, the nodes can represent Chinese character features with visual similarity, and finally point to candidate characters with the same or similar shape structures.
[0115] In the embodiment of the present disclosure, it is also possible to retrieve based on words to find candidate words related to the target character to be corrected. In the process of determining the candidate characters and / or candidate words in the recall set by using any one or more matching methods of pinyin matching, four-corner code matching, radical matching, and shape matching, it is possible to:
[0116] Based on the target character to be corrected and a preset number of characters after the target character to be corrected, searches are performed on one or more search trees respectively; in response to the existence of a matching node in the search tree, the word corresponding to the matching node is used as a candidate word, and after increasing the preset number, the search is re-executed based on the target character to be corrected and a preset number of characters after the target character to be corrected, and the subsequent steps are performed on one or more search trees respectively until there are no matching nodes in each search tree, and a recalled set is obtained.
[0117] At this time, the matching nodes in the search tree may also represent nodes that match a word composed of multiple characters, indicating related information of the word composed of multiple characters in different matching modes.
[0118] For example, for the target character "远", taking the radical matching method as an example, you can first search based on a single character, search in the trie tree to find all the characters containing "辶" as candidate characters, and then based on the word "远" composed of "远" and the character "远" after it, search for the node corresponding to "辶辶" in the trie tree. If there is a node of "辶辶" (for example, the candidate word "渊源" corresponding to the node is found), then add a character "体" to form a new word "远体", and continue to search for the node corresponding to "辶辶亻" in the retrieval trie tree. If the corresponding node is not found, stop searching, otherwise you can increase the number of characters and continue searching until the corresponding node is not found.
[0119] By using the above method to determine the recall set, the number of recalled characters or words can be maximized, and a high recall rate can be maintained. The embodiments of the present disclosure can achieve an average recall rate of over 93%. About 150 relevant characters or words can be recalled for a single character, which can significantly increase the amount of knowledge injection.
[0120] In the embodiments of the present disclosure, although four recall methods are adopted, through the trie search tree and the non-segmentation method (that is, no word segmentation is performed on the input Chinese text to be corrected), the recall time complexity is significantly reduced. The knowledge selection process is lightweight and efficient, facilitating integration into other networks, improving the scalability and applicability of the model in practical applications, and reducing the application cost and technical threshold.
[0121] Step S302: Based on the target character to be corrected and the recall set, determine the corrected result from the recall set.
[0122] Figure 4 The flowchart showing the use of the knowledge recall and selection network according to the embodiments of the present disclosure is as follows. Figure 4 As shown, through the above method, the knowledge recall model can output the recall set c i corresponding to the target character x to be corrected xi , which includes the target character x to be corrected i and the related candidate characters / candidate words cand i1 , cand i2 ... Thus, step S302 can be executed.
[0123] In step S302, the corrected result can be determined based on the vector representation corresponding to the target character to be corrected and the vector representations corresponding to each character and / or word in the recall set.
[0124] In one example, step S302 may include:
[0125] Determine the first vector representation corresponding to the target character to be corrected;
[0126] Determine the second vector representations corresponding to each character and / or word in the recall set;
[0127] Based on the first vector representation and the second vector representations, use the knowledge selection model to determine the corrected result of the target character to be corrected.
[0128] Among them, the first vector representation can be determined through related technologies. In the embodiments of the present application, the above language model can be used to determine the first vector representation corresponding to the target character to be corrected.
[0129] The language model can be constructed based on relevant technologies. The structure of the language model is not limited in the embodiments of the present disclosure, as long as the output of the last layer is a latent vector representation. The latent vector representation can include semantic information and context information of each character in the input Chinese text to be corrected.
[0130] Compared with the method of constructing candidate set vectors by character embedding in related technologies, the output of the last layer of the language model (i.e., latent vector representation) is directly used as the first vector representation in the embodiment of the present disclosure. Because compared with the first layer, the last layer contains more information at the error correction level, and the attention mechanism in the subsequent knowledge selection model can generate a word vector projected to a single character, which is more conducive to capturing the error correction relationship between characters, thereby more accurately expressing the intrinsic meaning of each character in the word.
[0131] See also Figure 4 , for the target character x to be corrected i (as shown in the figure, “far”), the language model can output the corresponding first vector representation
[0132] The knowledge representation model can be constructed with reference to the structure of the language model, and the output of the last layer can also be a latent vector representation. The above-mentioned knowledge representation model can be used to determine the second vector representation corresponding to each character and / or word in the recall set. In the embodiment of the present disclosure, the output of the last layer of the knowledge representation model (i.e., the latent vector representation) can also be used as the second vector representation corresponding to each character and / or word (such as Figure 4 middle These include And the vector representation corresponding to each candidate character / candidate word is an empty vector representation), the second vector representation may include semantic information of the corresponding character / word.
[0133] Figure 5 FIG. 1 is a schematic diagram showing a knowledge representation model according to an embodiment of the present disclosure. Figure 5 As shown in the figure, in the knowledge representation model, for the target character to be corrected "远", by inputting the candidate word "[cls]源源[sep]" and the candidate character "[cls]逼[sep]" ([cls] and [sep] are classification tags and segmentation tags, respectively) into the knowledge representation model, its last layer (LM Layer Nx in the figure) can output the latent vector corresponding to "源源" (i.e., the second vector representation), and the second vector representation corresponding to "源源" is aligned with the second vector representation corresponding to a single character, representing the overall semantics of the word "源源". This representation method provides a richer and more accurate information basis for subsequent knowledge selection and spelling correction, and outputs the latent vector related to "逼" to represent the semantics of the character "逼".
[0134] The design of the knowledge selection model aims to accurately select appropriate characters or words from a large number of candidate characters / candidate words in the recall set, which directly determines the error correction ability of the model and is used to solve the key problem of recall set screening in the Chinese spelling correction task. The knowledge selection model is the core component for improving Chinese spelling correction ability in the embodiments of the present disclosure. Through a carefully designed mechanism, it accurately screens correct information from the recall set, effectively improving the accuracy and reliability of error correction.
[0135] The knowledge selection model can be constructed based on the attention mechanism and use the attention mechanism to achieve knowledge selection. This knowledge selection model is a trained knowledge selection model. First, the training method of the knowledge selection model will be introduced below. This method also includes:
[0136] Obtain a training set;
[0137] Determine the third vector representation corresponding to the training character to be error-corrected in the training text;
[0138] Determine the fourth vector representations corresponding to each character and / or word in the recall set corresponding to the training character to be error-corrected;
[0139] Based on the third vector representation and the fourth vector representations, use the initial knowledge selection model to calculate the attention weights corresponding to each fourth vector representation;
[0140] After weighting each fourth vector representation based on the attention weights corresponding to each fourth vector representation, obtain the third fused vector representation;
[0141] After weighting based on the third vector representation and the third fused vector representation, obtain the fourth fused vector representation;
[0142] Based on the attention weights corresponding to each fourth vector representation, the fourth fused vector representation, and the labels, optimize the parameters of the initial knowledge selection model to obtain the trained knowledge selection model.
[0143] Among them, the training set may include training texts and labels. The labels can be used to indicate the actual error-corrected texts corresponding to the training texts. For example, for each character in the training text, the label can be used to indicate its correct character / word.
[0144] Figure 6 Show a schematic diagram of the knowledge selection model according to the embodiments of the present disclosure. As Figure 6 shown, the training character to be error-corrected can refer to the above-mentioned target training character to be error-corrected (denoted by x i ), and the third vector representation corresponding to the training character to be error-corrected in the training text can be determined by using the above-mentioned language model. Refer to the above (as corresponding to "yuan" in the figure).
[0145] The above knowledge representation model can be used to determine the fourth vector representations corresponding to each character and / or word in the recall set for the character to be corrected and trained. See the above (as corresponding to "yuan"..."yu", "encounter", "force", "origin" in the figure).
[0146] The attention weight of any fourth vector representation corresponding to the third vector representation can be calculated based on the third vector representation, the fourth vector representation, each fourth vector representation corresponding to the third vector representation, and a trainable projection matrix. One way to calculate the attention weights corresponding to each fourth vector representation using the initial knowledge selection model (corresponding to Figure 6 the "Attention" process in) can be seen in the formula:
[0147]
[0148] where, a i,j can represent the attention weight corresponding to the fourth vector representation of the j-th candidate character / candidate word for the i-th character to be corrected and trained, are respectively the trainable projection matrices in the knowledge selection model, is the set of real numbers, d is the preset matrix dimension, can represent the fourth vector representation of the j-th candidate character / candidate word for the i-th character to be corrected and trained, can represent the third vector representation of the i-th character to be corrected and trained.
[0149] When the recall set contains appropriate candidate characters / candidate words, the model will give priority to correcting the characters rather than simply retaining the original characters, so that the appropriate candidate character / candidate word even exceeds the score of the original input (i.e., x i ) in terms of score.
[0150] One way to calculate the fourth fusion vector representation (corresponding to Figure 6 the "Weighted Sum" process in) can be seen in the formula:
[0151]
[0152] where, represents the fourth fusion vector representation of the i-th character to be corrected and trained, λ fk is the parameter for preset fusion knowledge, is a trainable parameter. is the third fusion vector representation of the i-th character to be corrected and trained.
[0153] When calculating using the attention mechanism in the embodiments of the present disclosure, the vector representation of the original input (i.e., ) is incorporated into the composition of the key (such as W K ) and the value (W V ). Referring to the above formula, where is multiplied by W K corresponding to to calculate the correlation, obtaining a i,j . Then, a i,j is used to multiply V corresponding to W . This design takes into account the situation where the knowledge selection model may not be able to successfully retrieve appropriate candidate characters / candidate words. When there are no effective candidate characters / candidate words, the model can rely on the vector representation of the original input (i.e., ) to learn a stronger correlation with itself, ensuring reasonable judgment can still be made in complex situations.
[0154] During the process of optimizing the parameters of the initial knowledge selection model, the parameters of the initial knowledge selection model can be optimized and learned through the following loss function:
[0155] L = (1 - λ KS )L SC + λ KS L KS
[0156] where L is the loss function for optimizing the parameters of the initial knowledge selection model, λ KS is a preset weight parameter, and the knowledge selection loss function L KS can be used to represent the difference between the attention weight value and the label. By calculating the value of L KS , W Q , W K can be optimized and learned. The value of L KS can be calculated from the attention weights and labels corresponding to each fourth vector representation respectively. One calculation method is as follows:
[0157]
[0158] where, can correspond to the normalized attention weight calculated in the knowledge selection model (the attention weight can be referred to as a Figure 6 in i,j , and does not involve the calculation of the fusion vector representation). The normalized attention weight can be a value between 0 and 1, can represent the j-th candidate character / candidate word corresponding to the i-th character to be corrected and trained, It can represent a tag of length N, which is a one-hot tag representing a true candidate (see Figure 6 , indicating that "source" is the correct word, and the value corresponding to "source" is 1, and the rest are 0). If the true tag is not included in the recall set, the training character x to be corrected i can be used as the true tag to calculate the cross-entropy.
[0159] In the embodiments of the present disclosure, when both "source" and "yuan" appear in the recall set of "yuan", since "source" contains more information, "source" is preferentially selected as the correct tag during training, thereby improving the error correction effect of the model.
[0160] The spelling correction loss function L SC can be used to represent the difference between the fused vector representation and the tag. By calculating the value of L SC , it can be obtained by calculating the attention weights corresponding to each fourth vector representation, the fourth fused vector representation, and the tag respectively. One calculation method is as follows:
[0161]
[0162] Among them, can correspond to the output of the spelling correction model in the knowledge selection model. The output of the spelling correction model is the calculated and normalized fused vector representation (the fused vector representation can be seen in the above The normalized fused vector representation is a value between 0 and 1), and can be defined by the normalized (softmax) probability as can represent the parameters of the output layer (Output Layer). V represents the number of training characters to be corrected in the training text. By calculating the value of L SC , the values of W V , W O can be optimized and learned.
[0163] By jointly obtaining the loss function L to optimize the parameters W Q , W K , W V of the initial knowledge selection model, and the parameters W O of the output layer, a trained knowledge selection model and a trained output layer can be obtained.
[0164] Return to see Figure 3 , in the process of determining the error-corrected result of the target character to be corrected by using the knowledge selection model based on the first vector representation and the second vector representation, it is possible to:
[0165] Based on the first vector representation and the second vector representation, use the knowledge selection model to calculate the attention weights corresponding to each second vector representation;
[0166] After weighting each second vector representation based on the attention weights corresponding to each second vector representation respectively, obtain the first fusion vector representation;
[0167] After weighting based on the first vector representation and the first fusion vector representation, obtain the second fusion vector representation;
[0168] Based on the second fusion vector representation, determine the corrected result from the recall set.
[0169] Among them, the method of using the knowledge selection model to calculate the attention weights corresponding to each second vector representation can refer to the calculation method of the attention weights in the above training process; the calculation method of the first fusion vector representation can refer to the calculation method of the third fusion vector representation above, and the calculation method of the second fusion vector can refer to the calculation method of the fourth fusion vector representation above (as shown in the output of the knowledge selection model in the figure ).
[0170] Through the output layer in Figure 4 , perform normalization processing on the second fusion vector representation. The normalized second fusion vector can represent the probability of each character / word in the recall set as the corrected result. The character / word with the maximum corresponding probability can be used as the corrected result of the target character to be corrected, that is, as the target candidate character / target candidate word.
[0171] Step S303, based on the corrected result, correct the Chinese text to be corrected.
[0172] According to the embodiments of the present disclosure, by determining the recall set corresponding to the target character to be corrected in the Chinese text to be corrected, based on the target character to be corrected and the recall set, determine the corrected result from the recall set, and based on the corrected result, correct the Chinese text to be corrected. Among them, the recall set includes the target character to be corrected and / or candidate characters, and the candidate characters are similar to the target character to be corrected in pronunciation and / or glyph. This can enable the Chinese error correction process to have knowledge retrieval capabilities, efficiently integrate external knowledge, greatly enrich the knowledge injection volume, provide sufficient candidates for subsequent accurate error correction, and thus improve the accuracy and generalization ability of Chinese error correction.
[0173] In step S303, it is possible to:
[0174] Replace the target character to be corrected in the Chinese text to be corrected with the target candidate character; or, replace the target character to be corrected in the Chinese text to be corrected and one or more characters after the target character to be corrected with the target candidate word.
[0175] Among them, the result after error correction is the target candidate character or target candidate word in the recall set. For example, see Figure 4 , for the first "yuan" in the Chinese text to be error-corrected, the target candidate word "origin" can be obtained, so that "yuanyuan" in the Chinese text to be error-corrected can be corrected to "origin"; for the second "yuan" in the Chinese text to be error-corrected, the target candidate word "source body" can be obtained, so that "yuanti" in the Chinese text to be error-corrected can be corrected to "source body"; for the "xi" in the Chinese text to be error-corrected, the target candidate character "system" can be obtained, so that "xi" in the Chinese text to be error-corrected can be corrected to "system", and finally the corrected Chinese text "legal origin system" is output. Among them, in the embodiments of the present disclosure, the two "yuan" in the Chinese text to be error-corrected can be independently updated respectively, effectively solving the overlapping problem brought by nested words and avoiding a series of problems that may occur in similar decoder structures in the related art, such as the problem of inability to ensure character alignment.
[0176] To verify the effect of the ReSC in the embodiments of the present disclosure, performance test experiments are carried out in the embodiments of the present disclosure. In the experiments, the present disclosure uses two main data sets, ECSpell and SIGHAN, which provide multi-domain and multi-scale data support for model training, evaluation and comparison with other methods, and help to comprehensively measure the performance of the model in different scenarios.
[0177] 1. ECSpell data set: It is a specific benchmark data set in the field of Chinese spelling correction, including data in three different fields: law (LAW), medicine (MED), and official document writing (ODW). These field data are carefully organized and can reflect the unique language challenges and term characteristics of their respective fields. Each field in the data set has corresponding training sets and test sets. For example, there are 1,960 training data and 500 test data in the law field; 2,500 training data and 500 test data in the medical field; 1,728 training data and 500 test data in the official document writing field. This data set is used to evaluate the Chinese spelling correction ability of the model in a specific field, and in the experiment, to ensure fairness, the same field dictionary as Rspell is used.
[0178] 2. SIGHAN Dataset: It contains multiple sub-datasets such as SIGHAN13, SIGHAN14, and SIGHAN15, which have been used in multiple studies for Chinese spelling correction experiments. In the paper experiments, these sub-datasets were also tested. Their statistical information is as follows: SIGHAN13 has 350 training data and 1000 test data; SIGHAN14 has 3437 training data and 1062 test data; SIGHAN15 has 2338 training data and 1100 test data. In addition, the Wang27k dataset is also involved. It is a large CSC dataset generated by Wang et al. in 2018 and was used in the training process of the SIGHAN dataset. Specifically, the ReLM model was first trained on the Wang27k dataset, and then trained and fine-tuned on SIGHAN13 - 15 respectively.
[0179] The models of the embodiments of the present disclosure used in the experiments include:
[0180] ReSC char : It can represent the application form of the ReSC model of the embodiments of the present disclosure at the character level. When dealing with Chinese spelling correction tasks, it mainly performs knowledge recall and selection operations based on characters. When determining the recall set, it only recalls for characters. In the recall stage, the above four recall methods are used to find candidate characters related to the incorrect character; in the selection stage, these candidate characters are screened to determine the final correction result. In the SIGHAN dataset experiment, due to the relevant settings of this dataset, ReSC only obtained the character-level result ReSC char , which is used to verify the efficiency of the selection network.
[0181] ReSC word : It can represent the application form of the ReSC model of the embodiments of the present disclosure at the word level. It not only considers character information but also fuses word information for spelling correction. When determining the recall set, it not only recalls for characters but also for words. In the selection stage, it comprehensively considers the knowledge representation and attention mechanism of characters and words to determine the correct word or character at each character position. In the ECSpell dataset experiment, ReSC word performed well. Compared with ReSC only at the character level char , the fused word information enhanced the model's expressive ability and achieved better results in metrics such as F1 score, proving the effectiveness of word-level information integration in CSC tasks.
[0182] Other methods used for comparison in the experiments:
[0183] 1. Fine-tuning-based methods:
[0184] Masked-Fine-Tuning (MFT): When training for the CSC task, a simple masking technique is applied to characters, enabling BERT-based models to achieve good results. The principle is to guide the model to learn character error correction patterns through the masking mechanism.
[0185] BERT: Bidirectional Encoder Representations from Transformers. It directly uses the MFT technique to fine-tune the BERT model. As a pre-trained language model, BERT is widely used in natural language processing tasks. After fine-tuning, it can be used for Chinese spelling correction tasks to learn spelling error patterns in the text.
[0186] 2. Language Model-based Methods:
[0187] Baichuan2: Fine-tune the well-known Chinese large language model Baichuan2 using the MFT technique to adapt to the CSC task. Although large language models have powerful language understanding capabilities, there may be problems such as character alignment in the CSC task.
[0188] ChatGPT: Apply ChatGPT to the CSC task through the OpenAI API and use its language generation ability to try to correct spelling errors. However, it performs poorly in handling CSC tasks with character alignment, such as miswriting "icy drink" as "betel nut".
[0189] 3. Specific Architecture-based Methods:
[0190] MDCSpell: Proposed by Zhu et al. in 2022, it is an enhanced BERT-based model. It adopts a detector-corrector architecture, attempts to retain the key visual and phonetic clues of misspelled characters, and improves the error correction ability through multi-task learning.
[0191] ReLM: Uses a rewriting method for Chinese spelling correction. Different from the basic token method, it corrects errors by rewriting the entire sentence. An auxiliary task of randomly replacing tokens with error characters and correcting them is set during pre-training to learn error correction patterns.
[0192] Rspell: A retrieval-enhanced framework for the CSC task. It integrates relevant domain terms through pinyin fuzzy confusion sets to enhance the error correction ability in specific domains. It also has an adaptive control mechanism and an iterative strategy to improve the error correction effect.
[0193] ECSpell UD:Proposed by Lv et al. in 2023, it is an error-consistent masking strategy for data generation during pre-training. It comes with a user dictionary-guided inference module (UD), attached to a general token classification spell checker, using the user dictionary to improve the error correction performance of domain-specific datasets.
[0194] SpellGCN: A graph convolutional network designed for CSC that integrates speech and visual similarity knowledge into the language model. By constructing a Chinese character graph structure and transforming it into an interdependent character classifier, it enhances the error detection and correction capabilities of the language model.
[0195] GAD: That is, Global Attention Decoder, proposed by Guo et al. in 2021. By capturing the global context relationship between characters and candidates, it improves the accuracy of Chinese spelling error correction, focusing on using global information for error correction.
[0196] The experimental results based on the ECSpell dataset can be seen in Table 1:
[0197] Table 1
[0198]
[0199] Among them, Domain can represent the domain of the data, Method can represent different models, Prec. (Precision), Rec. (Recall), and F1 (F1-score) can represent different evaluation metrics.
[0200] According to Table 1, it can be seen that ReSC word performs more prominently. In the comprehensive performance of the three fields of law, medicine, and official document writing, the F1 score reaches 94.4%, showing a significant improvement compared to other baseline methods, indicating that the strategy of integrating character and word information effectively enhances the error correction ability of the model.
[0201] By comparing with Rspell, it can be seen that the recall result of ReSC far exceeds that of Rspell. Due to the insufficient number of retrieval items in Rspell and the method of segmenting words first and then retrieving, some words cannot be correctly recognized. In the legal field, the F1 score of ReSC is 11% higher than that of Rspell, with a significant gap, fully demonstrating the superiority of the ReSC recall method.
[0202] By comparing with ReLM, it can be seen that ReSC integrates richer character and word information and the performance improvement is obvious. In the three fields, the average F1 score of ReSC is 3.36% higher than that of ReLM, proving that ReSC has more advantages in information utilization and error correction ability.
[0203] By comparing with ECSpell UD It can be seen that although ECSpell UD uses a large number of dictionaries, due to insufficient mining of dictionary content, its results are relatively poor. While ReSC can utilize knowledge more effectively and performs better in various metrics.
[0204] The deficiencies of large language models further highlight the advantages of ReSC. For example, large language models such as ChatGPT and Baichuan2 perform poorly in the CSC task. Taking ChatGPT as an example, it cannot guarantee character alignment when rewriting answers, such as miswriting "icy drink" as "betel nut". In contrast, ReSC has obvious advantages in handling CSC tasks with character alignment and can accurately correct spelling.
[0205] The experimental results based on the SIGHAN dataset can be seen in Table 2:
[0206] Table 2
[0207]
[0208] Referring to Table 2, it can be seen that since the SIGHAN dataset does not have a comprehensive domain dictionary, ReSC in the embodiments of the present disclosure mainly conducts experiments with a recall set at the character granularity on this dataset, aiming to verify the efficiency of the selection network. ReSC was tested on three sub-datasets: SIGHAN13, SIGHAN14, and SIGHAN15.
[0209] The specific index results are as follows:
[0210] For SIGHAN13: ReSC char has a precision of 84.6%, a recall of 80.1%, and an F1 score of 82.3%. Compared with SpellGCN, the F1 score has an approximate 6% improvement; compared with ReLM, the results are similar. However, due to the small size of the SIGHAN13 training set, which limits the model's learning, the advantages of ReSC are more obvious when compared with SpellGCN and GAD.
[0211] For SIGHAN14: ReSC char has a precision of 64.8%, a recall of 73.1%, and an F1 score of 68.7%. Compared with SpellGCN and ReLM, there is a certain improvement in recall and F1 score, indicating that ReSC can more effectively screen candidates in the recall set when processing this dataset, improving the error correction ability.
[0212] For SIGHAN15: ReSC charThe precision rate is 76.0%, the recall rate is 81.1%, and the F1 score is 78.5%. It is also superior to SpellGCN in terms of the recall rate and F1 score, and there is also a certain improvement compared with ReLM, further verifying the verification results of the effectiveness of ReSC in selecting networks on datasets of different scales and characteristics.
[0213] In summary, in the absence of a domain dictionary, ReSC can better distinguish necessary and unnecessary candidates by using the recall set compared with other comparison methods, demonstrating the adaptability and effectiveness of this method under complex data conditions.
[0214] Figure 7 The structural diagram of a Chinese error correction device according to an embodiment of the present disclosure is shown. As Figure 7 shown, the device includes:
[0215] A first determination module 701, configured to determine a recall set corresponding to a target error-correction character in the Chinese text to be error-corrected, where the recall set includes the target error-correction character and / or candidate characters, and the candidate characters have similarity with the target error-correction character in terms of pronunciation and / or glyph;
[0216] A second determination module 702, configured to determine the error-corrected result from the recall set based on the target error-correction character and the recall set;
[0217] A correction module 703, configured to correct the Chinese text to be error-corrected based on the error-corrected result.
[0218] In a possible implementation manner, the recall set further includes candidate words, and each character in the candidate words corresponds to the target error-correction character and one or more characters after the target error-correction character in the Chinese text to be error-corrected, respectively.
[0219] In a possible implementation manner, the first determination module 701 is configured to:
[0220] Use any one or more of the matching methods of pinyin matching, four-corner code matching, radical matching, and shape matching to determine candidate characters and / or candidate words in the recall set;
[0221] Among them, pinyin matching determines candidate characters and / or candidate words based on the initials and finals of the characters, four-corner code matching determines candidate characters and / or candidate words based on the encoding of the Chinese character structure components of the characters, radical matching determines candidate characters and / or candidate words based on the radicals of the characters, and shape matching determines candidate characters and / or candidate words based on the shape structure of the characters.
[0222] In a possible implementation manner, using any one or more of the matching methods of pinyin matching, four-corner code matching, radical matching, and shape matching to determine candidate characters and / or candidate words in the recall set includes:
[0223] Search on one or more search trees respectively based on the target character to be corrected, and in response to the existence of a matching node in the search tree, use the character corresponding to the matching node as the candidate character; wherein, each search tree corresponds to a matching method respectively.
[0224] In a possible implementation, use any one or more of pinyin matching, four-corner code matching, radical matching, and shape matching to determine the candidate characters and / or candidate words in the recall set, including:
[0225] Search on one or more search trees respectively based on the target character to be corrected and a preset number of characters after the target character to be corrected;
[0226] In response to the existence of a matching node in the search tree, use the word corresponding to the matching node as the candidate word, and after increasing the preset number, re-execute the steps of searching on one or more search trees respectively based on the target character to be corrected and a preset number of characters after the target character to be corrected and the subsequent steps until there are no matching nodes in each search tree, and obtain the recall set.
[0227] In a possible implementation, the second determination module 702 is used for:
[0228] Determine the first vector representation corresponding to the target character to be corrected;
[0229] Determine the second vector representations corresponding to the respective characters and / or words in the recall set;
[0230] Based on the vector representation and the second vector representations, use the knowledge selection model to determine the corrected result of the target character to be corrected, and the knowledge selection model is constructed based on the attention mechanism.
[0231] In a possible implementation, the knowledge selection model is a trained knowledge selection model, and the device further includes:
[0232] An acquisition module, configured to acquire a training set, where the training set includes training texts and labels, and the labels are used to indicate the actual corrected texts corresponding to the training texts;
[0233] A third determination module, configured to determine the third vector representation corresponding to the training character to be corrected in the training text;
[0234] A fourth determination module, configured to determine the fourth vector representations corresponding to the respective characters and / or words in the recall set corresponding to the training character to be corrected;
[0235] A calculation module, configured to calculate the attention weights corresponding to the respective fourth vector representations respectively based on the third vector representation and the fourth vector representations by using the initial knowledge selection model;
[0236] A first weighting module, configured to weight each fourth vector representation based on the attention weights respectively corresponding to each fourth vector representation to obtain a third fused vector representation;
[0237] A second weighting module, configured to weight based on the third vector representation and the third fused vector representation to obtain a fourth fused vector representation;
[0238] An optimization module, configured to optimize the parameters of the initial knowledge selection model based on the attention weights respectively corresponding to each fourth vector representation, the fourth fused vector representation, and the label, to obtain a trained knowledge selection model.
[0239] In a possible implementation manner, based on the first vector representation and the second vector representation, using the knowledge selection model to determine the corrected result of the target character to be corrected, including:
[0240] Based on the first vector representation and the second vector representation, using the knowledge selection model to calculate the attention weights respectively corresponding to each second vector representation;
[0241] Weight each second vector representation based on the attention weights respectively corresponding to each second vector representation to obtain a first fused vector representation;
[0242] Weight based on the first vector representation and the first fused vector representation to obtain a second fused vector representation;
[0243] Based on the second fused vector representation, determine the corrected result from the recall set.
[0244] In a possible implementation manner, the corrected result is the target candidate character or the target candidate word in the recall set, and the correction module 703 is configured to:
[0245] Replace the target character to be corrected in the Chinese text to be corrected with the target candidate character; or,
[0246] Replace the target character to be corrected and one or more characters after the target character to be corrected in the Chinese text to be corrected with the target candidate word.
[0247] According to another aspect of the present disclosure, there is provided a Chinese error correction device, including a memory, a processor, and a computer program stored on the memory, and the processor executes the computer program to implement the steps of the above method.
[0248] According to another aspect of the present disclosure, there is provided a non-volatile computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above method are implemented.
[0249] According to another aspect of the present disclosure, there is provided a computer program product, including a computer program or a non-volatile computer-readable storage medium carrying the computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.
[0250] According to an embodiment of the present disclosure, by determining a recall set corresponding to a target character to be corrected in a Chinese text to be corrected, based on the target character to be corrected and the recall set, a corrected result is determined from the recall set, and based on the corrected result, the Chinese text to be corrected is corrected. Among them, the recall set includes the target character to be corrected and / or candidate characters, and the candidate characters have similarity with the target character to be corrected in terms of pronunciation and / or glyph, which can enable the Chinese error correction process to have knowledge retrieval ability, be able to efficiently integrate external knowledge, greatly enrich the knowledge injection amount, provide sufficient candidates for subsequent accurate error correction, and thus improve the accuracy and generalization ability of Chinese error correction.
[0251] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the methods described in the above method embodiments. Its specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.
[0252] The embodiments of the present disclosure further provide a Chinese error correction device, including a memory, a processor, and a computer program stored on the memory, and the processor executes the computer program to implement the steps of the above method.
[0253] The embodiments of the present disclosure further provide a non-volatile computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above method are implemented.
[0254] The embodiments of the present disclosure further provide a computer program product, including a computer program or a non-volatile computer-readable storage medium carrying the computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.
[0255] Figure 8 is a block diagram of a device 1900 for Chinese error correction shown according to an exemplary embodiment. For example, the device 1900 can be provided as a server or a terminal device. Refer to Figure 8 , the device 1900 includes a processing component 1922, which further includes one or more processors, and memory resources represented by a memory 1932 for storing instructions executable by the processing component 1922, such as application programs. The application programs stored in the memory 1932 can include one or more modules each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute instructions to execute the above method.
[0256] Device 1900 may also include a power supply component 1926 configured to perform power management of device 1900, a wired or wireless network interface 1950 configured to connect device 1900 to a network, and an input / output interface 1958 (I / O interface). Device 1900 may operate based on an operating system stored in memory 1932, such as Windows Server TM , MacOS X TM , Unix TM , Linux TM , FreeBSD TM or the like.
[0257] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as memory 1932 including computer program instructions, and the computer program instructions can be executed by the processing component 1922 of device 1900 to complete the above method.
[0258] A computer-readable storage medium may be a tangible device that can hold and store programs / instructions used by an instruction execution device. A computer-readable storage medium may be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the above. More specific examples (non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital versatile disc (DVD), memory stick, floppy disk, mechanically encoded device, such as a punched card or raised structures in grooves storing instructions thereon, and any suitable combination of the above. A computer-readable storage medium as used herein is not construed as an instantaneous signal itself, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission medium (e.g., optical pulses through an optical fiber cable), or electrical signals transmitted through wires.
[0259] The computer programs (or computer-readable program instructions) described herein can be downloaded from a computer-readable storage medium to various computing / processing devices, or downloaded to an external computer or external storage device through a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission cables, fiber optic transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in the computer-readable storage media in each computing / processing device.
[0260] A computer program (or computer program instructions) for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine - related instructions, microcode, firmware instructions, state - setting data, or source code or object code written in any combination of one or more programming languages, including object - oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer - readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand - alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via an Internet service provider through the Internet). In some embodiments, by using the state information of the computer - readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field - programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer - readable program instructions to implement various aspects of the present disclosure.
[0261] Aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer - readable program instructions.
[0262] These computer - readable program instructions can be provided to a processor of a general - purpose computer, a special - purpose computer, or other programmable data - processing apparatus to produce a machine such that the instructions, when executed by the processor of the computer or other programmable data - processing apparatus, create a means for implementing the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer - readable program instructions can also be stored in a computer - readable storage medium, which causes a computer, a programmable data - processing apparatus, and / or other devices to operate in a particular manner, so that the computer - readable medium storing the instructions includes a manufacture, which includes instructions for implementing various aspects of the functions / acts specified in one or more blocks of the flowchart and / or block diagram.
[0263] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other devices, causing a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other devices to generate a computer-implemented process, such that the instructions executed on the computer, other programmable data processing apparatus, or other devices implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.
[0264] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of code, or a portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending upon the functionality involved. It should also be noted that each block of the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or acts, or by a combination of dedicated hardware and computer instructions.
[0265] The embodiments of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, the practical application, or improvements made to the technology in the marketplace, or to enable other ordinary skilled artisans in the art to understand the embodiments disclosed herein.
Claims
1. A Chinese error correction method, characterized in that: The method comprises: Determine a recall set corresponding to a target character to be corrected in the Chinese text to be corrected, wherein the recall set includes the target character to be corrected and / or candidate characters, and the candidate characters have similarities with the target character to be corrected in pronunciation and / or shape; Based on the target character to be corrected and the recalled set, determining a corrected result from the recalled set; Based on the error-corrected result, the Chinese text to be corrected is corrected.
2. The method according to claim 1, characterized in that The recall set also includes candidate words, and each character in the candidate words corresponds to a target character to be corrected and one or more characters after the target character to be corrected in the Chinese text to be corrected.
3. The method according to claim 2, characterized in that The step of determining a recall set corresponding to a target character to be corrected in the Chinese text to be corrected includes: Determine the candidate characters and / or candidate words in the recall set by using any one or more matching methods of pinyin matching, four-corner code matching, radical matching, and shape matching; Among them, the pinyin matching determines candidate characters and / or candidate words based on the initial consonants and finals of the characters, the four-corner code matching determines candidate characters and / or candidate words based on the Chinese character structure component codes of the characters, the radical matching determines candidate characters and / or candidate words based on the radicals of the characters, and the shape matching determines candidate characters and / or candidate words based on the shape structure of the characters.
4. The method according to claim 3, characterized in that The method of determining the candidate characters and / or candidate words in the recall set by using any one or more matching methods of pinyin matching, four-corner code matching, radical matching, and shape matching includes: Based on the target character to be corrected, a search is performed on one or more search trees respectively, and in response to the existence of a matching node in the search tree, the character corresponding to the matching node is used as a candidate character; wherein each search tree corresponds to a matching method.
5. The method according to claim 4, characterized in that The method of determining the candidate characters and / or candidate words in the recall set by using any one or more matching methods of pinyin matching, four-corner code matching, radical matching, and shape matching includes: Based on the target character to be corrected and a preset number of characters after the target character to be corrected, searching on one or more search trees respectively; In response to the existence of a matching node in the search tree, the word corresponding to the matching node is used as a candidate word, and after increasing the preset number, the search is re-executed based on the target character to be corrected and the preset number of characters after the target character to be corrected, and the search and subsequent steps are performed on one or more search trees respectively, until there are no matching nodes in each search tree, and a recalled set is obtained.
6. The method according to claim 1, characterized in that The step of determining a correction result from the recalled set based on the target character to be corrected and the recalled set includes: Determine a first vector representation corresponding to the target character to be corrected; Determine a second vector representation corresponding to each character and / or word in the recall set; Based on the first vector representation and the second vector representation, a knowledge selection model is used to determine a correction result of a target character to be corrected, and the knowledge selection model is constructed based on an attention mechanism.
7. The method according to claim 6, characterized in that Based on the first vector representation and the second vector representation, determining the error correction result of the target character to be corrected by using the knowledge selection model includes: Based on the first vector representation and the second vector representation, using the knowledge selection model to calculate the attention weights corresponding to the second vector representations respectively; Weighting each second vector representation based on the attention weights corresponding to each second vector representation to obtain a first fused vector representation; Obtain a second fused vector representation by weighting the first vector representation and the first fused vector representation; Based on the second fusion vector representation, a correction result of the target character to be corrected is determined from the recall set.
8. The method according to claim 2, characterized in that: The error-corrected result is a target candidate character or a target candidate word in the recall set, and the correcting of the Chinese text to be corrected based on the error-corrected result includes: replacing the target character to be corrected in the Chinese text to be corrected with the target candidate character; or, The target character to be corrected in the Chinese text to be corrected and one or more characters after the target character to be corrected are replaced with the target candidate word.
9. The method according to claim 6, characterized in that The knowledge selection model is a trained knowledge selection model, and the method further includes: Obtaining a training set, wherein the training set includes training text and a label, wherein the label is used to indicate the actual error-corrected text corresponding to the training text; Determine a third vector representation corresponding to the training character to be corrected in the training text; Determine a fourth vector representation corresponding to each character and / or word in a recall set corresponding to the training character to be corrected; Based on the third vector representation and the fourth vector representation, using an initial knowledge selection model to calculate the attention weights corresponding to the fourth vector representations; Weighting each fourth vector representation based on the attention weights respectively corresponding to each fourth vector representation to obtain a third fused vector representation; Obtain a fourth fused vector representation after weighting based on the third vector representation and the third fused vector representation; Based on the attention weights respectively corresponding to the fourth vector representations, the fourth fusion vector representation and the label, the parameters of the initial knowledge selection model are optimized to obtain a trained knowledge selection model.
10. A Chinese error correction device, characterized in that: The device comprises: A first determination module is used to determine a recall set corresponding to a target character to be corrected in a Chinese text to be corrected, wherein the recall set includes the target character to be corrected and / or candidate characters, and the candidate characters have similarities with the target character to be corrected in pronunciation and / or shape; A second determination module is used to determine a correction result from the recall set based on the target character to be corrected and the recall set; A correction module is used to correct the Chinese text to be corrected based on the correction result.
11. A Chinese error correction device, comprising a memory, a processor and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 9.
12. A non-volatile computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 9 are implemented.
13. A computer program product, comprising a computer program, or a non-volatile computer-readable storage medium carrying a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 9 are implemented.
Citation Information
Patent Citations
Text error correction method, device and equipment
CN111931489A
Search text processing method, device, electronic equipment and medium
CN113535895A
Error correction method and device for Chinese text in power field, storage medium and computing equipment
CN114118065A
Text error correction method and device, equipment and storage medium
CN116702761A
Model training method, chinese text error correction method, electronic device, and storage medium
WO2023093525A1