Method, system and equipment for correcting wrongly written characters and storage medium
By replacing the typos in the text to be corrected with a mask, and calling the correction model respectively to correct the mask text and the original text, combining the two correction results to determine the correct words, the problems of inaccurate correction results and semantic changes in the prior art are solved, and the typos correction effect with more accurate and semantic reservations is achieved.
Patent Information
- Application Number
- CN202311738212.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-15
- Publication Date
- 2025-06-17
AI Technical Summary
The existing Chinese typo correction methods are easily affected by other wrong Chinese characters associated with wrong Chinese characters, resulting in inaccurate correction results; while the method of first checking errors and then correcting errors only considers the location of the wrong Chinese characters, and does not consider the information of the original Chinese characters, resulting in a change in the text semantics after correction.
By replacing the typos in the text to be corrected with a mask, the mask text is generated, and the mask text and the original text are corrected by calling the first correction model and the second correction model respectively, the correct word corresponding to the typos is determined based on the two correction results.
It improves the accuracy of typo correction, avoids the influence of associated errors, and takes into account the information of the original Chinese characters to ensure that the text semantics after the error is corrected remains unchanged.
Smart Images

Figure CN120163145A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the field of natural language processing, and more particularly to a method, system, device, and storage medium for correcting typos. Background Art
[0002] Chinese typo correction is widely used to correct phonetically similar and visually similar errors in the input text without changing the length of the input text. There are many methods for Chinese typo correction, which can be roughly divided into the following two types according to the combination of error detection and error correction:
[0003] (1) Direct error correction. Directly predict the Chinese character corresponding to each Chinese character position. If the predicted Chinese character is different from the original Chinese character, it is considered that the original Chinese character is incorrect. However, this method is easily affected by other incorrect Chinese characters associated with the incorrect Chinese character, resulting in the predicted Chinese character being possibly the same as the original Chinese character.
[0004] (2) Error detection first and then error correction. First, determine whether each Chinese character in the input text is incorrect. For the incorrect positions, then determine which Chinese character should be replaced. However, this method only considers the position of the incorrect Chinese character and does not consider the information of the original Chinese character, resulting in a change in the semantics of the text after error correction. Summary of the Invention
[0005] Therefore, embodiments of the present application provide a method, system, device, and storage medium for correcting typos. According to the correction result of the text to be corrected without considering typos and the correction result of the text to be corrected considering typos, the correct character corresponding to the typo in the text to be corrected can be determined more accurately.
[0006] To achieve the above object, embodiments of the present application provide the following technical solutions:
[0007] According to a first aspect of embodiments of the present application, there is provided a method for correcting typos, the method comprising:
[0008] Replacing the typos in the text to be corrected with masks to generate a masked text corresponding to the text to be corrected;
[0009] Invoking a first correction model to correct the masked text to obtain a first correction result;
[0010] Invoking a second correction model to correct the text to be corrected to obtain a second correction result;
[0011] Determining the correct character corresponding to the typo in the text to be corrected according to the first correction result and the second correction result.
[0012] Optionally, the replacing the typos in the text to be corrected with masks to generate a masked text corresponding to the text to be corrected includes:
[0013] Extract the character features and semantic features of each character in the text to be corrected;
[0014] Determine the position of the misspelled character in the text to be corrected according to the character features and semantic features of each character;
[0015] Replace the misspelled character with the mask to obtain the masked text.
[0016] Optionally, the determining the position of the misspelled character in the text to be corrected according to the character features and semantic features of each character includes:
[0017] Calculate the error probability of each character according to the character features and semantic features of each character;
[0018] Determine whether the character is a misspelled character based on the error probability of each character and the set misspelled character threshold;
[0019] After determining the misspelled character, obtain the position of the misspelled character in the text to be corrected.
[0020] Optionally, calling a first correction model to correct the masked text to obtain a first correction result includes:
[0021] Extract the semantic features corresponding to the mask based on the first correction model;
[0022] Calculate the first predicted probability distribution corresponding to the mask according to the semantic features of the mask. The first predicted probability distribution corresponding to the mask includes the probabilities of each word in the preset word list replacing the mask, and the first predicted probability distribution is the first correction result.
[0023] Optionally, calling a second correction model to correct the masked text to obtain a second correction result includes:
[0024] Extract the semantic features corresponding to the misspelled character based on the second correction model;
[0025] Calculate the second predicted probability distribution corresponding to the misspelled character according to the semantic features of the misspelled character. The second predicted probability distribution corresponding to the misspelled character includes the probabilities of each word in the preset word list replacing the misspelled character, and the second predicted probability distribution is the second correction result.
[0026] Optionally, the determining the correct character corresponding to the misspelled character in the text to be corrected according to the first correction result and the second correction result includes:
[0027] Determine the predicted probability distribution of this position according to the first predicted probability distribution corresponding to the position of the mask and the second predicted probability distribution corresponding to the misspelled character;
[0028] Determine the target probability of the position according to the predicted probability distribution of the position;
[0029] Determine the corrected character corresponding to the position based on a preset vocabulary and the target probability.
[0030] Optionally, after replacing the misspelled characters in the text to be corrected with masks to generate a masked text corresponding to the text to be corrected, the method further includes:
[0031] Call the first correction model to calculate the first probability distribution of the position of each character in the masked text;
[0032] Call the second correction model to calculate the second probability distribution of the position of each character in the text to be corrected;
[0033] Determine the target probability of the position according to the first probability distribution and the second probability distribution of the position of each character;
[0034] Compare one by one the target probability of the position corresponding to each character with the character at the corresponding position in the text to be corrected in the vocabulary corresponding to the preset vocabulary. If they do not match, update the character at the corresponding position in the text to be corrected to the character corresponding to the preset vocabulary until all character positions are compared.
[0035] Optionally, the method further includes:
[0036] Train an initial model with a training sample set to obtain a correction model, where the correction model includes the error detection model, the first correction model, and the second correction model. The training sample set includes multiple training sample pairs, and each training sample pair includes a misspelled text and a masked text corresponding to the misspelled text, and the position of the misspelled character in the masked text is a mask;
[0037] The training process includes:
[0038] Input the training sample pair into the initial model to obtain a training result; the training result includes an error detection training result and a misspelled character correction training result;
[0039] Calculate the cross-entropy total loss according to the error detection loss and the misspelled character correction loss. The error detection loss represents the loss between the error detection training result of the position of each character and a preset error detection annotation result, and the misspelled character correction loss refers to the loss between the misspelled character correction training result of the position of each character and a preset misspelled character correction annotation result;
[0040] If the cross-entropy total loss is less than or equal to a preset threshold, determine the initial model as the correction model;
[0041] If the total cross-entropy loss is greater than the preset threshold, update the parameters of the initial model to obtain a new initial model, and perform again the step of inputting the training sample pair into the initial model to obtain a training result.
[0042] According to a second aspect of the embodiments of the present application, there is provided a typo correction system, the system comprising:
[0043] An error detection module, configured to replace typos in the text to be corrected with masks to generate a masked text corresponding to the text to be corrected;
[0044] A first correction module, configured to call a first correction model to correct the masked text to obtain a first correction result;
[0045] A second correction module, configured to call a second correction model to correct the text to be corrected to obtain a second correction result;
[0046] A result fusion module, configured to determine the correct characters corresponding to the typos in the text to be corrected according to the first correction result and the second correction result.
[0047] According to a third aspect of the embodiments of the present application, there is provided an electronic device, comprising: a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor runs the computer program, it is configured to implement the method described in the first aspect above.
[0048] According to a fourth aspect of the embodiments of the present application, there is provided a computer-readable storage medium, on which computer-readable instructions are stored, and the computer-readable instructions can be executed by a processor to implement the method described in the first aspect above.
[0049] In summary, the embodiments of the present application provide a typo correction method, system, device, and storage medium. By replacing typos in the text to be corrected with masks to generate a masked text corresponding to the text to be corrected; calling a first correction model to correct the masked text to obtain a first correction result; calling a second correction model to correct the text to be corrected to obtain a second correction result; and determining the correct characters corresponding to the typos in the text to be corrected according to the first correction result and the second correction result. According to the correction result of the text to be corrected without considering typos and the correction result of the text to be corrected considering typos, the correct characters corresponding to the typos in the text to be corrected can be determined more accurately. Description of the Drawings
[0050] To more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only exemplary. For those of ordinary skill in the art, without creative efforts, other implementation drawings can also be obtained based on the provided drawings.
[0051] The structures, ratios, sizes, etc. shown in this specification are only used to cooperate with the content disclosed in the specification for those who are familiar with this technology to understand and read, and are not used to limit the limiting conditions for the implementation of the present invention. Therefore, they do not have a substantial technical meaning. Any modification of the structure, change in the proportional relationship, or adjustment of the size, without affecting the effects that the present invention can produce and the purposes that can be achieved, should still fall within the scope covered by the technical content disclosed in the present invention.
[0052] Figure 1 It is a schematic flow chart of a typo correction method provided by an embodiment of the present application;
[0053] Figure 2 It is a schematic flow chart of the typo correction process provided by an embodiment of the present application;
[0054] Figure 3 It is a schematic flow chart of the typo correction model training process provided by an embodiment of the present application;
[0055] Figure 4 It is a block diagram of a typo correction system provided by an embodiment of the present application;
[0056] Figure 5 It shows a schematic structural diagram of an electronic device provided by an embodiment of the present application;
[0057] Figure 6 It shows a schematic diagram of a computer-readable storage medium provided by an embodiment of the present application. Specific Embodiments
[0058] The following specific embodiments illustrate the embodiments of the present invention. Those who are familiar with this technology can easily understand other advantages and effects of the present invention from the content disclosed in this specification. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of them. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope protected by the present invention.
[0059] In the prior art, if the predicted Chinese character is different from the original Chinese character, the original Chinese character is considered to be wrong. However, this direct error correction method is easily affected by other wrong Chinese characters associated with the wrong Chinese character. For example, in "现在放假真的太高了", "放" and "假" are both wrong characters. When correcting the Chinese character in the position of "放", it may be affected by the wrong Chinese character "假" and not be corrected.
[0060] In the prior art, the error detection and correction scheme first determines whether each Chinese character in the input text is wrong, then masks the wrong position, and determines which correct character should be replaced by the masked position in the masked text. However, this method only considers the position of the wrong Chinese character, and does not consider the information of the original Chinese character, and cannot use the information of the original Chinese character, resulting in a change in the semantics of the text after error correction. Continuing with the example of "The current holiday is too high", the error position is first obtained through error detection, and the mask [MASK] is used to replace the error position, and the text after error detection is output as "The current [MASK][MASK] is too high". Further, the [MASK] position is corrected, because it does not contain the original Chinese character information, it may be changed to "The current oil price is too high" and so on.
[0061] In order to improve the problems of the first two methods, there is a solution in the prior art that detects errors first and then corrects them. First, it is determined whether each Chinese character in the input text is wrong, and then which Chinese character should be replaced for the wrong position. In the process of direct error correction, the error detection information is implicitly used. However, this method only considers the position of the wrong Chinese character, not the information of the original Chinese character, which leads to the change of the semantics of the text after error correction.
[0062] Based on the current state of the art, this application proposes a method for correcting wrong characters. First, the original text is checked for errors to obtain a masked text constructed after the error detection. In order to not be affected by other wrong Chinese characters in the associated context and to consider the information of the original Chinese characters at the wrong character position, the masked text and the original text are respectively input into the correction model to obtain the correction results without considering the wrong characters and considering the wrong characters. The two correction results are then summarized to obtain the final correction result.
[0063] Figure 1 The method for correcting typos provided in the embodiment of the present application is shown, which specifically includes the following steps:
[0064] Step 101: replacing the wrong characters in the text to be corrected with a mask to generate a mask text corresponding to the text to be corrected;
[0065] Step 102: calling a first correction model to correct the masked text to obtain a first correction result;
[0066] Step 103: calling a second correction model to correct the text to be corrected, and obtaining a second correction result;
[0067] Step 104: Determine the correct character corresponding to the wrong character in the text to be corrected according to the first correction result and the second correction result.
[0068] In a possible implementation, in step 101, replacing the wrong characters in the text to be corrected with a mask to generate a mask text corresponding to the text to be corrected includes:
[0069] Extract the character features and semantic features of each character in the text to be corrected; determine the position of the wrong character in the text to be corrected according to the character features and semantic features of each character; replace the wrong character with a mask to obtain the masked text.
[0070] The character feature of each character is a feature containing original Chinese character information based on the output of the 0th layer of the error detection model. The semantic feature of each character is a feature of predicted Chinese character information inferred from context information based on the output of the setting layer of the error detection model.
[0071] The mask involved in the embodiment of the present application is a means to shield wrong characters, and [MASK] is used as a form of expression of the mask in the embodiment of the present application. For example, in the text to be corrected, "This is a good example of return", it is detected that "返" is a wrong character in the text to be corrected, then [MASK] is used to replace the word to shield the wrong character, and a mask text corresponding to the text to be corrected is generated, "This is a good [MASK] example". It should be noted that the means of shielding wrong characters in the embodiment of the present application is not limited to masks, and the form of expression of masks is not limited to [MASK].
[0072] In a possible implementation manner, in the above step of determining the position of the wrong character in the text to be corrected according to the character features and semantic features of each character, the following steps are specifically included:
[0073] The error probability of each character is calculated according to the character features and semantic features of each character; whether each character is a wrong character is determined based on the error probability of each character and a set wrong character threshold; after determining that a wrong character is present, the position of the wrong character in the text to be corrected is obtained.
[0074] Specifically, the character features and semantic features of each character are concatenated to obtain the error probability of each character. If the error probability is greater than a set wrong character threshold, the character corresponding to the error probability is determined to be a wrong character. The position of the wrong character in the text to be corrected can be further obtained to facilitate the subsequent step of replacing the wrong character at that position.
[0075] In a possible implementation, in step 102, calling a first correction model to correct the mask text to obtain a first correction result includes:
[0076] Extract the semantic features corresponding to the mask based on the first correction model; calculate the first predicted probability distribution corresponding to the mask according to the semantic features of the mask, where the first predicted probability distribution corresponding to the mask includes the probabilities of each word in a preset word list replacing the mask, and the first predicted probability distribution is the first correction result. The preset word list consists of several words and the probabilities corresponding to the current positions of these words. The sum of the probabilities corresponding to the several words at the current position is 1.
[0077] In a possible implementation manner, in step 103, call a second correction model to correct the masked text to obtain a second correction result, including:
[0078] Extract the semantic features corresponding to the misspelled word based on the second correction model; calculate the second predicted probability distribution corresponding to the misspelled word according to the semantic features of the misspelled word, where the second predicted probability distribution corresponding to the misspelled word includes the probabilities of each word in a preset word list replacing the misspelled word, and the second predicted probability distribution is the second correction result.
[0079] It should be noted that the misspelled word is the word that replaces the mask, and it may be a misspelled word or a correct word.
[0080] In a possible implementation manner, in step 104, the step of determining the correct word corresponding to the misspelled word in the text to be error-corrected according to the first correction result and the second correction result includes:
[0081] Determine the predicted probability distribution of this position according to the first predicted probability distribution corresponding to the position of the mask and the second predicted probability distribution corresponding to the misspelled word; determine the target probability of this position according to the predicted probability distribution of this position; determine the corrected word corresponding to this position based on the preset word list and the target probability.
[0082] In a possible implementation manner, take the average of the first predicted probability distribution corresponding to the position of the mask and the second predicted probability distribution corresponding to the misspelled word to obtain the predicted probability distribution of this position.
[0083] By specifically fusing the first predicted probability distribution at the mask position and the second predicted probability distribution, the corrected word at this mask position can be predicted more accurately. However, the position of the mask is determined based on the error detection module. To further improve the error detection efficiency of misspelled words, the embodiments of the present application also provide a method to fuse the first probability distribution and the second probability distribution at each position in the text to be corrected to obtain the correction result of the entire text to be corrected, and then further determine the corrected word.
[0084] Specifically, after replacing the misspelled words in the text to be corrected with masks to generate a masked text corresponding to the text to be corrected, the method further includes:
[0085] Invoking a first correction model to calculate a first probability distribution of the position of each character in the masked text; invoking a second correction model to calculate a second probability distribution of the position of each character in the text to be corrected; determining a target probability for that position according to the first probability distribution and the second probability distribution of the position of each character.
[0086] Compare one by one the target probability of the position corresponding to each character with the character at the corresponding position in the text to be corrected in the corresponding word list of the preset word list. If they do not match, update the character at the corresponding position in the text to be corrected to the character corresponding in the preset word list until all character positions are compared.
[0087] In the above method, it may specifically include the following steps:
[0088] Step 1: Extract semantic features of each character in the masked text based on the first correction model; calculate a first probability distribution of the position of each character in the masked text according to the semantic features of each character in the masked text.
[0089] Step 2: Obtain semantic features of each character in the text to be corrected based on the second correction model; calculate a second probability distribution of the position of each character in the text to be corrected according to the semantic features of each character in the text to be corrected.
[0090] Step 3: Determine the target probability of the position corresponding to each character according to the first probability distribution and the second probability distribution of the position corresponding to each character; for example, calculate the mean value.
[0091] Step 4: Obtain the character corresponding to the target probability of the position corresponding to each character in the corresponding word list of the preset word list.
[0092] Step 5: Compare one by one the target probability of the position corresponding to each character with the character at the corresponding position in the text to be corrected in the corresponding word list of the preset word list. If they do not match, update the character at the corresponding position in the text to be corrected to the character corresponding in the preset word list until all character positions are compared.
[0093] Fusing the correction results without considering the original incorrect Chinese character information and the correction results considering the original incorrect Chinese character information improves the correction efficiency of text misspelled words.
[0094] In a possible implementation, the method further includes: training an initial model with a training sample set to obtain an error correction model, where the error correction model includes the error detection model, the first correction model, and the second correction model, the training sample set includes a plurality of training sample pairs, and each training sample pair includes a misspelled text and a masked text corresponding to the misspelled text, and the positions of the misspelled characters in the masked text are masked.
[0095] In a possible implementation, the above training process includes:
[0096] Step 1: Input the training sample pair into the initial model to obtain a training result; the training result includes an error detection training result and a misspelled character correction training result; the error detection training result includes the error detection training results of the positions of a plurality of characters; the misspelled character correction training result includes the correction training results of the positions of a plurality of characters.
[0097] Step 2: Calculate the cross-entropy total loss according to the error detection loss and the misspelled character correction loss. The error detection loss represents the loss between the error detection training result of each character position and a preset error detection annotation result, and the misspelled character correction loss refers to the loss between the misspelled character correction training result of each character position and a preset misspelled character correction annotation result; the error detection annotation result is the correct annotation result of this position set in advance, and the correction annotation result is the correct annotation result of this position set in advance.
[0098] Step 3: If the cross-entropy total loss is less than or equal to a preset threshold, determine the initial model as the error correction model; if the cross-entropy total loss is greater than the preset threshold, update the parameters of the initial model to obtain a new initial model, and execute the step of inputting the training sample pair into the initial model to obtain a training result again.
[0099] Figure 2 Shows the end-to-end misspelled character detection and correction model structure provided by the embodiments of the present application.
[0100] Assume that the text to be corrected is "This is a very good re example". On the one hand, perform an error detection task, and construct a masked text after error detection on the text to be corrected through the detection module: "This is a very good [MASK] example", and further input the masked text into the correction model 1 to obtain a first correction result without considering the information of the original incorrect Chinese characters. On the other hand, input the text to be corrected into the correction model 2 to obtain a second correction result considering the original incorrect Chinese characters; finally, summarize the two results to obtain the corrected text "This is a very good counterexample".
[0101] Among them, the correction model 1 and the correction model 2 can share model parameters or be independent of each other.
[0102] Based on the above typo detection and correction model, Figure 3 FIG. Figure 3 shows a schematic diagram of the typo correction model construction process provided by an embodiment of the present application, which mainly includes the following stages:
[0103] The first stage: Obtain a typo correction training set as a training sample set; the training sample set includes a plurality of training sample pairs, and each training sample pair includes a typo text and a masked text corresponding to the typo text. The position of the typo in the masked text is masked. The lengths of the two texts are the same.
[0104] For example: "I have seen many brave people who struggle without fear of setbacks. This spirit is worthy of our learning." and "I have seen many brave people who struggle without fear of frustration. This spirit is worthy of our learning."
[0105] The second stage: Based on the detection module, perform text error detection on the text to be corrected to obtain an error detection output;
[0106] Call the detection module to perform error detection on the context of the input text to determine whether the Chinese character at the current position is used incorrectly. The text to be corrected input is s, and the length of the text to be corrected s is m. For the position i of each character in the text to be corrected, the character features and semantic features of each Chinese character are concatenated and then error detection is performed to output the probability of Chinese character error d refers to the detect detection module. The detection module can be a binary classification network, such as network structures like lstm, cnn, transformer layer, etc. The detection module in the embodiment of the present application adopts the bert model. Each layer of the bert model contains context information. The more layers, the more parameters, and the greater the training cost. Different set layers are selected for feature output according to different actual application requirements.
[0107] Specifically, it includes the following steps:
[0108] (1) Calculate the character feature e of the text to be corrected s (The character feature contains the original Chinese character information):
[0109]
[0110] The character feature is a feature containing the original Chinese character information output based on the set layer of the bert model, and the set layer here can be set to the 0th layer according to the requirement of the original Chinese character information.
[0111] (2) Calculate the semantic feature h of the text to be corrected s (The semantic feature contains the correct Chinese character information):
[0112]
[0113] The semantic feature is a feature of the predicted Chinese character information inferred based on context information based on the setting layer output of the BERT model.
[0114] (3) Combine the character features and semantic features of Chinese characters to further calculate the probability of Chinese character errors
[0115]
[0116] Among them, w d Represents the network weight parameters of the BERT model, represents the semantic feature of a layer in the error detection model at the ith position, sigmoid is the sigmoid function, represents the word feature at the i-th position, b d Represents the network bias parameters of the BERT model.
[0117] When the probability If the value is greater than the set threshold thd, it is determined to be a typo, and error_ids represents the position set of the Chinese characters detected as typos:
[0118]
[0119] (4) Replace the Chinese characters that are judged to be wrong with the set characters (mask) to obtain the masked text. For example, if the input text is "this is a good example of return", and the error probability corresponding to "返" is detected to be greater than the threshold, then [MASK] is used to replace the Chinese character, and the original input becomes "this is a good [MASK] example", which is recorded as the masked text q.
[0120] The third stage: text error correction is performed based on correction model 1 and correction model 2 respectively;
[0121] Step 1: Input the masked text q obtained in the second stage into the correction model 1 to correct the wrong characters, and obtain the semantic features of the position of each word in the entire text and the probability distribution of the position of each word, the probability distribution p i q To predict the probability distribution of replacing a certain Chinese character as the correction result without considering the original erroneous Chinese character information.
[0122] The error correction model 1 can be any existing typo correction model. In the embodiment of the present application, a model composed of a BERT model and a fully connected classification layer is used as an example of a typo correction model.
[0123]
[0124]
[0125] Among them, hq represents the set of semantic features of the position of each character after being represented by the bert model, and w1 represents the weight parameter of the correction model 1; represents the semantic feature of the i-th character position; b1 represents the bias parameter of the correction model 1; Softmax represents the softmax function.
[0126] Step 2: Input the text s to be corrected into the correction model 2 for correcting typos, and predict the probability distribution of replacing it with a certain Chinese character as the correction result including the original wrong Chinese character information.
[0127] The correction model 2 can use the same structure as the correction model 1, and the two share parameters, or different typo correction models can also be used.
[0128]
[0129]
[0130] Among them, h s represents the set of semantic features of the position of each character after being represented by the bert model. w2 represents the weight parameter of the correction model 2; represents the semantic feature of the i-th character position; Softmax represents the softmax function; b2 represents the bias parameter of the correction model 2.
[0131] Fourth stage: Combine the correction results to obtain the final output.
[0132] The final output is a probability distribution, representing the probability of each Chinese character at the current position; for example, if the preset vocabulary size is 3, the probability distribution may be [0.1, 0.2, 0.7].
[0133] The methods for combining the correction results include but are not limited to taking the mean. The method for combining the correction results provided in the embodiments of the present application is exemplified by taking the mean. For a certain position i of the input text, the probability distribution of replacing it with a certain Chinese character is p i :
[0134]
[0135] After averaging, the probability distribution values of the positions of each character in the entire text are obtained.
[0136] Fifth stage: Determine the maximum value in the probability distribution values as the target probability, and complete the text correction of the text to be corrected according to the target probability.
[0137] Step 5-1: Determine the maximum value in the probability distribution values as the target probability. For example, the maximum probability value in the above probability distribution [0.1, 0.2, 0.7] is 0.7.
[0138] Step 5-2: According to the word list relationship, obtain the characters corresponding to the target probability in the preset word list as the correct characters to be replaced. For example, obtain the characters in the word list corresponding to 0.7 according to the word list relationship as the correct characters to be replaced.
[0139] Step 5-3: Compare one by one the characters corresponding to the target probability at each character position in the preset word list with the characters at the corresponding position in the text to be corrected. If they do not match, update the character at the corresponding position in the text to be corrected with the character corresponding to the preset word list. If they match, do not change. Until all character positions are compared.
[0140] Sixth stage: Calculate the cross-entropy loss and train and optimize the model. The model includes the detection module, the correction model 1 and the correction model 2.
[0141] Step 6-1: Calculate the error detection loss I through the cross-entropy loss det :
[0142]
[0143] The error detection loss I det represents the loss between the error detection training result and the error detection annotation result at each character position. The error detection annotation result is the preset correct annotation result at this position. Among them, N is the number of sentences in the training set, and the value of d i is 0 or 1, representing that there is no error or there is an error at this position.
[0144] Step 6-2: Calculate the misspelling correction loss I through the cross-entropy loss cor :
[0145]
[0146] The misspelling correction loss I cor refers to the loss between the misspelling correction training result and the misspelling correction annotation result at each character position. The correction annotation result is the preset correct annotation result at this position. Among them, N is the number of sentences in the training set, and l is the index of the correct Chinese character at the i-th position in the word list.
[0147] Step 6-3: Perform weighted summation on the above error detection loss I det and the misspelling correction loss I cor to obtain the final loss:
[0148] l = l cor + λl det
[0149] During the training process, minimize the value of the loss function l. λ is a hyperparameter that can be set manually, and save the model with the best recognition effect on the validation set.
[0150] Step 6-4: Optimize the model according to the loss. If the total cross-entropy loss is less than or equal to the preset threshold, determine the initial model as the model to be optimized; if the total cross-entropy loss is greater than the preset threshold, update the parameters of the initial model to obtain a new initial model, and execute again the step of inputting the training sample pair into the initial model to obtain the training result. That is, input the training sample pair into the initial model to obtain the training result; the training result includes an error detection training result and a typo correction training result; the error detection training result includes error detection training results at the positions of multiple characters; the typo correction training result includes correction training results at the positions of multiple characters.
[0151] In summary, the embodiment of the present application provides a typo correction method, which replaces the typos in the text to be corrected with masks to generate a mask text corresponding to the text to be corrected; calls a first correction model to correct the mask text to obtain a first correction result; calls a second correction model to correct the text to be corrected to obtain a second correction result; determines the correct characters corresponding to the typos in the text to be corrected according to the first correction result and the second correction result. Based on the two correction results of correcting the text to be corrected and the text after error detection, the text is corrected for typos more accurately and efficiently.
[0152] Based on the same technical concept, the embodiment of the present application also provides a typo correction system, as Figure 4 shown, the system includes:
[0153] An error detection module 401, configured to replace the typos in the text to be corrected with masks to generate a mask text corresponding to the text to be corrected;
[0154] A first correction module 402, configured to call a first correction model to correct the mask text to obtain a first correction result;
[0155] A second correction module 403, configured to call a second correction model to correct the text to be corrected to obtain a second correction result;
[0156] A result fusion module 404, configured to determine the correct characters corresponding to the typos in the text to be corrected according to the first correction result and the second correction result.
[0157] In a possible implementation manner, the error detection module 401 is specifically configured to: extract the character features and semantic features of each character in the text to be corrected; determine the positions of the typos in the text to be corrected according to the character features and semantic features of each character; replace the typos with the masks to obtain the mask text.
[0158] In a possible implementation, determining the position of the misspelled word in the text to be corrected according to the character features and semantic features of each character in the error detection module 401 specifically includes: calculating the error probability of each character according to the character features and semantic features of each character; determining whether the character is a misspelled word based on the error probability of each character and a set misspelled word threshold; and after determining the misspelled word, obtaining the position of the misspelled word in the text to be corrected.
[0159] In a possible implementation, the first correction module 402 is specifically configured to:
[0160] Extract the semantic features corresponding to the mask; calculate the first predicted probability distribution corresponding to the mask according to the semantic features of the mask, where the first predicted probability distribution corresponding to the mask includes the probabilities of each word in a preset word list replacing the mask, and the first predicted probability distribution is the first correction result.
[0161] In a possible implementation, the second correction module 403 is specifically configured to:
[0162] Extract the semantic features corresponding to the misspelled word; calculate the second predicted probability distribution corresponding to the misspelled word according to the semantic features of the misspelled word, where the second predicted probability distribution corresponding to the misspelled word includes the probabilities of each word in a preset word list replacing the misspelled word, and the second predicted probability distribution is the second correction result.
[0163] In a possible implementation, the result fusion module 404 is specifically configured to:
[0164] Determine the predicted probability distribution at this position according to the first predicted probability distribution corresponding to the position of the mask and the second predicted probability distribution corresponding to the misspelled word; determine the target probability at this position according to the predicted probability distribution at this position; and determine the corrected word corresponding to this position based on the preset word list and the target probability.
[0165] In a possible implementation, the system further includes:
[0166] The first correction module 402 is further configured to: calculate the first probability distribution of the position of each character in the masked text;
[0167] The second correction module 403 is further configured to: call the second correction model to calculate the second probability distribution of the position of each character in the text to be corrected;
[0168] The result fusion module 404 is further configured to: determine the target probability of each position according to the first probability distribution and the second probability distribution of the position of each character; compare one by one the target probability of the position corresponding to each character with the character at the corresponding position in the text to be corrected in the preset word list, and if they do not match, update the character at the corresponding position in the text to be corrected to the character corresponding to the preset word list until the comparison of all character positions is completed.
[0169] In a possible implementation manner, the system further includes:
[0170] A training module, configured to train an initial model with a training sample set to obtain an error correction model, where the error correction model includes the first correction model and the second correction model, the training sample set includes a plurality of training sample pairs, and each training sample pair includes a misspelled word text and a mask text corresponding to the misspelled word text, and the position of the misspelled word in the mask text is a mask.
[0171] The training process includes: inputting the training sample pair into the initial model to obtain a training result, where the training result includes an error detection training result and a misspelled word correction training result; calculating a cross-entropy total loss according to the error detection loss and the misspelled word correction loss, where the error detection loss represents the loss between the error detection training result of each character position and a preset error detection annotation result, and the misspelled word correction loss refers to the loss between the misspelled word correction training result of each character position and a preset misspelled word correction annotation result; if the cross-entropy total loss is less than or equal to a preset threshold, determining the initial model as the error correction model; if the cross-entropy total loss is greater than the preset threshold, updating the parameters of the initial model to obtain a new initial model, and then executing the step of inputting the training sample pair into the initial model to obtain a training result again.
[0172] The embodiment of the present application further provides an electronic device corresponding to the method provided in the foregoing embodiment. Please refer to Figure 5 , which shows a schematic diagram of an electronic device provided in some embodiments of the present application. The electronic device 20 may include: a processor 200, a memory 201, a bus 202, and a communication interface 203, and the processor 200, the communication interface 203, and the memory 201 are connected through the bus 202; a computer program that can run on the processor 200 is stored in the memory 201, and when the processor 200 runs the computer program, it executes the method provided in any foregoing embodiment of the present application.
[0173] Among them, the memory 201 may include high-speed random access memory (RAM), and may also include non-volatile memory, such as at least one disk memory. The communication connection between the system network element and at least one other network element is implemented through at least one physical port 203 (which can be wired or wireless), and the Internet, wide area network, local area network, metropolitan area network, etc. can be used.
[0174] The bus 202 can be an ISA bus, a PCI bus, an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. Among them, the memory 201 is used to store a program, and after receiving an execution instruction, the processor 200 executes the program. Any implementation manner of the method disclosed in any implementation manner of the foregoing embodiments of the present application can be applied to the processor 200 or implemented by the processor 200.
[0175] The processor 200 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit in the hardware of the processor 200 or by instructions in the form of software. The above-mentioned processor 200 may be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art, such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 201, and the processor 200 reads the information in the memory 201 and combines its hardware to complete the steps of the above method.
[0176] The electronic device provided by the embodiments of the present application and the method provided by the embodiments of the present application are based on the same inventive concept and have the same beneficial effects as the method adopted, run, or implemented by it.
[0177] The embodiments of the present application also provide a computer-readable storage medium corresponding to the method provided by the foregoing embodiments. Please refer to Figure 6, which shows that the computer-readable storage medium is an optical disc 30, on which a computer program (i.e., program product) is stored. When the computer program is run by a processor, it will execute the method provided by any of the foregoing embodiments.
[0178] It should be noted that examples of the computer-readable storage medium may also include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, or other optical and magnetic storage media, which will not be elaborated here one by one.
[0179] The computer-readable storage medium provided by the above embodiments of the present application and the method provided by the embodiments of the present application are based on the same inventive concept and have the same beneficial effects as the methods adopted, run, or implemented by the application programs stored therein.
[0180] It should be noted that:
[0181] The algorithms and displays provided herein are not inherently related to any particular computer, virtual device, or other equipment. Various general-purpose devices can also be used in conjunction with the teachings herein. The structures required to construct such devices are obvious from the above description. In addition, the present application is not directed to any specific programming language. It should be understood that the content of the present application described herein can be implemented using various programming languages, and the description of the specific language above is to disclose the best implementation mode of the present application.
[0182] In the specification provided herein, a large number of specific details are set forth. However, it can be understood that the embodiments of the present application can be practiced without these specific details. In some instances, well-known methods, structures, and technologies have not been shown in detail so as not to obscure the understanding of this specification.
[0183] Similarly, it should be understood that in order to streamline the present application and help understand one or more of the various inventive aspects, in the above description of the exemplary embodiments of the present application, the various features of the present application are sometimes grouped together into a single embodiment, figure, or description thereof. However, the disclosed method should not be construed as reflecting the intention that the claimed present application requires more features than those expressly recited in each claim. Rather, as reflected in the following claims, the inventive aspects lie in less than all the features of the single foregoing disclosed embodiment. Therefore, the claims following the detailed description are hereby expressly incorporated into the detailed description, where each claim itself serves as a separate embodiment of the present application.
[0184] Those skilled in the art can understand that the modules in the devices in the embodiments can be adaptively changed and arranged in one or more devices different from the embodiments. The modules or units or components in the embodiments can be combined into one module or unit or component, and in addition, they can be divided into multiple sub-modules or sub-units or sub-components. Except that at least some of such features and / or processes or units are mutually exclusive, any combination can be used to combine all the features disclosed in this specification (including the accompanying claims, abstract and drawings) and all the processes or units of any method or device so disclosed. Unless otherwise expressly stated, each feature disclosed in this specification (including the accompanying claims, abstract and drawings) can be replaced by an alternative feature that provides the same, equivalent or similar purpose.
[0185] In addition, those skilled in the art can understand that although some of the embodiments described herein include certain features included in other embodiments rather than other features, the combination of the features of different embodiments means that it is within the scope of this application and forms different embodiments. For example, in the following claims, any one of the claimed embodiments can be used in any combination.
[0186] Each component embodiment of the present application can be implemented in hardware, or in software modules running on one or more processors, or in a combination thereof. Those skilled in the art should understand that a microprocessor or a digital signal processor (DSP) can be used in practice to implement some or all of the functions of some or all of the components in the virtual machine creation device according to the embodiments of the present application. The present application can also be implemented as a device or device program (such as a computer program and a computer program product) for executing part or all of the methods described herein. Such a program implementing the present application can be stored on a computer-readable medium, or can be in the form of one or more signals. Such signals can be downloaded from an Internet website, or provided on a carrier signal, or provided in any other form.
[0187] It should be noted that the above embodiments are illustrative of the present application rather than restrictive of the present application, and those skilled in the art can design alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses shall not be construed as limiting the claim. The word "comprising" does not exclude the presence of elements or steps not listed in the claim. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present application can be implemented by means of hardware including several different elements and by means of a suitably programmed computer. In a unit claim listing several devices, several of these devices can be embodied by the same item of hardware. The use of the words first, second, and third, etc. does not denote any order. These words can be interpreted as names.
[0188] As described above, the above are only the preferred specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.
Claims
1. A method for correcting typos, characterized in that, The method includes: Replacing the misspelled words in the text to be corrected with masks to generate a masked text corresponding to the text to be corrected; Invoking a first correction model to correct the masked text to obtain a first correction result; Invoking a second correction model to correct the text to be corrected to obtain a second correction result; Determining the correct words corresponding to the misspelled words in the text to be corrected according to the first correction result and the second correction result.
2. The method according to claim 1, characterized in that, The replacing the misspelled words in the text to be corrected with masks to generate a masked text corresponding to the text to be corrected includes: Extracting the character features and semantic features of each character in the text to be corrected; Determining the positions of the misspelled words in the text to be corrected according to the character features and semantic features of each character; Replacing the misspelled words with the masks to obtain the masked text.
3. The method according to claim 2, characterized in that, The determining the positions of the misspelled words in the text to be corrected according to the character features and semantic features of each character includes: Calculating the error probability of each character according to the character features and semantic features of each character; Determining whether the character is a misspelled word based on the error probability of each character and a set misspelled word threshold; After determining the misspelled word, obtaining the position of the misspelled word in the text to be corrected.
4. The method according to claim 1, characterized in that, Invoking a first correction model to correct the masked text to obtain a first correction result includes: Extracting the semantic features corresponding to the mask based on the first correction model; Calculating a first predicted probability distribution corresponding to the mask according to the semantic features of the mask, where the first predicted probability distribution corresponding to the mask includes the probabilities of each word in a preset word list replacing the mask, and the first predicted probability distribution is the first correction result.
5. The method according to claim 1, characterized in that, Invoking a second correction model to correct the masked text to obtain a second correction result includes: Extracting the semantic features corresponding to the misspelled word based on the second correction model; Calculating a second predicted probability distribution corresponding to the misspelled word according to the semantic features of the misspelled word, where the second predicted probability distribution corresponding to the misspelled word includes the probabilities of each word in a preset word list replacing the misspelled word, and the second predicted probability distribution is the second correction result.
6. The method according to claim 4 or 5, characterized in that, The determining the correct words corresponding to the misspelled words in the text to be corrected according to the first correction result and the second correction result includes: Determining the predicted probability distribution of this position according to the first predicted probability distribution corresponding to the position of the mask and the second predicted probability distribution corresponding to the misspelled word; Determining the target probability of this position according to the predicted probability distribution of this position; Determining the corrected word corresponding to this position based on the preset word list and the target probability.
7. The method according to claim 1, characterized in that, After replacing the misspelled words in the text to be corrected with masks to generate a masked text corresponding to the text to be corrected, the method further includes: Invoking the first correction model to calculate the first probability distribution of each character position in the masked text; Invoking the second correction model to calculate the second probability distribution of each character position in the text to be corrected; Determining the target probability of this position according to the first probability distribution and the second probability distribution of each character position. Compare the target probability corresponding to each character at each position with the character at the corresponding position in the text to be corrected and the character corresponding to the preset vocabulary. If they do not match, update the character at the corresponding position in the text to be corrected to the character corresponding to the preset vocabulary until all character positions are compared.
8. The method according to claim 1, characterized in that, The method further includes: Training an initial model with a training sample set to obtain an error correction model, where the error correction model includes the first correction model and the second correction model. The training sample set includes multiple training sample pairs, and each training sample pair includes a misspelled text and a mask text corresponding to the misspelled text. The positions of the misspelled characters in the mask text are masks; The training process includes: Inputting the training sample pair into the initial model to obtain a training result, where the training result includes an error detection training result and a misspelled character correction training result; Calculating the cross-entropy total loss according to the error detection loss and the misspelled character correction loss. The error detection loss represents the loss between the error detection training result at each character position and the preset error detection annotation result, and the misspelled character correction loss refers to the loss between the misspelled character correction training result at each character position and the preset misspelled character correction annotation result; If the cross-entropy total loss is less than or equal to a preset threshold, determine the initial model as the error correction model; If the cross-entropy total loss is greater than the preset threshold, update the parameters of the initial model to obtain a new initial model, and execute the step of inputting the training sample pair into the initial model to obtain a training result again.
9. A typo correction system, characterized in that, The system includes: An error detection module for replacing the misspelled characters in the text to be corrected with masks to generate a mask text corresponding to the text to be corrected; A first correction module for calling the first correction model to correct the mask text to obtain a first correction result; A second correction module for calling the second correction model to correct the text to be corrected to obtain a second correction result; A result fusion module for determining the correct character corresponding to the misspelled character in the text to be corrected according to the first correction result and the second correction result.
10. An electronic device, comprising: A memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that when the processor runs the computer program, it is executed to implement the method according to any one of claims 1-8.
11. A computer-readable storage medium, characterized in that, Stored thereon are computer-readable instructions, and the computer-readable instructions can be executed by the processor to implement the method according to any one of claims 1-8.