Text error correction method, device, electronic device and storage medium
By combining the feature vector sequence and error correction label sequence of text fragments and historical text fragments, and using position detection networks and character prediction networks for text error correction processing, the problems of low error correction efficiency and low accuracy in the prior art are solved, and more efficient and accurate text error correction effects are achieved.
Patent Information
- Application Number
- CN202111402572.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-23
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2041-11-23
AI Technical Summary
The existing text error correction method determines the position to be corrected and the characters after the error correction based on text fragments, and does not combine historical information, resulting in low error correction efficiency and low accuracy.
By obtaining the text fragments to be processed and their historical text fragments, the feature vector sequence and error correction label sequence of the text fragments are determined, combined with the position detection network and the character prediction network, error correction processing is performed, and finally the error correction text fragment is generated.
Improves the efficiency and accuracy of text error correction, and can detect and correct erroneous characters that rely on historical information.
Smart Images

Figure CN114282546B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technologies, and particularly to the fields of natural language processing and deep learning technologies. In particular, it relates to a text error correction method, apparatus, electronic device, and storage medium. Background Art
[0002] Currently, the main text error correction method is to determine the position to be corrected in a text segment based on the text segment; and then, in combination with the text segment and the position to be corrected, determine the corrected character at the position to be corrected. In the above solution, only the position to be corrected and the corrected character are determined based on the text segment, resulting in low error correction efficiency and low error correction accuracy. Summary of the Invention
[0003] The present disclosure provides a text error correction method, apparatus, device, and storage medium.
[0004] According to one aspect of the present disclosure, a text error correction method is provided, including: obtaining a text segment to be processed and a corresponding historical text segment of the text segment; determining a feature vector sequence of the text segment according to the text segment and the historical text segment; determining an error correction label sequence of the text segment according to the feature vector sequence; where an error correction label in the error correction label sequence indicates whether the corresponding position in the text segment is a position to be corrected; and determining a corrected text segment corresponding to the text segment according to the text segment and the error correction label sequence.
[0005] According to another aspect of the present disclosure, a training method for a text error correction model is provided, including: constructing an initial text error correction model, where the text error correction model includes: a position detection network and a character prediction network connected in sequence; the position detection network includes: a first embedding network layer and a first attention network layer connected in sequence, a second embedding network layer and a first coding network layer connected in sequence, a first weighted network layer and a first fully connected network layer connected in sequence; the first weighted network layer is respectively connected to the first attention network layer and the first coding network layer; obtaining training data, where the training data includes: a sample text segment, a corresponding sample historical text segment, and a sample corrected text segment; using the sample text segment and the sample historical text segment as inputs of the text error correction model, and using the sample corrected text segment as the output of the text error correction model, adjusting the coefficients of the text error correction model to achieve training.
[0006] According to another aspect of the present disclosure, there is provided a text error correction device, including: an acquisition module configured to acquire a text segment to be processed and a corresponding historical text segment of the text segment; a first determination module configured to determine a feature vector sequence of the text segment according to the text segment and the historical text segment; a second determination module configured to determine an error correction label sequence of the text segment according to the feature vector sequence; wherein, the error correction label in the error correction label sequence indicates whether the corresponding position in the text segment is a position to be error-corrected; a third determination module configured to determine a corrected text segment corresponding to the text segment according to the text segment and the error correction label sequence.
[0007] According to another aspect of the present disclosure, there is provided a training device for a text error correction model, including: a construction module configured to construct an initial text error correction model, wherein the text error correction model includes: a position detection network and a character prediction network connected in sequence; the position detection network includes: a first embedding network layer and a first attention network layer connected in sequence, a second embedding network layer and a first coding network layer connected in sequence, a first weighting network layer and a first fully connected network layer connected in sequence; the first weighting network layer is respectively connected to the first attention network layer and the first coding network layer; an acquisition module configured to acquire training data, wherein the training data includes: a sample text segment, a corresponding sample historical text segment, and a sample corrected text segment; a training module configured to adjust the coefficients of the text error correction model with the sample text segment and the sample historical text segment as the input of the text error correction model and the sample corrected text segment as the output of the text error correction model to implement training.
[0008] According to yet another aspect of the present disclosure, there is provided an electronic device, including:
[0009] at least one processor; and
[0010] a memory communicatively connected to the at least one processor; wherein,
[0011] the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the text error correction method or the training method of the text error correction model proposed above in the present disclosure.
[0012] According to still another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the text error correction method or the training method of the text error correction model proposed above in the present disclosure.
[0013] According to another aspect of the present disclosure, there is provided a computer program product including a computer program which, when executed by a processor, implements the steps of the text error correction method or the training method of the text error correction model proposed above in the present disclosure.
[0014] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0016] Figure 1 is a schematic diagram according to the first embodiment of the present disclosure;
[0017] Figure 2 is a schematic diagram according to the second embodiment of the present disclosure;
[0018] Figure 3 is a schematic diagram according to the third embodiment of the present disclosure;
[0019] Figure 4 is a schematic diagram of a position detection network;
[0020] Figure 5 is a schematic diagram of a character prediction network;
[0021] Figure 6 is a schematic diagram according to the fourth embodiment of the present disclosure;
[0022] Figure 7 is a schematic diagram according to the fifth embodiment of the present disclosure;
[0023] Figure 8 shows a schematic block diagram of an exemplary electronic device 800 that can be used to implement the embodiments of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0024] The following makes an explanation of the exemplary embodiments of the present disclosure with reference to the drawings. Various details of the embodiments of the present disclosure are included to help understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described here without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, the description of well-known functions and structures is omitted below.
[0025] Currently, the text error correction method mainly includes: determining the position to be corrected in the text segment based on the text segment; and then combining the text segment and the position to be corrected to determine the corrected character at the position to be corrected. In the above solution, historical information is not combined, and errors that depend on historical information cannot be detected and corrected. Only the position to be corrected and the corrected character are determined based on the text segment, resulting in low error correction efficiency and low error correction accuracy.
[0026] In view of the above problems, the present disclosure proposes a text error correction method, apparatus, electronic device, and storage medium.
[0027] Figure 1 FIG. is a schematic diagram according to the first embodiment of the present disclosure. It should be noted that the text error correction method of the embodiments of the present disclosure can be applied to a text error correction device, which can be configured in an electronic device so that the electronic device can perform the text error correction function.
[0028] Among them, the electronic device can be any device with computing capabilities, such as a personal computer (PC for short), a mobile terminal, a server, etc. The mobile terminal can be, for example, a vehicle-mounted device, a mobile phone, a tablet computer, a personal digital assistant, a wearable device, etc., which are hardware devices with various operating systems, touch screens, and / or display screens.
[0029] As Figure 1 shown, the text error correction method may include the following steps:
[0030] Step 101, obtain the text segment to be processed and the historical text segment corresponding to the text segment.
[0031] In the embodiments of the present disclosure, the process of the text error correction device executing step 101 may be, for example, obtaining the text segment to be processed and the text to which the text segment belongs; obtaining at least one segment in the text before the text segment; and determining the at least one segment as the historical text segment.
[0032] In the embodiments of the present disclosure, the historical text segment may be the historical text segment before the text segment to be processed, where the historical text segment may be one or more. The historical text segment has a certain relevance to the text segment. Combining the historical text segment for error correction processing of the text segment can improve the error correction efficiency of the text segment.
[0033] In the embodiments of the present disclosure, for example, the text segment may be "She once wanted to recover this batch of goods, but...", and the corresponding historical text segment may be "Mr. Wang often...". Through "Mr. Wang" in the historical text segment, error characters such as "She" in the text segment can be detected, and the position corresponding to the error character is the position to be corrected.
[0034] Step 102: Determine the feature vector sequence of the text segment according to the text segment and the historical text segment.
[0035] In the embodiment of the present disclosure, the process of the text error correction device executing Step 102 may be, for example, splicing the historical text segment and the text segment to obtain a spliced segment; determining the feature vector sequence of the spliced segment; and using a preset number of vectors at the back in the feature vector sequence of the spliced segment as the feature vector sequence of the text segment. The preset number may be the number of characters or words in the text segment, which is the same as the dimension of the text vector sequence of the text segment.
[0036] Wherein, when the preset number is the number of characters in the text segment, the text vector sequence of the text segment may be a vector sequence obtained by vectorizing each character in the text segment.
[0037] Step 103: Determine the error correction label sequence of the text segment according to the feature vector sequence; wherein, the error correction label in the error correction label sequence indicates whether the corresponding position in the text segment is a position to be error-corrected.
[0038] In the embodiment of the present disclosure, the error correction label may be, for example, F or T. Wherein F indicates that the corresponding position in the text segment is a position to be error-corrected; T indicates that the corresponding position in the text segment is not a position to be error-corrected. For another example, the error correction label may be 0 or 1. Wherein, 0 indicates that the corresponding position in the text segment is a position to be error-corrected; 1 indicates that the corresponding position in the text segment is not a position to be error-corrected.
[0039] In one example, when the number of feature vectors in the feature vector sequence is the number of characters in the text segment, each position in the error correction label sequence may correspond to a character, for example.
[0040] In another example, when the number of feature vectors in the feature vector sequence is the number of words in the text segment, each position in the error correction label sequence may correspond to a word, for example.
[0041] Step 104: Determine the error-corrected text segment corresponding to the text segment according to the text segment and the error correction label sequence.
[0042] In an embodiment of the present disclosure, in one example, taking that each position of the error correction label sequence corresponds to a character as an example, the process for the text error correction device to execute step 104 may be, for example, performing a marking process on the text segment for the positions to be error corrected according to the error correction label sequence to obtain the marked text segment; determining a sequence of feature vectors of the marked text segment according to the marked text segment and the historical text segment; determining the error corrected character corresponding to the position to be error corrected in the marked text segment according to the sequence of feature vectors of the marked text segment; and generating the error corrected text segment corresponding to the text segment according to the marked text segment and the error corrected character corresponding to the position to be error corrected.
[0043] In another example, taking that each position of the error correction label sequence corresponds to a character as an example, the process for the text error correction device to execute step 104 may be, for example, performing a marking process on the text segment for the positions to be error corrected according to the error correction label sequence to obtain the marked text segment; determining a sequence of feature vectors of the marked text segment according to the marked text segment; determining the error corrected character corresponding to the position to be error corrected in the marked text segment according to the sequence of feature vectors of the marked text segment; and generating the error corrected text segment corresponding to the text segment according to the marked text segment and the error corrected character corresponding to the position to be error corrected.
[0044] In an embodiment of the present application, by combining the historical text segment and determining the sequence of feature vectors of the marked text segment according to the marked text segment, the error corrected character corresponding to the position to be error corrected in the marked text segment can be accurately determined, improving the accuracy of text segment error correction.
[0045] The text error correction method of the embodiment of the present disclosure obtains the text segment to be processed and the corresponding historical text segment of the text segment; determines the sequence of feature vectors of the text segment according to the text segment and the historical text segment; determines the error correction label sequence of the text segment according to the sequence of feature vectors; wherein, the error correction label in the error correction label sequence indicates whether the corresponding position in the text segment is a position to be error corrected; determines the error corrected text segment corresponding to the text segment according to the text segment and the error correction label sequence, and refers to the historical text segment to detect and correct the error characters in the text segment that depend on the historical text segment, improving the error correction efficiency and the error correction accuracy.
[0046] In order to accurately obtain the sequence of feature vectors of the text segment according to the text segment and the corresponding historical text segment, as Figure 2 shown, Figure 2 is a schematic diagram according to the second embodiment of the present disclosure. In the embodiment of the present disclosure, the historical sequence of feature vectors of the text segment can be determined according to the sequence of text vectors of the historical text segment and the initial sequence of feature vectors, and then the sequence of feature vectors of the text segment can be determined. Figure 2 The embodiment shown may include the following steps:
[0047] Step 201: Obtain the text segment to be processed and the corresponding historical text segment.
[0048] Step 202: Determine the initial feature vector sequence of the text segment and the text vector sequence of the historical text segment.
[0049] In an embodiment of the present disclosure, in one example, the process by which the text error correction device determines the initial feature vector sequence of the text segment may be, for example, to vectorize each character in the text segment to generate a text vector sequence of the text segment; perform feature extraction processing based on the text vector sequence of the text segment to obtain the initial feature vector sequence of the text segment. The process by which the text error correction device determines the text vector sequence of the historical text segment may be, for example, to vectorize each character in the historical text segment to generate a text vector sequence of the historical text segment.
[0050] In another example, the process by which the text error correction device determines the initial feature vector sequence of the text segment may be, for example, to vectorize each word in the text segment to generate a text vector sequence of the text segment; perform feature extraction processing based on the text vector sequence of the text segment to obtain the initial feature vector sequence of the text segment. The process by which the text error correction device determines the text vector sequence of the historical text segment may be, for example, to vectorize each word in the historical text segment to generate a text vector sequence of the historical text segment.
[0051] Among them, the process of vectorizing each character in the text segment is, for example, to input each character in the text segment into the second embedding network layer to obtain the text vector sequence of the text segment output by the second embedding network layer. The process of performing feature extraction processing on the text vector sequence of the text segment is, for example, to input the text vector sequence of the text segment into the first encoding network layer for feature extraction processing to obtain the initial feature vector sequence of the text segment output by the first encoding network layer. Among them, the process of vectorizing each character in the historical text segment is, for example, to input each character in the historical text segment into the first embedding network layer to obtain the text vector sequence of the historical text segment output by the first embedding network layer.
[0052] Among them, the process of vectorizing each word in the text segment is, for example, inputting each word in the text segment into the second embedding network layer to obtain the text vector sequence of the text segment output by the second embedding network layer. The process of feature extraction processing on the text vector sequence of the text segment is, for example, inputting the text vector sequence of the text segment into the first encoding network layer for feature extraction processing to obtain the initial feature vector sequence of the text segment output by the first encoding network layer. Among them, the process of vectorizing each word in the historical text segment is, for example, inputting each word in the historical text segment into the first embedding network layer to obtain the text vector sequence of the historical text segment output by the first embedding network layer.
[0053] In the embodiments of the present disclosure, the length of the text vector sequence is determined according to the number of characters in the historical text segment, and the length of the initial feature vector sequence is determined according to the number of characters in the text segment. In the case of vectorizing the characters in the text segment and the characters in the historical text segment, if the number of characters in the text segment is the same as the number of characters in the historical text segment, the number of vectors in the initial feature vector sequence is the same as the number of vectors in the text vector sequence of the historical text segment; if the number of characters in the text segment is different from the number of characters in the historical text segment, the number of vectors in the initial feature vector sequence is different from the number of vectors in the text vector sequence of the historical text segment.
[0054] Step 203: Determine the historical feature vector sequence of the text segment according to the text vector sequence and the initial feature vector sequence.
[0055] In the embodiments of the present disclosure, the process for the text error correction device to execute step 203 may be, for example, determining the weight sequence of the initial feature vector sequence according to the text vector sequence and the initial feature vector sequence; and determining the historical feature vector sequence according to the initial feature vector sequence and the weight sequence.
[0056] In the embodiments of the present disclosure, the initial feature vector sequence and the weight sequence have the same length. The feature vector at the j-th position in the initial feature vector sequence is multiplied by the weight at the j-th position in the weight sequence to obtain the feature vector at the j-th position in the historical feature vector sequence. Wherein, the value of j ranges from 1 to N, and N is the number of feature vectors in the initial feature vector sequence.
[0057] In the embodiments of the present disclosure, according to the text vector sequence and the initial feature vector sequence, the weight sequence of the initial feature vector sequence is determined; according to the initial feature vector sequence and the weight sequence, the historical feature vector sequence is determined, so as to combine the text vector sequence of the historical text segment to determine the influence on the initial feature vector sequence, and then determine the historical feature vector sequence. Combining the historical features for error correction processing can further improve the efficiency of text error correction.
[0058] Step 204: Determine the feature vector sequence of the text segment according to the initial feature vector sequence and the historical feature vector sequence.
[0059] In the embodiments of the present disclosure, the process of the text error correction device executing Step 204 may be, for example, determining the first weight corresponding to the historical text segment and the second weight corresponding to the text segment; performing weighted summation processing on the initial feature vector sequence and the historical feature vector sequence according to the first weight and the second weight to obtain the feature vector sequence.
[0060] In the embodiments of the present disclosure, the first weight and the second weight may be determined according to the importance of the historical text segment and the text segment. For example, if the historical text segment is an adjacent segment of the text segment and has a high correlation, the first weight may be increased; if the distance between the historical text segment and the text segment is large or the correlation is low, the first weight may be decreased. By adjusting the first weight and the second weight, the influence of the historical text segment on the text segment can be adjusted, further improving the efficiency and accuracy of text error correction.
[0061] Step 205: Determine the error correction label sequence of the text segment according to the feature vector sequence; wherein, the error correction label in the error correction label sequence indicates whether the corresponding position in the text segment is a position to be error-corrected.
[0062] Step 206: Determine the error-corrected text segment corresponding to the text segment according to the text segment and the error correction label sequence.
[0063] It should be noted that for the detailed content of Step 201, Step 205, and Step 206, reference may be made to Figure 1 Steps 101, 103, and 104 in the illustrated embodiments, and details are not described herein again.
[0064] The text error correction method of the embodiments of the present disclosure obtains the text segment to be processed and the corresponding historical text segment of the text segment; determines the initial feature vector sequence of the text segment and the text vector sequence of the historical text segment; determines the historical feature vector sequence of the text segment according to the text vector sequence and the initial feature vector sequence; determines the feature vector sequence of the text segment according to the initial feature vector sequence and the historical feature vector sequence. Determine the error correction label sequence of the text segment according to the feature vector sequence; wherein, the error correction label in the error correction label sequence indicates whether the corresponding position in the text segment is a position to be error-corrected; determine the error-corrected text segment corresponding to the text segment according to the text segment and the error correction label sequence, refer to the historical text segment, detect and correct the error characters in the text segment that depend on the historical text segment, improve the efficiency of error correction, and improve the accuracy of error correction.
[0065] Figure 3 is a schematic diagram according to the third embodiment of the present disclosure. As Figure 3 shown, the training method of the text error correction model includes:
[0066] Step 301, construct an initial text error correction model, where the text error correction model includes: a position detection network and a character prediction network connected in sequence; the position detection network includes: a first embedding network layer and a first attention network layer connected in sequence, a second embedding network layer and a first encoding network layer connected in sequence, a first weighted network layer and a first fully connected network layer connected in sequence; the first weighted network layer is respectively connected to the first attention network layer and the first encoding network layer.
[0067] In the embodiment of the present disclosure, the character prediction network includes: a third embedding network layer and a second attention network layer connected in sequence, a fourth embedding network layer and a second encoding network layer connected in sequence, a second weighted network layer and a second fully connected network layer connected in sequence; the second weighted network layer is respectively connected to the second attention network layer and the second encoding network layer.
[0068] In the embodiment of the present disclosure, the historical text segment can be input into the first embedding network layer to obtain the text vector sequence of the historical text segment output by the first embedding network layer; the text segment can be input into the second embedding network layer to obtain the text vector sequence of the text segment output by the second embedding network layer, and the output of the second embedding network layer is used as the input of the first encoding network layer to obtain the initial feature vector sequence of the text segment; the output of the first embedding network layer and the output of the first encoding network layer are used as the input of the first attention network layer to obtain the historical feature vector sequence; the output of the first encoding network layer and the output of the first attention network layer are used as the input of the first weighted network layer to perform weighted summation processing on the initial feature vector sequence and the historical feature vector sequence to obtain the feature vector sequence, which is input into the first fully connected network layer, and after being processed by the first fully connected network layer, the error correction label sequence is determined.
[0069] In an embodiment of the present disclosure, a historical text segment can be input into a third embedding network layer to obtain a text vector sequence of the historical text segment output by the third embedding network layer; the text segment can be processed for marking positions to be corrected according to an error correction label sequence to obtain a marked text segment; the marked text segment is input into a fourth embedding network layer to obtain a text vector sequence of the text segment output by the fourth embedding network layer, and the output of the fourth embedding network layer is used as the input of a second encoding network layer to obtain an initial feature vector sequence of the text segment; the output of the third embedding network layer and the output of the second encoding network layer are used as the input of a second attention network layer to obtain a historical feature vector sequence; the output of the second encoding network layer and the output of the second attention network layer are used as the input of a second weighting network layer to perform weighted summation processing on the initial feature vector sequence and the historical feature vector sequence to obtain a feature vector sequence, which is input into a second fully connected network layer, and after being processed by the second fully connected network layer, the corrected character at the position to be corrected is determined, and then the corrected text segment of the text segment is generated.
[0070] Step 302, obtain training data, where the training data includes: sample text segments, corresponding sample historical text segments, and sample corrected text segments.
[0071] Step 303, use the sample text segment and the sample historical text segment as the input of the text error correction model, and use the sample corrected text segment as the output of the text error correction model to adjust the coefficients of the text error correction model to achieve training.
[0072] In an embodiment of the present disclosure, the coefficients of the first embedding network layer, the second embedding network layer, the first encoding network layer, the third embedding network layer, the fourth embedding network layer, and the second encoding network layer in the initial text error correction model are obtained by migrating the coefficients in a pre-trained semantic representation model. During the process of training the text error correction model, the coefficients of the above-mentioned each network layer can be fixed, and the coefficients of other network layers can be adjusted; or, the gradient of the coefficient adjustment of the above-mentioned each network layer can be reduced, and the gradient of the coefficient adjustment of other network layers can be increased, so as to accelerate the training of the text error correction model and ensure the accuracy of the trained text error correction model.
[0073] In summary, by constructing an initial text error correction model, where the text error correction model includes: a position detection network and a character prediction network connected in sequence; the position detection network includes: a first embedding network layer and a first attention network layer connected in sequence, a second embedding network layer and a first encoding network layer connected in sequence, a first weighting network layer and a first fully connected network layer connected in sequence; the first weighting network layer is respectively connected to the first attention network layer and the first encoding network layer; obtaining training data, where the training data includes: sample text segments, corresponding sample historical text segments, and sample corrected text segments; using the sample text segments and the sample historical text segments as the input of the text error correction model, and using the sample corrected text segments as the output of the text error correction model, adjusting the coefficients of the text error correction model to achieve training, referring to the historical text segments, detecting and correcting the error characters in the text segments that depend on the historical text segments, improving the efficiency of error correction and the accuracy of error correction.
[0074] To illustrate the above embodiments more clearly, examples are given below.
[0075] For example, the text error correction model includes: a position detection network (detection module) and a character prediction network (error correction module), as Figure 4 shown, is a schematic diagram of the position detection network, as Figure 5 shown, is a schematic diagram of the character prediction network. In Figure 4 it, first, the character sequence of the text segment obtains the initial feature vector sequence h i through the second embedding network layer and the first network encoding layer. Before the text error correction model predicts the error correction label, first, the character sequence of the historical text segment is processed through the first embedding network layer to extract the text vector sequence of the historical text segment. Based on the first attention network layer, the initial feature vector sequence, and the text vector sequence, the historical feature vector sequence is determined. The specific calculation process is shown in formula (1).
[0076] d i = Atten(h i , Emb(c), Emb(c)) (1)
[0077] where c = {c 1 , c 2 ,..., c m}, represents the character sequence of length m of the historical text segment, and Atten represents the processing of the first attention network layer.
[0078] After determining the historical feature vector sequence, the historical feature vector sequence is fused with the initial feature vector sequence through the first weighted network layer to obtain a feature vector sequence. Then, the first fully connected network layer of the position detection network is used to predict the feature vector sequence to obtain an error correction label sequence, and the calculation process is shown in formulas (2), (3), and (4).
[0079] λ i =σ(W h h i +W d d i ) (2)
[0080]
[0081]
[0082] Among them, λ i represents the first weight of the initial feature vector sequence; 1 - λ i represents the second weight of the historical feature vector sequence; W h and W d are the coefficients of the first attention network layer; σ is the sigmoid function; represents the feature vector sequence; y i represents the error correction label sequence.
[0083] In Figure 5 , first, the text segment is marked with the positions to be error-corrected according to the error correction label sequence to obtain the marked text segment; the character sequence of the marked text segment is passed through the fourth embedding network layer and the second network encoding layer to obtain the initial feature vector sequence. Before the error-corrected character corresponding to the position to be error-corrected in the text error correction model, the character sequence of the historical text segment is first processed through the third embedding network layer to extract the text vector sequence of the historical text segment. Based on the second attention network layer, the initial feature vector sequence, and the text vector sequence, the historical feature vector sequence is determined; the historical feature vector sequence and the initial feature vector sequence are fused through the second weighted network layer to determine the feature vector sequence; then, the second fully connected network layer in the character prediction network is used to predict the feature vector sequence to obtain the error-corrected character corresponding to the predicted position to be error-corrected, and then the error-corrected text segment is generated.
[0084] To implement the above embodiments, the present disclosure also proposes a text error correction device.
[0085] As Figure 6 shown, Figure 6 According to the schematic diagram of the fourth embodiment of the present disclosure. The text error correction device 600 includes: an acquisition module 610, a first determination module 620, a second determination module 630, and a third determination module 640.
[0086] Among them, an acquisition module 610 is configured to acquire a text segment to be processed and a corresponding historical text segment; a first determination module 620 is configured to determine a feature vector sequence of the text segment according to the text segment and the historical text segment; a second determination module 630 is configured to determine an error correction label sequence of the text segment according to the feature vector sequence; wherein, an error correction label in the error correction label sequence indicates whether a corresponding position in the text segment is a position to be error-corrected; a third determination module 640 is configured to determine an error-corrected text segment corresponding to the text segment according to the text segment and the error correction label sequence.
[0087] As a possible implementation manner of an embodiment of the present disclosure, the first determination module 620 includes: a first determination unit, a second determination unit, and a third determination unit; wherein, the first determination unit is configured to determine an initial feature vector sequence of the text segment and a text vector sequence of the historical text segment; the second determination unit is configured to determine a historical feature vector sequence of the text segment according to the text vector sequence and the initial feature vector sequence; the third determination unit is configured to determine the feature vector sequence of the text segment according to the initial feature vector sequence and the historical feature vector sequence.
[0088] As a possible implementation manner of an embodiment of the present disclosure, the second determination unit is specifically configured to determine a weight sequence of the initial feature vector sequence according to the text vector sequence and the initial feature vector sequence; and determine the historical feature vector sequence according to the initial feature vector sequence and the weight sequence.
[0089] As a possible implementation manner of an embodiment of the present disclosure, the third determination unit is specifically configured to determine a first weight corresponding to the historical text segment and a second weight corresponding to the text segment; and perform weighted summation processing on the initial feature vector sequence and the historical feature vector sequence according to the first weight and the second weight to obtain the feature vector sequence.
[0090] As a possible implementation manner of an embodiment of the present disclosure, the third determination module 640 is specifically configured to perform a position marking process for positions to be error-corrected on the text segment according to the error correction label sequence to obtain a marked text segment; determine a feature vector sequence of the marked text segment according to the marked text segment and the historical text segment; determine an error-corrected character corresponding to a position to be error-corrected in the marked text segment according to the feature vector sequence of the marked text segment; and generate an error-corrected text segment corresponding to the text segment according to the marked text segment and the error-corrected character corresponding to the position to be error-corrected.
[0091] As a possible implementation manner of the embodiments of the present disclosure, the obtaining module 610 is specifically configured to obtain a text segment to be processed and the text to which the text segment belongs; obtain at least one segment in the text that is before the text segment; and determine the at least one segment as the historical text segment.
[0092] The text error correction device of the embodiments of the present disclosure obtains a text segment to be processed and a corresponding historical text segment of the text segment; determines a feature vector sequence of the text segment according to the text segment and the historical text segment; determines an error correction label sequence of the text segment according to the feature vector sequence; wherein, the error correction label in the error correction label sequence indicates whether the corresponding position in the text segment is a position to be error-corrected; determines a corrected text segment corresponding to the text segment according to the text segment and the error correction label sequence, and refers to the historical text segment to detect and correct the error characters in the text segment that depend on the historical text segment, improving the efficiency of error correction and the accuracy of error correction.
[0093] To implement the above embodiments, the present disclosure also proposes a training device for a text error correction model.
[0094] As Figure 7 shown, Figure 7 According to the schematic diagram of the fifth embodiment of the present disclosure. The training device 700 for the text error correction model includes: a construction module 710, an obtaining module 720, and a training module 730.
[0095] Among them, the construction module 710 is configured to construct an initial text error correction model, where the text error correction model includes: a position detection network and a character prediction network connected in sequence; the position detection network includes: a first embedding network layer and a first attention network layer connected in sequence, a second embedding network layer and a first coding network layer connected in sequence, a first weighted network layer and a first fully connected network layer connected in sequence; the first weighted network layer is respectively connected to the first attention network layer and the first coding network layer;
[0096] The obtaining module 720 is configured to obtain training data, where the training data includes: a sample text segment, a corresponding sample historical text segment, and a sample corrected text segment;
[0097] The training module 730 is configured to use the sample text segment and the sample historical text segment as the input of the text error correction model, and use the sample corrected text segment as the output of the text error correction model to adjust the coefficients of the text error correction model to achieve training.
[0098] As a possible implementation manner of the embodiments of the present disclosure, the character prediction network includes: a third embedding network layer and a second attention network layer connected in sequence, a fourth embedding network layer and a second encoding network layer connected in sequence, and a second weighting network layer and a second fully connected network layer connected in sequence; the second weighting network layer is respectively connected to the second attention network layer and the second encoding network layer.
[0099] The training device of the text error correction model of the embodiments of the present disclosure constructs an initial text error correction model, where the text error correction model includes: a position detection network and a character prediction network connected in sequence; the position detection network includes: a first embedding network layer and a first attention network layer connected in sequence, a second embedding network layer and a first encoding network layer connected in sequence, and a first weighting network layer and a first fully connected network layer connected in sequence; the first weighting network layer is respectively connected to the first attention network layer and the first encoding network layer; obtains training data, where the training data includes: a sample text segment, a corresponding sample historical text segment, and a sample corrected text segment; uses the sample text segment and the sample historical text segment as the input of the text error correction model, and uses the sample corrected text segment as the output of the text error correction model to adjust the coefficients of the text error correction model to implement training, refers to the historical text segment, detects and corrects the error characters in the text segment that depend on the historical text segment, improves the efficiency of error correction, and improves the accuracy of error correction.
[0100] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information and other processing are all carried out on the premise of obtaining the user's consent, and all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0101] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0102] Figure 8 FIG. shows a schematic block diagram of an exemplary electronic device 800 that can be used to implement the embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processing, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0103] As Figure 8As shown, device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0104] Multiple components in device 800 are connected to the I / O interface 805, including: an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a magnetic disk, an optical disc, etc.; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0105] The computing unit 801 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 801 executes the various methods and processes described above, such as a text correction method or a training method for a text correction model. For example, in some embodiments, the text correction method or the training method for the text correction model can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the text correction method or the training method for the text correction model described above can be executed. Alternatively, in other embodiments, the computing unit 801 can be configured to execute the text correction method or the training method for the text correction model in any other appropriate manner (e.g., by means of firmware).
[0106] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0107] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program code can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0108] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media would include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0109] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).
[0110] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0111] A computer system can include a client and a server. The client and the server are generally far from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, or a server of a distributed system, or a server incorporating a blockchain.
[0112] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution disclosed in this disclosure can be achieved, and no limitation is imposed herein.
[0113] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.
Claims
1. A text error correction method, comprising: obtaining a text segment to be processed and a corresponding historical text segment of the text segment; determining an initial feature vector sequence of the text segment and a text vector sequence of the historical text segment; determining a historical feature vector sequence of the text segment according to the text vector sequence and the initial feature vector sequence; determining the feature vector sequence of the text segment according to the initial feature vector sequence and the historical feature vector sequence; determining an error correction label sequence of the text segment according to the feature vector sequence; wherein, the error correction label in the error correction label sequence indicates whether the corresponding position in the text segment is a position to be error corrected; determining a corrected text segment corresponding to the text segment according to the text segment and the error correction label sequence.
2. The method according to claim 1, wherein, the determining the historical feature vector sequence of the text segment according to the text vector sequence and the initial feature vector sequence includes: determining a weight sequence of the initial feature vector sequence according to the text vector sequence and the initial feature vector sequence; determining the historical feature vector sequence according to the initial feature vector sequence and the weight sequence.
3. The method according to claim 1, wherein, the determining the feature vector sequence of the text segment according to the initial feature vector sequence and the historical feature vector sequence includes: determining a first weight corresponding to the historical text segment and a second weight corresponding to the text segment; performing a weighted summation process on the initial feature vector sequence and the historical feature vector sequence according to the first weight and the second weight to obtain the feature vector sequence.
4. The method according to claim 1, wherein, the determining the corrected text corresponding to the text segment according to the text segment and the error correction label sequence includes: performing a marking process on the text segment at the positions to be error corrected according to the error correction label sequence to obtain a marked text segment; determining a feature vector sequence of the marked text segment according to the marked text segment and the historical text segment; determining a corrected character corresponding to the position to be error corrected in the marked text segment according to the feature vector sequence of the marked text segment; generating a corrected text segment corresponding to the text segment according to the marked text segment and the corrected character corresponding to the position to be error corrected.
5. The method according to claim 1, wherein, the obtaining the text segment to be processed and the corresponding historical text segment includes: obtaining a text segment to be processed and the text to which the text segment belongs; obtaining at least one segment in the text before the text segment; determining the at least one segment as the historical text segment.
6. A text error correction device, comprising: an obtaining module for obtaining a text segment to be processed and a corresponding historical text segment of the text segment; The first determination module is configured to determine a feature vector sequence of the text segment according to the text segment and the historical text segment; The second determination module is configured to determine an error correction label sequence of the text segment according to the feature vector sequence; wherein, the error correction label in the error correction label sequence indicates whether the corresponding position in the text segment is a position to be error-corrected; The third determination module is configured to determine a corrected text segment corresponding to the text segment according to the text segment and the error correction label sequence; Wherein, the first determination module includes: a first determination unit, a second determination unit, and a third determination unit; The first determination unit is configured to determine an initial feature vector sequence of the text segment and a text vector sequence of the historical text segment; The second determination unit is configured to determine a historical feature vector sequence of the text segment according to the text vector sequence and the initial feature vector sequence; The third determination unit is configured to determine the feature vector sequence of the text segment according to the initial feature vector sequence and the historical feature vector sequence.
7. The apparatus according to claim 6, wherein, The second determination unit is specifically configured to, determine a weight sequence of the initial feature vector sequence according to the text vector sequence and the initial feature vector sequence; determine the historical feature vector sequence according to the initial feature vector sequence and the weight sequence.
8. The apparatus according to claim 6, wherein, The third determination unit is specifically configured to, determine a first weight corresponding to the historical text segment and a second weight corresponding to the text segment; perform a weighted summation process on the initial feature vector sequence and the historical feature vector sequence according to the first weight and the second weight to obtain the feature vector sequence.
9. The apparatus according to claim 6, wherein, The third determination module is specifically configured to, perform a position marking process for positions to be error-corrected on the text segment according to the error correction label sequence to obtain a marked text segment; determine a feature vector sequence of the marked text segment according to the marked text segment and the historical text segment; determine an error-corrected character corresponding to a position to be error-corrected in the marked text segment according to the feature vector sequence of the marked text segment; generate a corrected text segment corresponding to the text segment according to the marked text segment and the error-corrected character corresponding to the position to be error-corrected.
10. The apparatus according to claim 6, wherein, The acquisition module is specifically configured to, acquire a text segment to be processed and the text to which the text segment belongs; acquire at least one segment in the text that is before the text segment; determine the at least one segment as the historical text segment.
11. An electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1-5.
12. A non-transitory computer-readable storage medium storing computer instructions, wherein, the computer instructions are used to cause the computer to execute the method according to any one of claims 1-5.
13. A computer program product comprising a computer program which, when executed by a processor, implements the steps of the method according to any one of claims 1-5.
Citation Information
Patent Citations
Text error correction model training and text error correction method and device
CN113255332A
Text error correction method and device, terminal equipment and computer storage medium
CN113297833A