Text error correction method and device based on GRU-Transform, electronic equipment and storage medium

By adopting a GRU-Transformer-based method in the field of text error correction, a multi-dataset error correction model is constructed, which solves the problems of low accuracy and poor adaptability of text error correction in the prior art, and achieves a more efficient and comprehensive and accurate text error correction effect.

CN120106053APending Publication Date: 2025-06-06WUHAN FIBERHOME PUTIAN INFORMATION TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510062280.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The prior art has problems such as low accuracy, poor adaptability and complex calculations in the field of text error correction, especially in open field text error correction, which is difficult to effectively deal with complex semantics and contextual relationships.

Method used

The text error correction method based on GRU-Transformer is adopted to obtain multiple data sets (spelling, grammar, punctuation and similar word error data sets), a text error correction data set is constructed, and a text error correction model is trained. The model includes feature extraction submodel, feature fusion submodel, GRU-Transformer encoding-decoding submodel, and loss calculation submodel, which can extract features from multiple dimensions and process context information.

Benefits of technology

It improves the accuracy and comprehensiveness of text error correction, can effectively deal with various text error types, identify and correct various types of text errors such as spelling, grammar, word use and punctuation, and enhances the error correction ability of open-field texts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120106053A_ABST
    Figure CN120106053A_ABST
Patent Text Reader

Abstract

The invention discloses a text error correction method and device based on GRU-Transform, electronic equipment and a storage medium, and belongs to the technical field of text processing. The method comprises the following steps: acquiring a plurality of data sets, wherein the data sets comprise a spelling error data set, a grammar error data set, a punctuation error data set and a similar word error data set; constructing a text error correction data set based on the plurality of data sets; a text error correction model is trained based on the text error correction data set, the text error correction model comprises a feature extraction sub-model, a feature fusion sub-model, a coding-decoding sub-model and a loss calculation sub-model, and the coding-decoding sub-model is a GRU-Transform model; and inputting the to-be-corrected text data into the trained text error correction model to obtain an error-corrected text output result, thereby improving the accuracy and applicability of text error correction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of text processing technology, and in particular, relates to a text error correction method, device, electronic device and storage medium based on GRU-Transformer. Background Art

[0002] With the continuous advancement of information technology, the speed of text data generation and dissemination is accelerating, resulting in errors in the text, especially spelling, grammar, word choice and punctuation errors, which frequently affect the quality and readability of the text. Especially in the open field, text error correction faces more severe challenges because this type of text has no fixed domain norms and the error types are complex and diverse.

[0003] Existing text error correction technologies can be divided into rule-based methods, statistical methods, and deep learning-based methods. Rule-based methods manually design error detection rules and are suitable for simple spelling and grammatical errors. However, when dealing with complex error types and open domain texts, the coverage and flexibility of the rules are relatively limited. Statistical methods rely on a large amount of training data to count the probability of error occurrence, but have certain deficiencies in dealing with complex semantics and contextual relationships. Deep learning-based methods often rely on a single type of data set and ignore the fusion of multi-feature information when dealing with text errors in different fields and of different types.

[0004] The application of related technologies in the field of text error correction still has certain limitations. In open-domain text error correction, there are problems such as low error correction accuracy, poor adaptability, and complex calculations. Summary of the invention

[0005] The present application aims to solve at least one of the technical problems existing in the prior art. To this end, the present application proposes a text error correction method, device, electronic device and storage medium based on GRU-Transformer, which improves the accuracy and applicability of text error correction.

[0006] In a first aspect, the present application provides a text error correction method based on GRU-Transformer, the method comprising:

[0007] Acquire multiple data sets, the data sets including a spelling error data set, a grammatical error data set, a punctuation error data set, and a similar word error data set;

[0008] Based on the multiple data sets, construct a text error correction data set;

[0009] Training a text error correction model based on the text error correction data set, wherein the text error correction model includes a feature extraction sub-model, a feature fusion sub-model, an encoding-decoding sub-model and a loss calculation sub-model, and the encoding-decoding sub-model is a GRU-Transformer model;

[0010] The text data to be corrected is input into the trained text correction model to obtain the corrected text output result.

[0011] According to one embodiment of the present application, the feature extraction sub-model includes:

[0012] Pinyin feature extraction module, used to convert Chinese characters into pinyin to obtain pinyin features;

[0013] The character feature extraction module is used to split the sentence into individual characters and obtain character features;

[0014] The word feature extraction module is used to split the sentence into words and obtain word features;

[0015] The stroke feature extraction module is used to obtain the stroke sequence of Chinese characters and obtain stroke features.

[0016] According to one embodiment of the present application, the feature fusion sub-model includes:

[0017] Vector embedding module, used to convert features into feature vectors through the embedding layer;

[0018] The weighted fusion module is used to perform weighted fusion on the feature vectors to obtain a fusion vector.

[0019] According to one embodiment of the present application, the encoding-decoding sub-model includes:

[0020] An encoder module, used for processing the fused vector to obtain context information of the fused vector;

[0021] A decoder module, used to capture the dependency between various positions in the context information and obtain the text output result after error correction;

[0022] Among them, the encoder module is a GRU encoder, and the decoder module is a Transformer decoder.

[0023] According to one embodiment of the present application, the loss calculation sub-model includes:

[0024] The cross entropy loss module is used to calculate the difference between the predicted probability and the true label at each moment;

[0025] A back propagation module, used to calculate the gradient of the cross entropy loss module to the text error correction model parameters;

[0026] The parameter optimization module is used to optimize the parameters of the text error correction model.

[0027] According to an embodiment of the present application, the step of inputting the text data to be corrected into the trained text correction model to obtain a text output result after correction includes:

[0028] Inputting the text data to be corrected into the feature extraction sub-model to obtain phonetic features, character features, word features and stroke features;

[0029] Inputting the pinyin feature, the character feature, the word feature and the stroke feature into the feature fusion sub-model to obtain the fusion vector;

[0030] The fusion vector is input into the encoding-decoding sub-model to obtain the error-corrected text output result.

[0031] According to one embodiment of the present application, the training of a text error correction model based on the text error correction dataset includes:

[0032] Build a preset text error correction model;

[0033] Based on the loss calculation sub-model, the preset text error correction model is trained according to the text error correction data set to obtain the text error correction model.

[0034] In a second aspect, the present application provides a text error correction device based on GRU-Transformer, the device comprising:

[0035] An acquisition module, used to acquire multiple data sets, wherein the data sets include a spelling error data set, a grammatical error data set, a punctuation error data set, and a similar word error data set;

[0036] A construction module, used to construct a text error correction dataset based on the multiple datasets;

[0037] A training module, used for training a text error correction model based on the text error correction data set, wherein the text error correction model includes a feature extraction sub-model, a feature fusion sub-model, an encoding-decoding sub-model and a loss calculation sub-model, and the encoding-decoding sub-model is a GRU-Transformer model;

[0038] The processing module is used to input the text data to be corrected into the trained text correction model to obtain the text output result after correction.

[0039] In a third aspect, the present application provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the GRU-Transformer-based text error correction method as described in the first aspect above is implemented.

[0040] In a fourth aspect, the present application provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the GRU-Transformer-based text error correction method as described in the first aspect above.

[0041] In a fifth aspect, the present application provides a chip, comprising a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the GRU-Transformer-based text error correction method as described in the first aspect.

[0042] In a sixth aspect, the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the GRU-Transformer-based text error correction method as described in the first aspect above.

[0043] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through the practice of the present application.

[0044] The text error correction method based on GRU-Transformer provided by the present invention has the following advantages over the prior art:

[0045] Beneficial effects:

[0046] (1) The present invention obtains multiple data sets, including spelling errors, grammatical errors, punctuation errors and similar word error data sets, constructs a text correction data set, combines the feature extraction sub-model, the feature fusion sub-model and the GRU-Transformer encoding-decoding sub-model to process the text data to be corrected, and obtains the text output result after error correction. It can comprehensively capture the error information in the text, effectively handle various types of text errors, identify and correct various types of text errors such as spelling, grammar, word usage and punctuation, and improve the accuracy and comprehensiveness of text correction.

[0047] (2) The present invention combines a pinyin feature extraction module, a character feature extraction module, a word feature extraction module and a stroke feature extraction module to extract pinyin, character, word and stroke features respectively, and can extract features of the text from multiple dimensions, which helps to comprehensively capture the information in the text, improve the ability to recognize and correct spelling, grammatical and structural errors, improve the accuracy and comprehensiveness of text error correction, reduce the propagation of errors, and optimize the text processing effect.

[0048] (3) The present invention uses GRU as an encoder to process the input information after feature fusion and extract the context information in the text. The Transformer decoder further generates the text output after error correction. By combining the sequence modeling ability of GRU and the long-range dependency capture ability of Transformer, the text error correction performance can be effectively improved. The advantages of both can be fully utilized in the text error correction task, the accuracy of text error correction and the ability to correct various complex errors can be improved, and the error types of various open field texts can be effectively dealt with. It has strong practicality and adaptability. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the description of the embodiments in conjunction with the following drawings, in which:

[0050] Figure 1 This is one of the flow charts of the text error correction method based on GRU-Transformer provided in the embodiment of the present application;

[0051] Figure 2 It is a schematic diagram of constructing a text error correction data set provided in an embodiment of the present application;

[0052] Figure 3 This is the second flow chart of the text error correction method based on GRU-Transformer provided in the embodiment of the present application;

[0053] Figure 4 It is a structural diagram of a text error correction device based on GRU-Transformer provided in an embodiment of the present application;

[0054] Figure 5 It is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0055] The following will be combined with the drawings in the embodiments of the present application to clearly describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments in the present application belong to the scope of protection of this application.

[0056] The terms "first", "second", etc. in the specification and claims of this application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described here, and the objects distinguished by "first", "second", etc. are generally of one type, and the number of objects is not limited. For example, the first object can be one or more. In addition, "and / or" in the specification and claims represents at least one of the connected objects, and the character " / " generally indicates that the objects associated with each other are in an "or" relationship.

[0057] In the following, in combination with the accompanying drawings, the GRU-Transformer-based text correction method, the GRU-Transformer-based text correction device, the electronic device and the readable storage medium provided in the embodiments of the present application are described in detail through specific embodiments and their application scenarios.

[0058] Among them, the text error correction method based on GRU-Transformer can be applied to the terminal, and can be specifically executed by hardware or software in the terminal.

[0059] The terminal includes, but is not limited to, a portable communication device such as a mobile phone or tablet computer with a touch-sensitive surface (e.g., a touch screen display and / or a touch pad). It should also be understood that in some embodiments, the terminal may not be a portable communication device, but a desktop computer with a touch-sensitive surface (e.g., a touch screen display and / or a touch pad).

[0060] In the following various embodiments, a terminal including a display and a touch-sensitive surface is described. However, it should be understood that the terminal may include one or more other physical user interface devices such as a physical keyboard, a mouse and a joystick.

[0061] The GRU-Transformer-based text correction method provided in the embodiment of the present application may be executed by an electronic device or a functional module or functional entity in the electronic device that can implement the GRU-Transformer-based text correction method. The electronic devices mentioned in the embodiment of the present application include but are not limited to mobile phones, tablet computers, computers, cameras, and wearable devices. The GRU-Transformer-based text correction method provided in the embodiment of the present application is described below using an electronic device as an example of an execution subject.

[0062] Figure 1 This is one of the flowcharts of the text error correction method based on GRU-Transformer provided in the embodiment of the present application, such as Figure 1 As shown, the text error correction method based on GRU-Transformer includes: step 110, step 120, step 130 and step 140.

[0063] Step 110, obtaining multiple data sets, wherein the data sets include a spelling error data set, a grammatical error data set, a punctuation error data set, and a similar word error data set;

[0064] It is easy to understand that the electronic device obtains multiple general text error correction datasets. Table 1 is a public general text error correction dataset. Table 1 is shown below. These datasets contain text data from various open fields, including spelling error datasets, grammatical error datasets, punctuation error datasets, and similar word error datasets.

[0065] Table 1

[0066]

[0067]

[0068] Step 120: construct a text error correction dataset based on the multiple datasets;

[0069] It is easy to understand that in order to further enhance the diversity and generalization ability of text error correction, the electronic device constructs a text error correction dataset based on multiple datasets. The text error correction dataset includes multiple datasets of specific error types. Each dataset is specially constructed for one error type and can perform detailed corrections for different error types.

[0070] Exemplarily, the error types of the text error correction dataset include the following types:

[0071] (1) Spelling errors: This is the most common type of error in text correction, mainly including similar characters;

[0072] (2) errors, homophone errors, and complete non-existent vocabulary errors were randomly replaced with words / characters in the correct data by using a homophone / homophone dictionary;

[0073] (3) Grammatical errors: This dataset includes common grammatical problems such as disordered word order and unreasonable sentence structure. The word order of correct sentences is randomly disrupted.

[0074] (4) Similar word errors: including word errors that are similar in shape, by randomly replacing words / characters in the correct sentences;

[0075] (5) Punctuation errors: manifested as improper use or incorrect placement of punctuation marks, constructed by randomly deleting and inserting punctuation marks in the correct matrix to construct error-correct sentence pairs;

[0076] (6) Other types of errors: including some errors that are difficult to classify, such as missing words and repetitions, etc., which are constructed by randomly deleting words / phrases, randomly adding reduplications of words, etc.

[0077] Figure 2 is a schematic diagram of constructing a text error correction data set provided by an embodiment of the present application, such as Figure 2 As shown, sentences are obtained from the news set of the news site, the sentences are segmented, and a text correction dataset is constructed based on the segmentation to generate a text correction dataset containing different error types.

[0078] Step 130: training a text error correction model based on the text error correction data set, wherein the text error correction model includes a feature extraction sub-model, a feature fusion sub-model, an encoding-decoding sub-model and a loss calculation sub-model, and the encoding-decoding sub-model is a GRU-Transformer model;

[0079] Furthermore, a preset text error correction model is constructed and trained based on a text error correction data set. The preset text error correction model includes a feature extraction sub-model, a feature fusion sub-model, an encoding-decoding sub-model and a loss calculation sub-model. The feature extraction sub-model is used to perform multi-feature extraction to obtain a feature vector. The feature fusion sub-model is used to use a weighted fusion strategy to weightedly fuse various feature vectors to obtain a fusion vector. The encoding-decoding sub-model includes a GRU encoder and a Transformer decoder. The GRU encoder is used to process the fusion vector and extract the contextual information of the text. The Transformer decoder is used to generate the final error correction text output and combine the self-attention mechanism to capture the dependency of each position in the text. The loss calculation sub-model is used to optimize the parameters of the preset text error correction model.

[0080] Step 140: input the text data to be corrected into the trained text correction model to obtain a corrected text output result.

[0081] Finally, the text to be corrected is input into the trained text correction model, and the text to be corrected is corrected by the text correction model to obtain the corrected text output result.

[0082] According to the GRU-Transformer-based text correction method provided in the embodiment of the present application, a text correction dataset is constructed by acquiring multiple data sets, including spelling errors, grammatical errors, punctuation errors and similar word error datasets, and the feature extraction sub-model, feature fusion sub-model and GRU-Transformer encoding-decoding sub-model are combined to process the text data to be corrected to obtain a corrected text output result. The method can comprehensively capture the error information in the text, effectively handle various types of text errors, effectively identify and correct various types of text errors such as spelling, grammar, vocabulary and punctuation, and improve the accuracy and comprehensiveness of text correction.

[0083] In some embodiments, the feature extraction sub-model includes:

[0084] Pinyin feature extraction module, used to convert Chinese characters into pinyin to obtain pinyin features;

[0085] The character feature extraction module is used to split the sentence into individual characters and obtain character features;

[0086] The word feature extraction module is used to split the sentence into words and obtain word features;

[0087] The stroke feature extraction module is used to obtain the stroke sequence of Chinese characters and obtain stroke features.

[0088] It is easy to understand that when building a text error correction model, feature selection and extraction are crucial. The feature extraction sub-model includes pinyin feature extraction module, character feature extraction module, word feature extraction module and stroke feature extraction module.

[0089] The pinyin feature extraction module is used to convert Chinese characters into pinyin and obtain pinyin features. Pinyin features are helpful to correct homophone errors. The pinyin library can be used to convert Chinese characters into corresponding pinyin representations and extract pinyin features.

[0090] The character feature extraction module is used to split the sentence into individual characters to obtain character features. The similarity between characters can be captured by character feature extraction. The similarity between characters can be captured by splitting the sentence into individual characters to obtain character features.

[0091] The word feature extraction module is used to split sentences into words and obtain word features. Word features can help understand the relationship between wrong words and correct words at the semantic level. It can be implemented using an efficient word segmentation tool library to obtain word features after sentence segmentation.

[0092] The stroke feature extraction module is used to obtain the stroke order of Chinese characters and obtain stroke features. By calculating the stroke order of Chinese characters, the problem of correcting errors in similar-looking characters can be effectively handled. For example, "wood" and "grain" are similar in appearance but have different stroke orders, and stroke features can help identify such errors.

[0093] In this embodiment, by combining the pinyin feature extraction module, the character feature extraction module, the word feature extraction module, and the stroke feature extraction module to extract pinyin, character, word, and stroke features respectively, the text can be feature-extracted from multiple dimensions, which helps to comprehensively capture the information in the text, improves the ability to identify and correct spelling, grammar, and structural errors, enhances the accuracy and comprehensiveness of text error correction, reduces the propagation of errors, and optimizes the text processing effect.

[0094] In some embodiments, the feature fusion sub-model includes:

[0095] A vector embedding module, which is used to convert features into feature vectors through an embedding layer;

[0096] A weighted fusion module, which is used to perform weighted fusion on the feature vectors to obtain a fusion vector.

[0097] It is easy to understand that the vector embedding module is used to convert features into feature vectors through an embedding layer. Exemplarily, the pinyin feature, character feature, word feature, and stroke feature are respectively converted into high-dimensional vector representations through corresponding embedding layers. The pinyin feature obtains a pinyin vector by looking up the pinyin embedding layer, and the character feature and word feature obtain a character embedding layer and a word embedding layer through pre-trained word vectors, and the stroke feature represents the morphological information of characters through the stroke embedding layer.

[0098] Furthermore, the weighted fusion module is used to perform weighted fusion on the feature vectors to obtain a fusion vector. Exemplarily, a weighted fusion strategy is adopted to fuse the pinyin feature, character feature, word feature, and stroke feature. Each feature embedding will be assigned a weight according to its importance, and the weight can be learned by the feature fusion sub-model during the training process. The final fusion vector is the weighted sum of various feature vectors. The calculation formula of the fusion vector is as follows:

[0099] E total (t i )=w p ·E pinyin (t i )+w c ·E char (t i )+w w ·E word (t i )+w s ·E stroke (ti )

[0100] Among them, E pinyin (t i ) is the pinyin feature vector, E char (t i ) is the character feature vector, E word (t i ) is the word feature vector, E stroke (t i ) is the stroke feature vector, E total (t i ) is the fusion vector, w p is the weight of the pinyin feature vector, w c is the weight of the character feature vector, w w is the weight of the word feature vector, w s is the weight of the stroke feature vector.

[0101] In this embodiment, by converting various features such as pinyin, characters, words and strokes into feature vectors through a vector embedding module, and combining the weighted fusion module to perform weighted fusion on the feature vectors, various types of text errors can be effectively processed, erroneous information in the text can be captured, and the accuracy and robustness of open-domain text error correction can be improved.

[0102] In some embodiments, the encoding-decoding sub-model includes:

[0103] An encoder module, used for processing the fused vector to obtain context information of the fused vector;

[0104] A decoder module, used to capture the dependency between various positions in the context information and obtain the text output result after error correction;

[0105] Among them, the encoder module is a GRU encoder, and the decoder module is a Transformer decoder.

[0106] It is easy to understand that the encoding-decoding sub-model includes an encoder module and a decoder module. The encoder module is a GRU encoder and the decoder module is a Transformer decoder.

[0107] The GRU encoder is responsible for processing the input information after feature fusion. The feature fusion sub-model extracts multi-dimensional features such as pinyin, characters, words and strokes, and generates a fusion vector E through weighted fusion. total (t i ), the GRU encoder is used to process the fused vector and obtain the context information of the fused vector.

[0108] In one embodiment, the task of the GRU encoder is to model the sequence of the input fusion vector, transforming the feature vector E at each time step into total (t i ) is transformed into a hidden state h t , thereby extracting the contextual information of the text and using it as the basis for subsequent decoding. GRU can effectively capture long-distance dependencies through the gating mechanism. In text error correction, especially when dealing with grammatical and spelling errors, GRU can capture potential error patterns from the sequence. The calculation formula of the GRU encoder is as follows:

[0109] h t =GRU(E total (t t ),h t-1 )

[0110] Among them, E total (t t ) represents the fusion vector at the tth position, h t Represents the context information of the text, h t-1 Represents the context information of the previous time step of the text.

[0111] Furthermore, after being processed by the GRU encoder, the hidden state h t It is passed to the Transformer decoder as a context vector, which further processes the context information and generates the final error-corrected text.

[0112] In one embodiment, the Transformer decoder uses a self-attention mechanism to capture the dependencies between positions in the input sequence and gradually generate the corrected output text. During the decoding process, the decoder not only pays attention to the context information provided by the encoder, but also uses the self-attention mechanism to generate output predictions for each position. The calculation formula of the Transformer decoder is as follows:

[0113] y t = Transformer(h t ,y t-1 )

[0114] Among them, y t is the text output generated by the decoder at step t, y t-1 is the output of the decoder at the previous moment, h t The context information of the text.

[0115] In this embodiment, GRU is used as an encoder to process the input information after feature fusion and extract contextual information in the text, and the Transformer decoder further generates the corrected text output. By combining the sequence modeling capability of GRU and the long-range dependency capture capability of Transformer, the text correction performance can be effectively improved, and the advantages of both can be fully utilized in the text correction task, thereby improving the accuracy of text correction and the ability to correct various complex errors. It can effectively deal with various error types of open domain texts and has strong practicality and adaptability.

[0116] In some embodiments, the loss calculation sub-model includes:

[0117] The cross entropy loss module is used to calculate the difference between the predicted probability and the true label at each moment;

[0118] A back propagation module, used to calculate the gradient of the cross entropy loss module to the text error correction model parameters;

[0119] The parameter optimization module is used to optimize the parameters of the text error correction model.

[0120] It is easy to understand that the loss calculation sub-model includes the cross entropy loss module, the back propagation module and the parameter optimization module. The cross entropy loss module is used to calculate the difference between the predicted probability and the true label at each moment. The cross entropy loss module uses the cross entropy loss function to measure the difference between the predicted text and the true text, and encourages the model to generate words similar to the target text. The calculation formula of the cross entropy loss function is as follows:

[0121]

[0122] Where V is the vocabulary size, y t,i It is the real text. is predictive text, is the cross entropy loss function, and t represents the time step.

[0123] For the entire sequence, the total loss is the average of the losses at each time step. The total loss is calculated as follows:

[0124]

[0125] Where T is the sequence length, is the total loss, is the cross entropy loss function.

[0126] The back-propagation module is used to calculate the gradient of the cross-entropy loss module to the parameters of the text error correction model, calculate the gradient of the loss function to the model parameters through the chain rule, calculate the gradient of the total loss to the model output, and calculate the gradient of the loss to the Transformer decoder, GRU encoder and feature fusion sub-model in turn, and finally obtain the gradient of the loss to all trainable parameters of the model.

[0127] The parameter optimization module is used to optimize the parameters of the text error correction model. For example, the Adam optimizer can be used to update the model parameters. To reduce gradient explosion, the gradient clipping technology can be used to limit the gradient to a reasonable range. The gradient calculation formula is as follows:

[0128]

[0129]

[0130] in, represents the gradient of the total loss function with respect to the model parameters θ, represents the norm of the gradient, threshold represents the maximum allowed norm of the gradient, and θ is the parameter of the text error correction model.

[0131] In this embodiment, the difference between the predicted probability and the true label at each moment is calculated through the cross entropy loss module, so that errors can be captured and adjusted in time. The model parameters are further optimized through the back propagation module and the parameter optimization module, thereby improving the accuracy and stability of the text error correction model.

[0132] In some embodiments, the step of inputting the text data to be corrected into the trained text correction model to obtain a text output result after correction includes:

[0133] Inputting the text data to be corrected into the feature extraction sub-model to obtain phonetic features, character features, word features and stroke features;

[0134] Inputting the pinyin feature, the character feature, the word feature and the stroke feature into the feature fusion sub-model to obtain the fusion vector;

[0135] The fusion vector is input into the encoding-decoding sub-model to obtain the error-corrected text output result.

[0136] Figure 3 This is the second flow chart of the text error correction method based on GRU-Transformer provided in the embodiment of the present application, such as Figure 3As shown in the figure, the text to be corrected is "because we should give up our seats for the elderly", the text to be corrected is input into the feature extraction sub-model to obtain pinyin features, character features, word features and stroke features, the pinyin features, character features, word features and stroke features are input into the feature fusion sub-model to obtain a fusion vector, the fusion vector is passed into the GRU encoder to generate context information, the context information is passed into the Transformer decoder, and the corrected text output is generated as "we should give up our seats for the elderly".

[0137] In this embodiment, by inputting the text data to be corrected into the feature extraction sub-model, pinyin features, character features, word features and stroke features can be extracted to provide rich feature information for subsequent text correction. After these features are input into the feature fusion sub-model, a fusion vector is obtained, which further strengthens the correlation between features; then, the fusion vector is input into the encoding-decoding sub-model to generate the corrected text output result by capturing the context dependency. This improves the accuracy and robustness of text correction.

[0138] In some embodiments, the training of a text error correction model based on the text error correction dataset includes:

[0139] Build a preset text error correction model;

[0140] Based on the loss calculation sub-model, the preset text error correction model is trained according to the text error correction data set to obtain the text error correction model.

[0141] It is easy to understand that model training includes four main steps: forward calculation, loss calculation, back propagation and parameter optimization. The process of model training is as follows:

[0142] (1) Construct a preset text error correction model, input the text error correction data set into the preset text error correction model, and perform forward calculation. In the forward calculation stage, the input text processed by the feature extraction sub-model and the feature fusion sub-model is passed to the GRU encoder for encoding to generate a context information vector. The encoded feature vector is passed to the Transformer decoder to generate the corrected text output. The output of the preset text error correction model is the predicted word at each moment. The predicted word is the text generated based on the input feature vector and context information.

[0143] (2) Based on the loss calculation sub-model, the cross entropy loss function is used to measure the difference between the predicted text and the true text.

[0144] (3) Through the back-propagation algorithm, the gradient of the loss with respect to the parameters of the preset text error correction model is calculated, and the gradient is transferred to each part of the model through the chain rule.

[0145] (4) Optimize parameters. For example, use the Adam optimizer to update model parameters and apply gradient clipping to prevent gradient explosion. During each training, the model continuously adjusts parameters to reduce cross entropy loss, thereby improving the error correction effect.

[0146] (5) Repeat the above steps until the model loss converges or reaches the predetermined number of training rounds, and then end the training.

[0147] It should be noted that the training process is usually divided into multiple epochs. By gradually adjusting the model parameters, each epoch contains multiple iterations of the training data. After each iteration, the model will calculate the loss through the loss function and update the parameters until the predetermined stopping condition is reached, for example, the loss converges or the maximum number of iterations is reached.

[0148] In this embodiment, by constructing a preset text error correction model and training the text error correction data set based on the loss calculation sub-model, a trained text error correction model is obtained, which can effectively reduce the complexity of the text error correction model calculation process and improve the accuracy and reliability of the text error correction model.

[0149] The GRU-Transformer-based text error correction method provided in the embodiment of the present application can be executed by a GRU-Transformer-based text error correction device. In the embodiment of the present application, a GRU-Transformer-based text error correction device executing the GRU-Transformer-based text error correction method is taken as an example to illustrate the GRU-Transformer-based text error correction device provided in the embodiment of the present application.

[0150] The present application also provides a text error correction device based on GRU-Transformer. Figure 4 As shown, the text error correction device based on GRU-Transformer includes: an acquisition module 410, a construction module 420, a training module 430 and a processing module 440.

[0151] An acquisition module 410 is used to acquire multiple data sets, wherein the data sets include a spelling error data set, a grammatical error data set, a punctuation error data set, and a similar word error data set;

[0152] A construction module 420 is used to construct a text error correction dataset based on the multiple datasets;

[0153] A training module 430 is used to train a text error correction model based on the text error correction data set, wherein the text error correction model includes a feature extraction sub-model, a feature fusion sub-model, an encoding-decoding sub-model and a loss calculation sub-model, and the encoding-decoding sub-model is a GRU-Transformer model;

[0154] The processing module 440 is used to input the text data to be corrected into the trained text correction model to obtain a text output result after correction.

[0155] According to the GRU-Transformer-based text correction device provided in the embodiment of the present application, a text correction dataset is constructed by acquiring multiple data sets, including spelling errors, grammatical errors, punctuation errors and similar word error datasets, and the feature extraction sub-model, feature fusion sub-model and GRU-Transformer encoding-decoding sub-model are combined to process the text data to be corrected to obtain a corrected text output result. The device can comprehensively capture error information in the text, effectively handle various types of text errors, effectively identify and correct various types of text errors such as spelling, grammar, vocabulary and punctuation, and improve the accuracy and comprehensiveness of text correction.

[0156] The text error correction device based on GRU-Transformer provided in the embodiment of the present application can achieve Figures 1 to 3 To avoid repetition, the various processes implemented in the embodiment of the text error correction method based on GRU-Transformer are not repeated here.

[0157] In some embodiments, Figure 5 As shown, an embodiment of the present application also provides an electronic device 500, including a processor 501, a memory 502, and a computer program stored in the memory 502 and executable on the processor 501. When the program is executed by the processor 501, each process of the above-mentioned text error correction method embodiment based on GRU-Transformer is implemented, and the same technical effect can be achieved. To avoid repetition, it will not be described here.

[0158] It should be noted that the electronic devices in the embodiments of the present application include the mobile electronic devices and non-mobile electronic devices mentioned above.

[0159] An embodiment of the present application also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the various processes of the above-mentioned GRU-Transformer-based text error correction method embodiment are implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.

[0160] The processor is the processor in the electronic device in the above embodiment. The readable storage medium includes a computer readable storage medium, such as a computer read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0161] An embodiment of the present application also provides a computer program product, including a computer program, which implements the above-mentioned GRU-Transformer-based text error correction method when executed by a processor.

[0162] The processor is the processor in the electronic device in the above embodiment. The readable storage medium includes a computer readable storage medium, such as a computer read-only memory ROM, a random access memory RAM, a magnetic disk or an optical disk.

[0163] An embodiment of the present application further provides a chip, which includes a processor and a communication interface, wherein the communication interface and the processor are coupled, and the processor is used to run programs or instructions to implement the various processes of the above-mentioned GRU-Transformer-based text error correction method embodiment, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.

[0164] It should be understood that the chip mentioned in the embodiments of the present application can also be called a system-level chip, a system chip, a chip system or a system-on-chip chip, etc.

[0165] It should be noted that, in this article, the terms "comprise", "include" or any other variant thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise one..." do not exclude the presence of other identical elements in the process, method, article or device including the element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in reverse order according to the functions involved, for example, the described method may be performed in an order different from that described, and various steps may also be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.

[0166] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a computer software product, which is stored in a storage medium (such as ROM / RAM, disk, CD), including a number of instructions for a terminal (which can be a mobile phone, computer, server, or network device, etc.) to execute the GRU-Transformer-based text error correction method of each embodiment of the present application.

[0167] In the description of this application, "first feature" or "second feature" may include one or more of the features.

[0168] In the description of the present application, “plurality” means two or more.

[0169] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present application, ordinary technicians in this field can also make many forms without departing from the purpose of the present application and the scope of protection of the claims, all of which are within the protection of the present application.

[0170] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "illustrative embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials, or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.

[0171] Although the embodiments of the present application have been shown and described, those skilled in the art will appreciate that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present application, and that the scope of the present application is defined by the claims and their equivalents.

Claims

1. A text error correction method based on GRU-Transformer, characterized in that: The method comprises: Acquire multiple data sets, the data sets including a spelling error data set, a grammatical error data set, a punctuation error data set, and a similar word error data set; Based on the multiple data sets, construct a text error correction data set; Training a text error correction model based on the text error correction data set, wherein the text error correction model includes a feature extraction sub-model, a feature fusion sub-model, an encoding-decoding sub-model and a loss calculation sub-model, and the encoding-decoding sub-model is a GRU-Transformer model; The text data to be corrected is input into the trained text correction model to obtain the corrected text output result.

2. The text error correction method based on GRU-Transformer according to claim 1, characterized in that: The feature extraction sub-model include: Pinyin feature extraction module, used to convert Chinese characters into pinyin to obtain pinyin features; The character feature extraction module is used to split the sentence into individual characters and obtain character features; The word feature extraction module is used to split the sentence into words and obtain word features; The stroke feature extraction module is used to obtain the stroke sequence of Chinese characters and obtain stroke features.

3. The text error correction method based on GRU-Transformer according to claim 1, characterized in that: The feature fusion sub-model includes: Vector embedding module, used to convert features into feature vectors through the embedding layer; The weighted fusion module is used to perform weighted fusion on the feature vectors to obtain a fusion vector.

4. The text error correction method based on GRU-Transformer according to claim 1, characterized in that: The encoding-decoding sub-model includes: An encoder module, used for processing the fused vector to obtain context information of the fused vector; A decoder module, used to capture the dependency between various positions in the context information and obtain the text output result after error correction; Among them, the encoder module is a GRU encoder, and the decoder module is a Transformer decoder.

5. The text error correction method based on GRU-Transformer according to claim 1, characterized in that: The loss calculation sub-model includes: The cross entropy loss module is used to calculate the difference between the predicted probability and the true label at each moment; A back propagation module, used to calculate the gradient of the cross entropy loss module to the text error correction model parameters; The parameter optimization module is used to optimize the parameters of the text error correction model.

6. The text error correction method based on GRU-Transformer according to claim 1, characterized in that: The step of inputting the text data to be corrected into the trained text correction model to obtain a text output result after correction includes: Inputting the text data to be corrected into the feature extraction sub-model to obtain phonetic features, character features, word features and stroke features; Inputting the pinyin feature, the character feature, the word feature and the stroke feature into the feature fusion sub-model to obtain the fusion vector; The fusion vector is input into the encoding-decoding sub-model to obtain the error-corrected text output result.

7. The text error correction method based on GRU-Transformer according to claim 1, characterized in that: The training of the text error correction model based on the text error correction data set includes: Build a preset text error correction model; Based on the loss calculation sub-model, the preset text error correction model is trained according to the text error correction data set to obtain the text error correction model.

8. A text error correction device based on GRU-Transformer, implemented by the text error correction method based on GRU-Transformer according to any one of claims 1 to 7, characterized in that: The device comprises: An acquisition module, used to acquire multiple data sets, wherein the data sets include a spelling error data set, a grammatical error data set, a punctuation error data set, and a similar word error data set; A construction module, used to construct a text error correction dataset based on the multiple datasets; A training module, used for training a text error correction model based on the text error correction data set, wherein the text error correction model includes a feature extraction sub-model, a feature fusion sub-model, an encoding-decoding sub-model and a loss calculation sub-model, and the encoding-decoding sub-model is a GRU-Transformer model; The processing module is used to input the text data to be corrected into the trained text correction model to obtain the text output result after correction.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the GRU-Transformer-based text error correction method as described in any one of claims 1 to 7 is implemented.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the GRU-Transformer-based text error correction method according to any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Chinese text error correction method based on model fusion and scene self-adaption

    CN120975078A

  • A Chinese text correction method based on model fusion and scene self-adaption

    CN120975078B