Error correction method for multi-thread sequencing with copying mechanism and dynamic weighting
By introducing multi-threaded sorting with replication mechanism and dynamic weighting error correction methods in the ERNIE error correction model, the problems of overcorrection, error correction and single-threaded low efficiency in the existing ERNIE error correction model are solved, and more efficient and accurate text error correction effects are achieved.
Patent Information
- Application Number
- CN202510592090.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-05-09
AI Technical Summary
The existing ERNIE error correction model has problems such as overcorrection, error correction and low single-threaded job efficiency.
The multi-threaded sorting, replication mechanism and dynamic weighting error correction method are adopted. By preprocessing the text to be corrected into a sentence array and dynamically allocating it to a multi-threaded channel for error correction, the Copying-ERNIE error correction model with replication mechanism and dynamic weighting is used to reduce overcorrection and error correction and improve error correction efficiency.
It effectively reduces the generation of over-correction of low-frequency expression to high-frequency expression, improves the accuracy and efficiency of error correction, and performs better especially when dealing with complex texts and low-frequency expressions.
Smart Images

Figure CN120106016A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular to a multi-threaded sorting with replication mechanism and dynamic weighted error correction method. Background Art
[0002] Chinese typo detection is a meaningful but challenging task. It has many application scenarios in practice, including assisting content editors in pre-publishing article detection and text correction in chat conversations. In addition, typo correction is also a very important part of the pre-processing process of many natural language processing (NLP, Neuro-Linguistic Programming) tasks (such as word segmentation, part-of-speech tagging, syntactic analysis, etc.), affecting the performance of many downstream tasks in NLP.
[0003] like Figure 2 As shown in Figure 1, the existing ERNIE error correction model. The existing ERNIE error correction model has an over-correction problem: the ERNIE model tends to mistakenly correct low-frequency expressions as high-frequency expressions. The basic idea of the ERNIE error correction model is to use ERNIE to extract the semantic encoding vector of the text, use softmax to perform multi-classification prediction of the vocabulary size, and take the label with the highest probability as the model output result.
[0004] Based on the existing ERNIE error correction model, there are mainly three technical shortcomings: The existing ERNIE error correction model is a MASK language model. Its pre-training method leads to a tendency to predict high-frequency expressions. The existing ERNIE error correction model generally has the problem of over-correcting low-frequency expressions. For example, "This is not to say" is corrected to "This is not to say"; The generated word list space is very large, and the existing ERNIE error correction model has a low prediction probability for the original word, which easily leads to miscorrection. The ERNIE error correction model can be regarded as a restricted generation model. The model generates a target character from the entire word list space for each character in the input sentence. Since the word list space is very large, it is easy to generate characters that are different from the input. Since most of the characters input in the CSC (Chinese Spelling Check) task are correct, it will lead to over-correction problems; The efficiency of single-threaded job execution is low. The existing ERNIE error correction model corrects errors in a single statement in sequence, and executes a paragraph, an article or a long text to be corrected in sequence, which is inefficient. Summary of the invention
[0005] In view of the above analysis, an embodiment of the present invention aims to provide a multi-threaded sorting with replication mechanism and dynamic weighted error correction method to solve the technical problems of over-correction, miscorrection and low single-thread operation error correction efficiency in existing error correction methods.
[0006] The purpose of the present invention is mainly achieved through the following technical solutions: The present invention provides a multi-threaded sorting with replication mechanism and dynamic weighted error correction method, comprising: Preprocess the text to be corrected and assemble it into a sentence array; Each sentence in the sentence array is cyclically assigned to a multi-threaded channel that is dynamically created, destroyed, and horizontally scalable; each thread reads the sentence input from the corresponding data input channel and performs error correction on the trained Copying-ERNIE error correction model with a copying mechanism and dynamic weighting to obtain the output probability distribution of each character in the sentence, selects the character corresponding to the maximum output probability of each character, obtains the corrected sentence and outputs it to the corresponding result output channel; The error-corrected sentences are read in the index order of the sentence array for fault-tolerant aggregation to obtain the error-corrected text.
[0007] Furthermore, the output probability distribution of each character in the sentence is obtained, including: Perform ERNIE encoding conversion on the sentence read from the data input channel to obtain the semantic feature vector of each character position in the sentence; Based on the semantic feature vector of each character position, a generation probability distribution of each character position is obtained, based on the generation probability distribution, a maximum candidate probability and a generation probability of the original character are obtained, and based on the maximum candidate probability and the generation probability of the original character, a copy probability is calculated; The generation probability distribution and the copy probability are dynamically weighted to obtain the output probability distribution of the characters.
[0008] Furthermore, the ERNIE model is used to perform ERNIE encoding conversion to obtain the semantic feature vector of each character position in the sentence; The ERNIE model includes an input layer, an embedding layer and an encoding layer; The input layer concatenates the sentences read from the corresponding data input channel into a character sequence; The embedding layer maps each character in the character sequence to a character vector sequence of a preset dimension; The encoding layer calculates the contextual semantic features of each character vector sequence through the self-attention mechanism, and generates a semantic feature vector for each character position , , is the number of characters in the sentence.
[0009] Furthermore, generating the generation probability and copy probability of the semantic feature vector of each character position includes: Inputting the semantic feature vector of each character position into a Softmax classifier, performing a linear transformation on the semantic feature vector of each character position to obtain a linearly transformed vector, converting the linearly transformed vector into a probability distribution to obtain a generation probability distribution of a word in a vocabulary corresponding to each character; Based on the generation probability distribution, the generation probability corresponding to the original character is obtained, and the maximum generation probability is selected as the maximum candidate probability; The copy probability of each character position is obtained by dynamically adjusting the difference between the maximum candidate probability and the generation probability of the original character.
[0010] Furthermore, the replication probability is calculated as follows: ; in, , are the weight matrices of the linear transformation layer and the Relu activation layer of the Softmax classifier, is the maximum candidate probability for each generated position, is the generation probability of the original character at each generation position, is the smoothing coefficient.
[0011] Furthermore, the generation probability distribution and the copy probability are dynamically weighted to obtain the output probability distribution of each character position, as follows: ; in, is the output probability distribution, is the replication probability distribution, To generate the probability distribution, is the probability of replication.
[0012] Furthermore, each sentence in the sentence array is dynamically and cyclically assigned to a multi-thread channel, including: Initialize the thread pool; Dynamic creation includes A thread pool of threads, each thread is bound to an independent data input channel and result output channel; When the number of sentences is less than the product of the initial number of threads and the number of thread channel queues, destroy the redundant threads or scale the number of thread channel queues horizontally; When the number of data input channel indexes of the calculated sentences is greater than the thread channel queue capacity, dynamically add threads or horizontally expand to increase the number of channel queues; Calculate the data input channel index to be input for each sentence in the sentence array based on the number of threads and the number of sentences; Each sentence poll is assigned to the data input channel corresponding to the data input channel index.
[0013] Furthermore, the fault-tolerant convergence includes: Each thread sets the timeout period T and the number of retries; After the thread completes the error correction process, it caches the result to the corresponding result output channel; For abnormal threads that have timed out and whose execution times are greater than or equal to the number of retries, determine whether the error correction task in the abnormal thread has been completed; if not, directly output the original sentence in the data input channel corresponding to the abnormal thread to the corresponding result output channel, and continue to poll and read the corrected sentence in the next result output channel; When all sentences are corrected, the corrected sentences in all result output channels are read, and the corrected sentences are reassembled according to the sentence array indexes to obtain the corrected text.
[0014] Furthermore, the preprocessing of the text to be corrected and assembling it into a sentence array includes: The text to be corrected is divided into Sentences, ; Remove spaces and line breaks, and add start and end markers for each sentence; According to the order of the sentences in the text to be corrected, Sentences are assembled into sentence arrays , ; Each array element is a sentence; The dividing elements include commas, periods, exclamation marks and question marks.
[0015] Furthermore, the Copying-ERNIE error correction model is obtained by the following training method: Construct a sample training set with sentences containing incorrect characters as samples and correct sentences as sample labels; Construct the Copying-ERNIE error correction model with replication mechanism and dynamic weighting; Weight matrices for the linear transformation layer and Relu activation layer of the Softmax classifier , Random initialization; smoothing coefficient Initialize as a hyperparameter; The sample training set is input into the model, forward propagation is used to generate the generation probability distribution and replication probability of the character position in each sample, and the output probability distribution of the character is calculated by dynamic weighting; Calculate the cross entropy loss between the model output and the true label; According to the loss value, back propagate to calculate the gradient; use the Adam optimization algorithm to update the model weights and parameters; Iterate forward propagation, loss calculation, backpropagation and gradient update until the loss function is minimized; Use accuracy, recall, and F1 score to evaluate model performance; After the training is completed, the Copying-ERNIE error correction model weights and hyperparameters are saved to obtain the trained Copying-ERNIE error correction model.
[0016] Compared with the prior art, the present invention can achieve at least one of the following beneficial effects: 1. The existing text error correction method detects and corrects typos through pre-trained semantic understanding and generation capabilities, and does not have the copy mechanism and copy probability of the present invention; the present invention introduces a copy mechanism algorithm, obtains the maximum candidate probability and the generation probability of the original character according to the generation probability distribution, and then dynamically generates the copy probability distribution, effectively reducing the output probability of over-correcting low-frequency character expressions into high-frequency character expressions; increases the prediction probability of the original low-frequency character expression, copies the original character to the output generated word with a certain probability, increases the generation probability of the original low-frequency character, and reduces the over-correction problem in the existing model; 2. In the Copying-ERNIE error correction model of the present invention, if the maximum candidate probability is close to the generation probability of the original input word, that is, the confidence of the candidate word with the maximum generation probability is low, it is often impossible to determine whether to output the generated character or copy the original character. Adding a dynamic weighting algorithm can increase the generation probability of the original input character, so that the Copying-ERNIE error correction model of the present invention tends to output the original input word; determining the copy probability based on the maximum candidate probability and the generation probability of the original character reduces the problem of miscorrection, especially when the word list space is large; 3. The present invention introduces sentence segmentation, distribution and aggregation, adopts a multi-threaded concurrent processing mechanism, supports dynamic creation and destruction of thread pools, can dynamically adjust the number of threads (including creation and destruction of threads) according to task load and system resources, and horizontally scale the number of queues in the thread channel, significantly improving the error correction efficiency; it can concurrently execute the entire paragraph or long text for error correction in multiple threads at the same time, while ensuring the order of sentences, improving the error correction execution efficiency; 4. Through the fault-tolerant convergence mechanism, when some threads time out or the threads cannot work (such as threads hang up), the fault-tolerant convergence mechanism in the present invention can quickly handle the exception to avoid long-term thread congestion, continuously monitor the thread status and result data channel, and determine whether the error correction task is completed. If not, the data corresponding to the abnormal thread is input into the original sentence in the channel, and re-input into the newly created multi-threaded channel for error correction again, so as to ensure the integrity, sequence, timeliness and reliability of the output text after error correction; 5. The dynamic weighted algorithm combines the generation probability distribution and the replication probability to optimize the output probability distribution, improve the accuracy and adaptability of error correction, especially when processing complex texts and low-frequency expressions.
[0017] In the present invention, the above-mentioned technical solutions can also be combined with each other to achieve more preferred combination solutions. Other features and advantages of the present invention will be described in the subsequent description, and some advantages can become obvious from the description, or can be understood by practicing the present invention. The purpose and other advantages of the present invention can be realized and obtained through the contents particularly pointed out in the description and the drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The accompanying drawings are only used for the purpose of illustrating specific embodiments and are not to be considered as limiting the present invention. In the entire drawings, the same reference symbols represent the same components; Figure 1 It is a flow chart of a multi-threaded sorting with replication mechanism and dynamic weighted error correction method in an embodiment of the present invention; Figure 2 It is a schematic diagram of the existing ERNIE model; Figure 3 A schematic diagram of multi-thread concurrent processing in an embodiment of the present invention; Figure 4 This is a diagram of the Copying-ERNIE error correction architecture in an embodiment of the present invention; Figure 5 Schematic diagram of the Softmax classifier structure in an embodiment of the present invention. DETAILED DESCRIPTION
[0019] The preferred embodiments of the present invention are described in detail below in conjunction with the accompanying drawings, wherein the accompanying drawings constitute a part of this application and are used together with the embodiments of the present invention to illustrate the principles of the present invention, but are not used to limit the scope of the present invention.
[0020] In order to solve the technical problems of over-correction, miscorrection and low error correction efficiency in existing methods, this paper proposes a multi-threaded sorting error correction method with copying mechanism and dynamic weighting. The Copying-ERNIE error correction model in the present invention introduces a variable copying mechanism based on the existing ERNIE model. When generating for each position, if the maximum candidate probability generated is close to the generation probability of the original input word, the original input is weighted during generation to increase the model's prediction probability for the original word, thereby reducing over-correction and miscorrection; as well as dynamically creating, destroying and horizontally scalable multi-threaded channels to solve the problem of low error correction efficiency in existing methods.
[0021] A specific embodiment of the present invention discloses a multi-threaded sorting with replication mechanism and dynamic weighted error correction method, such as Figure 1As shown, the following steps are included: Step S1, preprocessing the text to be corrected and assembling it into a sentence array; Step S2, cyclically assigning each sentence in the sentence array to a multi-threaded channel that is dynamically created, destroyed and horizontally scalable; each thread reads the sentence input from the corresponding data input channel and performs error correction on the trained Copying-ERNIE error correction model with a copying mechanism and dynamic weighting, obtains the output probability distribution of each character in the sentence, selects the character corresponding to the maximum output probability of each character, obtains the corrected sentence and outputs it to the corresponding result output channel; Step S3, reading the corrected sentences according to the index order of the sentence array for fault-tolerant aggregation to obtain the corrected text.
[0022] Step S1, specifically.
[0023] The text to be corrected can be a sentence, a paragraph or a long text.
[0024] First, the text to be corrected is preprocessed, including: The text to be corrected is segmented using the dividing element as a delimiter to obtain at least one sentence; For each sentence obtained, remove spaces and line breaks, and add a start tag [CLS] and an end tag [SEP] to each sentence; Assemble the preprocessed sentences into a sentence array in sequence.
[0025] The method of preprocessing the text to be corrected and assembling it into a sentence array includes: The text to be corrected is divided into Sentences, ; Remove spaces and line breaks, and add start and end markers for each sentence; According to the order of the sentences in the text to be corrected, Sentences are assembled into sentence arrays , ; Each array element is a sentence; The dividing elements include commas, periods, exclamation marks and question marks.
[0026] Each array element in the sentence array is a sentence, for example: is the first sentence of the text to be corrected, The last sentence of the text to be corrected. The sentence array retains the original order of the sentences in the text to be corrected.
[0027] The function of step S1 is to preprocess the text to be corrected, including sentence segmentation, removal of spaces and line breaks, adding start and end marks, and assembling the text into a sentence array according to the original sentence order in the text to be corrected for subsequent error correction processing.
[0028] Step S2 is divided into steps S21-S22.
[0029] Step S21, cyclically assigning each sentence in the sentence array to a multi-threaded channel that is dynamically created, destroyed, and horizontally scalable.
[0030] Dynamically and cyclically assigning each sentence in the sentence array to a multi-thread channel includes: Initialize the thread pool; Dynamic creation includes A thread pool of threads, each thread is bound to an independent data input channel and result output channel; When the number of sentences is less than the product of the initial number of threads and the number of thread channel queues, destroy the redundant threads or scale the number of thread channel queues horizontally; When the number of data input channel indexes of the calculated sentences is greater than the thread channel queue capacity, dynamically add threads or horizontally expand to increase the number of channel queues; Calculate the data input channel index to be input for each sentence in the sentence array; Each sentence poll is assigned to the data input channel corresponding to the data input channel index.
[0031] For example, use SpringBoot's ThreadPoolTaskExecutor configuration The Springboot thread pool is the hreadPoolTaskExecutor thread pool that comes with Springboot. It is a secondary encapsulation of Spring based on the Java thread pool ThreadPoolExecutor. It is the default thread pool in Spring.
[0032] Initialize the thread pool based on the performance and resource utilization of the thread host, including multiple threads. Define the maximum number of error correction queues in each thread according to demand. The purpose is to start error correction as soon as the error correction sentence is input, thereby improving efficiency.
[0033] Each dynamically created thread is bound to an independent data input channel and result output channel, such as Figure 3 shown.
[0034] When the number of sentences is less than the product of the initial number of threads and the number of thread channel queues, the explanation is as follows: For example, the number of sentences to be corrected is 30, the initial number of threads is 4, and the number of correction queues for each thread is 10; First, the solution to reduce resources is to dynamically destroy the fourth thread; Secondly, the solution to improve efficiency is to horizontally reduce the number of channel queues to 8 without destroying threads; or horizontally reduce the number of queues to 5, and create two more threads and corresponding data input channels and result output channels; Finally, you can also increase the number of thread channel queues. For example, if you increase the number of thread channel queues to 15, it is enough to use two threads for 30 sentences to be corrected, and destroy the third and fourth threads at the same time.
[0035] You can also choose one of the methods to dynamically destroy redundant threads or horizontally scale the number of thread channel queues.
[0036] Arrays Each array element in is distributed sequentially and passed to the corresponding data input channel in a loop.
[0037] Calculate the index of the data input channel to be assigned to each sentence in the sentence array as follows: ; in, is the number of thread channels, .
[0038] Take 4 data input channels as an example: Distributed to data input channel A1, Distributed to data channel A2, Distributed to data channel A3, Distributed to data channel A4, Sent to data channel A1, Distribute to data input channel A2, and recursively.
[0039] The function of step S21 is to dynamically create, destroy and horizontally expand the thread pool, and cyclically allocate each sentence in the sentence array to the multi-thread channel for subsequent efficient error correction processing.
[0040] Step S22, each thread reads the sentence input from the corresponding data input channel and performs error correction using the trained Copying-ERNIE error correction model with copying mechanism and dynamic weighting to obtain the output probability distribution of each character in the sentence, selects the character corresponding to the maximum output probability of each character, obtains the corrected sentence and outputs it to the corresponding result output channel.
[0041] Specifically.
[0042] Each thread reads the sentence input from the corresponding data input channel and uses the trained multi-threaded sorting with copying mechanism and dynamically weighted Copying-ERNIE error correction model to correct errors, and obtains the output probability distribution of each character in the sentence, including: Perform ERNIE encoding conversion on the sentence read from the data input channel to obtain the semantic feature vector of each character position in the sentence; Based on the semantic feature vector of each character position, a generation probability distribution of each character position is obtained, based on the generation probability distribution, a maximum candidate probability and a generation probability of the original character are obtained, and based on the maximum candidate probability and the generation probability of the original character, a copy probability is calculated; The generation probability distribution and the copy probability are dynamically weighted to obtain the output probability distribution of the characters.
[0043] First, perform ERNIE encoding conversion: The semantic feature vector of each character position in the sentence is obtained by encoding conversion using the ERNIE model; The ERNIE model includes an input layer, an embedding layer and an encoding layer; The input layer concatenates the sentences read from the corresponding data input channel into a character sequence; The embedding layer maps each character in the character sequence to a character vector sequence of a preset dimension; The encoding layer calculates the contextual semantic features of each character vector sequence through the self-attention mechanism, and generates a semantic feature vector for each character position , , is the number of characters in the sentence.
[0044] To enter the data input channel A1 corresponding to thread 1 For example, generate Semantic feature vector, , the semantic feature vector represents the semantic information of each character in the sentence to be corrected.
[0045] Second, obtain the generation probability and copy probability of the semantic feature vector of each character position.
[0046] Generating the generation probability and copy probability of the semantic feature vector of each character position includes: Inputting the semantic feature vector of each character position into a Softmax classifier, performing a linear transformation on the semantic feature vector of each character position to obtain a linearly transformed vector, converting the linearly transformed vector into a probability distribution to obtain a generation probability distribution of a word in a vocabulary corresponding to each character; Based on the generation probability distribution, the generation probability corresponding to the original character is obtained, and the maximum generation probability is selected as the maximum candidate probability; The copy probability of each character position is obtained by dynamically adjusting the difference between the maximum candidate probability and the generation probability of the original character.
[0047] The semantic feature vector of each character is input into the softmax classifier for multi-classification. When generating at each position, the corresponding generation probability is obtained , and at the same time copy the input word with a certain probability. The copy probability is determined based on the difference between the maximum candidate probability and the generation probability of the input word.
[0048] like Figure 5 As shown, the Softmax classifier includes an input layer, a linear transformation layer, a Relu activation layer, a Sigmoid layer, and an output layer.
[0049] (1) Input layer, which is used to receive the semantic feature vector of each character position of the sentence output by ERNIE encoding conversion ; (2) Linear transformation layer: Linearly transform the semantic feature vector of each character position and map it to the dimension of the vocabulary size.
[0050] Assuming the vocabulary size is V and the feature vector dimension is D, the linear transformation is expressed as follows: ; in, is the vector after linear transformation, with dimension V; is the weight matrix of the linear transformation layer, and the shape of the weight matrix is (V, D); is the semantic feature vector of the character position, with dimension D.
[0051] (3) ReLU activation layer: introduces nonlinearity, enabling the classifier to learn more complex features.
[0052] ; in, is the vector after Relu activation.
[0053] (4) Sigmoid layer: Converted to probability values for calculating replication probability.
[0054] ; Formula (4) in, is the Relu activation layer weight matrix.
[0055] Convert the linearly transformed vector into a generated probability distribution as follows: ; Among them, is the generation probability of character i, is the i-th probability value after the Sigmoid layer transforms .
[0056] The vocabulary contains approximately 28,000 Chinese characters. The Softmax classifier converts the semantic feature vector output by the ERNIE encoding layer into a generated probability distribution of characters, and performs multi-class prediction of the vocabulary size.
[0057] Determine the copy probability based on the difference between the maximum candidate probability and the generation probability of the input character.
[0058] The said copy probability is calculated as follows: ; Among them, , are the weight matrices of the linear transformation layer and the Relu activation layer of the Softmax classifier respectively, is the maximum candidate probability for each generation position, is the generation probability of the original character for each generation position, is the smoothing coefficient.
[0059] The smaller the difference between and is the smoothing coefficient determined by experiments, which determines the smoothing degree of the generated probability distribution. Exemplarily, takes a value of 0.1.
[0060] Perform a weighted sum of the generated probability distribution of generating characters from the vocabulary and the copy probability distribution of copying from the original characters to obtain the output probability distribution for each character position.
[0061] Dynamically weight the said generated probability distribution and the copy probability to obtain the output probability distribution for each character position as follows: ; Among them, is the output probability distribution, is the copy probability distribution, is the generated probability distribution, is the copy probability.
[0062] is the copy probability distribution generated by the copy operation and is a one-hot vector. For example, for the character "悔" in "患者错过后悔哭", at When represented as a vector, the probability of the character "悔" at its position is 1, and the probabilities of other characters in the vocabulary at their positions are 0.
[0063] It is the generation probability distribution obtained by the Softmax classifier.
[0064] Select the character corresponding to the maximum output probability at each character position, and output the corrected sentence to the corresponding result output channel.
[0065] Exemplarily, the original character is "痛"; Generation probability distribution :
[0066] Copy probability :
[0067] Copy probability distribution :
[0068] Calculation result: For "痛": 0.7 * 1.0 + 0.3 * 0.8 = 0.7 + 0.24 = 0.94 For "通": 0.7 * 0.0 + 0.3 * 0.1 = 0.0 + 0.03 = 0.03 For "铜": 0.7 * 0.0 + 0.3 * 0.05 = 0.0 + 0.015 = 0.015 The final output probability distribution is: p = "痛": 0.94, "通": 0.03, "铜": 0.015,... Select the character "痛" corresponding to the maximum output probability of 0.94.
[0069] The dynamic weighting algorithm makes a trade-off between generation and copying, thereby reducing the over-correction problem of existing error correction models.
[0070] The function of the copying mechanism and the dynamic weighting algorithm is to prevent the model from over-correcting. For example, for the text to be error-corrected "患者错过后会哭", if the existing ERNIE model is used, the error correction result output is "患者错过后悔哭", and the existing ERNIE model over-corrects and corrects "后会" to "后悔"; while the Copying-ERNIE error correction model in the present invention uses the copying mechanism and the dynamic weighting algorithm, successfully retains the character "会" after copying, and avoids the situation of over-correction.
[0071] Among them, the Copying-ERNIE error correction model is obtained through the following training method: Construct a sample training set with sentences containing incorrect characters as samples and correct sentences as sample labels; Construct the Copying-ERNIE error correction model with replication mechanism and dynamic weighting; Weight matrices for the linear transformation layer and Relu activation layer of the Softmax classifier , Random initialization; smoothing coefficient Initialize as a hyperparameter; The sample training set is input into the model, forward propagation is used to generate the generation probability distribution and replication probability of the character position in each sample, and the output probability distribution of the character is calculated by dynamic weighting; Calculate the cross entropy loss between the model output and the true label; According to the loss value, back propagate to calculate the gradient; use the Adam optimization algorithm to update the model weights and parameters; Iterate forward propagation, loss calculation, backpropagation and gradient update until the loss function is minimized; Use accuracy, recall, and F1 score to evaluate model performance; After the training is completed, the Copying-ERNIE error correction model weights and hyperparameters are saved to obtain the trained Copying-ERNIE error correction model.
[0072] The cross entropy loss function is as follows: ; in, is the sentence in the text to be corrected, N is the total number of sentences in the text to be corrected, and the output of the Copying-ERNIE error correction model is , The actual result is .
[0073] The Copying-ERNIE error correction model is measured using accuracy, precision, recall, and F1-score, as shown in formulas (9)-(12): ; ; ; ; in, , , and Represent accuracy, precision, recall and Fraction; Indicates the number of samples that are actually positive examples and are correctly corrected as positive examples by the Copying-ERNIE error correction model; Indicates the number of samples that are actually negative examples and are correctly predicted as negative examples by the Copying-ERNIE error correction model; Indicates the number of samples that are actually negative examples but are mistakenly predicted as positive examples by the Copying-ERNIE error correction model; It indicates the number of samples that are actually positive but are mistakenly predicted as negative by the Copying-ERNIE error correction model.
[0074] The following are two examples of over-correction in existing error correction methods: Example 1: We went to the market to buy vegetables after get off work.
[0075] The existing error correction method is to correct "buy vegetables" to "sell vegetables".
[0076] Example 2: Young people like to play video games.
[0077] With the existing error correction method, "electric" is over-corrected to "telephone".
[0078] Since the fixed collocation "selling vegetables" is more frequent than "buying vegetables" and "telephone" is more frequent than "electric", the model corrects the correct word input into an incorrect word. This requires the model to mitigate the overcorrection problem.
[0079] In the method of this embodiment, the over-correction problem in the existing ERNIE error correction model is solved mainly through the following technical details: (1) Introducing a copying mechanism, the Copying-ERNIE error correction model in the present invention directly copies the original character with a certain probability during the error correction generation process, rather than always generating new characters. This increases the generation probability of the original character, especially when the generated word list space is large, and reduces the possibility of over-correction; (2) Through a dynamic weighting algorithm, the output results of the Copying-ERNIE error correction model are optimized by combining the generation probability and the copy probability. The copy probability is dynamically adjusted according to the difference between the maximum candidate probability and the original character probability. The smaller the difference between the two, the greater the copy probability. Dynamically adjusting the copy probability means that Copying-ERNIE error correction is more inclined to retain the original character, thereby reducing over-correction; (3) Softmax classifier is combined with Sigmoid function. The Sigmoid function is introduced into the Softmax classifier to calculate the replication probability. Converted to a probability value, indicating the probability of copying the original character at the current character position. This probability value determines whether the model directly copies the original character instead of generating a new character.
[0080] These technical details work together to enable the Copying-ERNIE error correction model to more accurately retain correct characters and reduce unnecessary corrections when processing low-frequency expressions and long texts.
[0081] The function of step S22 is to correct each sentence using the trained Copying-ERNIE error correction model through multi-threaded processing and dynamic weighting algorithm, generate the output probability distribution of each character position, and select the character with the maximum probability to reduce the over-correction problem.
[0082] The function of step S2 is to process the sentences to be corrected in parallel through a dynamically scalable thread pool, and to use the Copying-ERNIE model with a replication mechanism and dynamic weighting for error correction, thereby improving efficiency while reducing over-correction and miscorrection problems.
[0083] Step S3, specifically.
[0084] According to the original index order of the sentences in the sentence array, the corrected sentences are read and combined for fault tolerance to obtain the corrected text.
[0085] Processing thread 1 reads data from data channel A1 After the error correction is completed using the Copying-ERNIE error correction model, the result data Send to the result output channel B1; Recursively, processing thread 2 reads data from data channel A2 After error correction, the result data Send to data channel B2; Processing thread 3 reads data from data channel A3 After error correction, the result data Send to data channel B3; repeat in sequence.
[0086] The fault-tolerant convergence includes: Each thread sets the timeout period T and the number of retries; After the thread completes the error correction process, it caches the result to the corresponding result output channel; For abnormal threads that have timed out and whose execution times are greater than or equal to the number of retries, determine whether the error correction task in the abnormal thread has been completed; if not, directly output the original sentence in the data input channel corresponding to the abnormal thread to the corresponding result output channel, and continue to poll and read the corrected sentence in the next result output channel; When all sentences are corrected, the corrected sentences in all result output channels are read, and the corrected sentences are reassembled according to the sentence array indexes to obtain the corrected text.
[0087] For threads that time out and have a read count greater than the retry count, there are two exceptional situations: (1) Timeout exception; (2) The thread cannot work, for example, the thread hangs up.
[0088] Record the exception log and continue to read the corrected sentence in the next result output channel.
[0089] Check and determine whether the error correction task in the channel for these two exceptions has been completed; If it has been completed, record the exception log to the log file; If it has not been completed, record the exception log to the log file, and directly output the original sentence in the data input channel corresponding to the exception thread to the corresponding result output channel.
[0090] Fault-tolerant aggregation avoids the situation where the aggregation thread is blocked for a long time.
[0091] Exemplarily, the aggregation sorting algorithm sequentially reads data b[0], b[1], b[2], b[3], b[4] from data channels B1, B2, B3, B4, B1 in a loop, splices each array element, and outputs it externally.
[0092] Implement multi-threaded concurrent processing of data streams, improving the processing speed while ensuring the order of the data stream before and after.
[0093] Exemplarily, a typical error correction case is shown in Table 1: Table 1: Typical Error Correction Cases
[0094] Here, "悔" is more appropriate, with the meaning of "regretting because of not achieving". When the replication mechanism is not introduced, based on the existing ERNIE model, "悔" is corrected to "会". Compared with other positions, the difference between the maximum generation candidate probability of "悔" and the generation probability of the input character is smaller; in the Copying-ERNIE error correction model of the present invention, the replication probability of the input character is increased during generation, so that the input character ranks first in the generation candidates, thus predicting correctly; Gold usually represents the ideal and accurate error correction result in the field of text error correction. In Table 1, the sentence "患者错过后悔哭" shown by Gold is the correct result after manual proofreading and verification, representing the ideal error correction effect.
[0095] The function of step S3 is to ensure the integrity and order consistency of the corrected text by aggregating the multi-threaded error correction results in the original order of the text to be corrected and implementing a fault-tolerant retry aggregation mechanism.
[0096] In summary, the multi-threaded sorting with replication mechanism and dynamic weighted error correction method according to the embodiment of the present invention has the following beneficial effects: 1. The existing text error correction method detects and corrects typos through pre-trained semantic understanding and generation capabilities, and does not have the copy mechanism and copy probability of the present invention; the present invention introduces a copy mechanism algorithm, obtains the maximum candidate probability and the generation probability of the original character according to the generation probability distribution, and then dynamically generates the copy probability distribution, effectively reducing the output probability of over-correcting low-frequency character expressions into high-frequency character expressions; increases the prediction probability of the original low-frequency character expression, copies the original character to the output generated word with a certain probability, increases the generation probability of the original low-frequency character, and reduces the over-correction problem in the existing model; 2. In the Copying-ERNIE error correction model of the present invention, if the maximum candidate probability is close to the generation probability of the original input word, that is, the confidence of the candidate word with the maximum generation probability is low, it is often impossible to determine whether to output the generated character or copy the original character. Adding a dynamic weighting algorithm can increase the generation probability of the original input character, so that the Copying-ERNIE error correction model of the present invention tends to output the original input word; determining the copy probability based on the maximum candidate probability and the generation probability of the original character reduces the problem of miscorrection, especially when the word list space is large; 3. The present invention introduces sentence segmentation, distribution and aggregation, adopts a multi-threaded concurrent processing mechanism, supports dynamic creation and destruction of thread pools, can dynamically adjust the number of threads (including creation and destruction of threads) according to task load and system resources, and horizontally scale the number of queues in the thread channel, significantly improving the error correction efficiency; it can concurrently execute the entire paragraph or long text for error correction in multiple threads at the same time, while ensuring the order of sentences, improving the error correction execution efficiency; 4. Through the fault-tolerant convergence mechanism, when some threads time out or the threads cannot work (such as threads hang up), the fault-tolerant convergence mechanism in the present invention can quickly handle the exception to avoid long-term thread congestion, continuously monitor the thread status and result data channel, and determine whether the error correction task is completed. If not, the data corresponding to the abnormal thread is input into the original sentence in the channel, and re-input into the newly created multi-threaded channel for error correction again, so as to ensure the integrity, sequence, timeliness and reliability of the output text after error correction; 5. The dynamic weighted algorithm combines the generation probability distribution and the replication probability to optimize the output probability distribution, improve the accuracy and adaptability of error correction, especially when processing complex texts and low-frequency expressions.
[0097] Those skilled in the art will appreciate that all or part of the processes of the above-mentioned embodiments can be implemented by instructing related hardware through a computer program, and the program can be stored in a computer-readable storage medium, wherein the computer-readable storage medium is a disk, an optical disk, a read-only storage memory, or a random access memory, etc.
[0098] The above description is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by any technician familiar with the technical field within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention.
Claims
1. A multi-threaded sorting error correction method with replication mechanism and dynamic weighting, characterized in that: include: Preprocess the text to be corrected and assemble it into a sentence array; Each sentence in the sentence array is cyclically assigned to a multi-threaded channel that is dynamically created, destroyed, and horizontally scalable; each thread reads the sentence input from the corresponding data input channel and performs error correction on the trained Copying-ERNIE error correction model with a copying mechanism and dynamic weighting to obtain the output probability distribution of each character in the sentence, selects the character corresponding to the maximum output probability of each character, obtains the corrected sentence and outputs it to the corresponding result output channel; The error-corrected sentences are read in the index order of the sentence array for fault-tolerant aggregation to obtain the error-corrected text.
2. The method according to claim 1, characterized in that: Get the output probability distribution of each character in the sentence, including: Perform ERNIE encoding conversion on the sentence read from the data input channel to obtain the semantic feature vector of each character position in the sentence; Based on the semantic feature vector of each character position, a generation probability distribution of each character position is obtained, based on the generation probability distribution, a maximum candidate probability and a generation probability of the original character are obtained, and based on the maximum candidate probability and the generation probability of the original character, a copy probability is calculated; The generation probability distribution and the copy probability are dynamically weighted to obtain the output probability distribution of the characters.
3. The method according to claim 2, characterized in that: The ERNIE model is used to perform ERNIE encoding conversion to obtain the semantic feature vector of each character position in the sentence; The ERNIE model includes an input layer, an embedding layer and an encoding layer; The input layer concatenates the sentences read from the corresponding data input channel into a character sequence; The embedding layer maps each character in the character sequence to a character vector sequence of a preset dimension; The encoding layer calculates the contextual semantic features of each character vector sequence through the self-attention mechanism, and generates a semantic feature vector for each character position , , is the number of characters in the sentence.
4. The method according to claim 3, characterized in that: Generating the generation probability and copy probability of the semantic feature vector of each character position includes: Inputting the semantic feature vector of each character position into a Softmax classifier, performing a linear transformation on the semantic feature vector of each character position to obtain a linearly transformed vector, converting the linearly transformed vector into a probability distribution to obtain a generation probability distribution of a word in a vocabulary corresponding to each character; Based on the generation probability distribution, the generation probability corresponding to the original character is obtained, and the maximum generation probability is selected as the maximum candidate probability; The copy probability of each character position is obtained by dynamically adjusting the difference between the maximum candidate probability and the generation probability of the original character.
5. The method according to claim 4, characterized in that: The replication probability is calculated as follows: ; in, , are the weight matrices of the linear transformation layer and the Relu activation layer of the Softmax classifier, is the maximum candidate probability for each generated position, is the generation probability of the original character at each generation position, is the smoothing coefficient.
6. The method according to claim 4, characterized in that: The generation probability distribution and the copy probability are dynamically weighted to obtain the output probability distribution of each character position, as follows: ; in, is the output probability distribution, is the replication probability distribution, To generate the probability distribution, is the probability of replication.
7. The method according to claim 1, characterized in that: Dynamically and cyclically assigning each sentence in the sentence array to a multi-thread channel includes: Initialize the thread pool; Dynamic creation includes A thread pool of threads, each thread is bound to an independent data input channel and result output channel; When the number of sentences is less than the product of the initial number of threads and the number of thread channel queues, destroy the redundant threads or scale the number of thread channel queues horizontally; When the number of data input channel indexes of the calculated sentences is greater than the thread channel queue capacity, dynamically add threads or horizontally expand to increase the number of channel queues; Calculate the data input channel index to be input for each sentence in the sentence array based on the number of threads and the number of sentences; Each sentence poll is assigned to the data input channel corresponding to the data input channel index.
8. The method according to claim 1, characterized in that: The fault-tolerant convergence includes: Each thread sets the timeout period T and the number of retries; After the thread completes the error correction process, it caches the result to the corresponding result output channel; For abnormal threads that have timed out and whose execution times are greater than or equal to the number of retries, determine whether the error correction task in the abnormal thread has been completed; if not, directly output the original sentence in the data input channel corresponding to the abnormal thread to the corresponding result output channel, and continue to poll and read the corrected sentence in the next result output channel; When all sentences are corrected, the corrected sentences in all result output channels are read, and the corrected sentences are reassembled according to the sentence array indexes to obtain the corrected text.
9. The method according to claim 1, characterized in that: The method of preprocessing the text to be corrected and assembling it into a sentence array includes: The text to be corrected is divided into Sentences, ; Remove spaces and line breaks, and add start and end markers for each sentence; According to the order of the sentences in the text to be corrected, Sentences are assembled into sentence arrays , ; Each array element is a sentence; The dividing elements include commas, periods, exclamation marks and question marks.
10. The method according to any one of claims 1 to 9, characterized in that: The Copying-ERNIE error correction model is obtained by the following training method: Construct a sample training set with sentences containing incorrect characters as samples and correct sentences as sample labels; Construct the Copying-ERNIE error correction model with replication mechanism and dynamic weighting; Weight matrices for the linear transformation layer and Relu activation layer of the Softmax classifier , Random initialization; smoothing coefficient Initialize as a hyperparameter; The sample training set is input into the model, forward propagation is used to generate the generation probability distribution and replication probability of the character position in each sample, and the output probability distribution of the character is calculated by dynamic weighting; Calculate the cross entropy loss between the model output and the true label; According to the loss value, back propagate to calculate the gradient; use the Adam optimization algorithm to update the model weights and parameters; Iterate forward propagation, loss calculation, backpropagation and gradient update until the loss function is minimized; Use accuracy, recall, and F1 score to evaluate model performance; After the training is completed, the Copying-ERNIE error correction model weights and hyperparameters are saved to obtain the trained Copying-ERNIE error correction model.
Citation Information
Patent Citations
Chinese text error correction method based on prefix tree merging
CN112597771A
Chinese spelling error correction method and device based on comparative learning and medium
CN116127953A
Text error correction model training method and device and text error correction method and device
CN116956877A
Speech recognition text error correction method based on pronunciation guidance
CN118038873A
Intelligent error processing method and system for cross-platform RPA robot
CN119537082A