A non-canonical text error correction model training method and device, and electronic equipment

By constructing a non-standard text correction training set and training a lightweight student model using the knowledge distillation method, the problems of low accuracy and high resource consumption in existing non-standard text correction methods are solved, achieving efficient non-standard text correction in resource-constrained or real-time-critical scenarios.

CN122347196APending Publication Date: 2026-07-07BEIJING YUNSHANG TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING YUNSHANG TECH CO LTD
Filing Date
2026-04-10
Publication Date
2026-07-07

AI Technical Summary

Technical Problem

In existing technologies, rule-based non-standard text correction methods have low accuracy and robustness when faced with complex and ever-changing non-standard text, while methods based on large language models consume huge amounts of computational resources and have high deployment costs, making them unable to effectively correct errors in scenarios with high real-time requirements or limited resources.

Method used

A non-standard text correction training set is constructed. A lightweight student model is trained by fine-tuning the large language teacher model and using knowledge distillation methods. The feature learning ability of the large teacher model is transferred to obtain a lightweight student model, which is used to obtain standard text results.

Benefits of technology

It improves the accuracy and robustness of lightweight student models, enabling effective error correction of non-standard text in resource-constrained or real-time-critical scenarios, while reducing computational resource consumption and deployment costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122347196A_ABST
    Figure CN122347196A_ABST
Patent Text Reader

Abstract

This invention discloses a training method for a non-standard text error correction model. It involves constructing a non-standard text error correction training set, which includes multiple error correction training instructions and corresponding error correction training text pairs for each instruction. Each error correction training text pair includes both non-standard and standard training texts. Using this training set, a large language classroom model is fine-tuned to obtain a target large language teacher model. Based on the target large language teacher model and the non-standard text error correction training set, a lightweight student model is trained using knowledge distillation until convergence, resulting in a trained target lightweight student model. This target lightweight student model is used to obtain the standard text result of the non-standard text to be corrected. This method improves the accuracy of the standard text result of the non-standard text to be corrected, effectively achieving error correction processing of non-standard text in scenarios with high real-time requirements or limited resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence natural language processing technology, and in particular to a method, apparatus and electronic device for training a non-standard text error correction model. Background Technology

[0002] With the rapid development of mobile internet and social media, non-standard text has become a mainstream form of daily digital communication. Non-standard text is widely found in instant messaging, online comments, forum posts, and voice transcriptions. Its significant characteristic is that it deviates from standard written language norms, specifically manifested in morphological errors, grammatical structure problems, and semantic and pragmatic inconsistencies.

[0003] In existing technologies, on the one hand, traditional rule-based non-standard text correction methods or non-standard text correction methods based on small-scale machine learning models are used to correct non-standard text and obtain its corresponding standard text. On the other hand, non-standard text correction methods based on large language models, such as GPT, LLaMA, ChatGLM, and Qwen, are used to correct non-standard text and obtain its corresponding standard text.

[0004] However, existing technologies, particularly traditional rule-based or small-scale machine learning-based non-standard text correction methods, suffer from low accuracy and robustness when dealing with large amounts of complex and varied non-standard text. Non-standard text correction methods based on large language models require massive datasets for training, and their numerous and complex parameters lead to enormous computational resource consumption and high deployment costs. This makes them unsuitable for scenarios with high real-time requirements or limited resources. Summary of the Invention

[0005] The purpose of this invention is to provide a training method for a non-standard text error correction model, addressing the shortcomings of existing traditional rule-based or small-scale machine learning-based non-standard text error correction methods, which suffer from low accuracy and robustness when faced with large amounts of complex and varied non-standard text. Non-standard text error correction methods based on large language models require massive datasets for training, and their numerous parameters and complex structures lead to huge computational resource consumption and high deployment costs. Furthermore, they are unable to achieve error correction of non-standard text in scenarios with high real-time requirements or limited resources.

[0006] To achieve the above objectives, in a first aspect, embodiments of the present invention provide a method for training a non-standard text error correction model, the method comprising: Construct a non-standard text error correction training set, wherein the non-standard text error correction training set includes: multiple error correction training instructions and error correction training text pairs corresponding to each error correction training instruction, wherein the error correction training text pairs include: non-standard training text and standard training text; Using the aforementioned non-standard text error correction training set, the large language classroom model is fine-tuned to obtain a well-trained target large language teacher model. Based on the target large language teacher model and the non-standard text error correction training set, a lightweight student model is trained using the knowledge distillation method until the model converges, resulting in a trained target lightweight student model. The target lightweight student model is used to obtain the standard text result corresponding to the non-standard text to be corrected.

[0007] In one embodiment, constructing the non-canonical text correction training set includes: Retrieve multiple specification texts and the corresponding error correction instructions for each specification text; Noise interference is processed on multiple standard texts to obtain the non-standard texts corresponding to each standard text. Based on the multiple standard texts and the non-standard texts corresponding to the multiple standard texts, multiple error correction training text pairs are determined; The non-standard text error correction training set is constructed based on multiple error correction instructions and multiple error correction training text pairs.

[0008] In one embodiment, the step of training a lightweight student model using a knowledge distillation method based on the target large language teacher model and the non-standard text error correction training set until the model converges to obtain a trained target lightweight student model includes: The non-standard text error correction training set is input into the target large language teacher model to obtain the first hidden coding feature, the first logits distribution feature, and the first hidden decoding feature; The non-standard text error correction training set is input into the lightweight student model to obtain a second hidden coding feature aligned with the first hidden coding feature, a second logits distribution feature aligned with the first logits distribution feature, and a second hidden decoding feature aligned with the first hidden decoding feature. The lightweight student model is trained based on the first hidden coding feature, the first logits distribution feature, the first hidden decoding feature, the second hidden coding feature, the second logits distribution feature, and the second hidden decoding feature until the model converges, thus obtaining the trained target lightweight student model.

[0009] In one embodiment, the large language teacher model includes a hidden encoder and a first decoder. The step of inputting the non-standard text correction training set into the target large language teacher model to obtain first hidden encoding features, first logits distribution features, and first hidden decoding features includes: The non-standard text error correction training set is input into the target large language teacher model. The first hidden encoding feature is obtained according to the hidden encoder, and the first logits distribution feature and the first hidden decoding feature are obtained according to the first decoder.

[0010] In one embodiment, the lightweight student model includes an encoder and a second decoder. The step of inputting the non-canonical text correction training set into the lightweight student model to obtain second hidden coding features aligned with the first hidden coding features, second logits distribution features aligned with the first logits distribution features, and second hidden decoding features aligned with the first hidden decoding features includes: The non-standard text error correction training set is input into the lightweight student model, and the second hidden encoding feature is obtained according to the encoder, and the second logits distribution feature and the second hidden decoding feature are obtained according to the second decoder.

[0011] In one embodiment, training the lightweight student model based on the first hidden coding feature, the first logits distribution feature, the first hidden decoding feature, the second hidden coding feature, the second logits distribution feature, and the second hidden decoding feature until the model converges to obtain the trained target lightweight student model includes: Based on the first hidden coding feature and the second hidden coding feature, determine the encoder alignment loss function corresponding to the second encoder; Based on the first logits distribution characteristics and the second logits distribution characteristics, determine the divergence distribution loss function corresponding to the second decoder; Based on the first hidden decoding feature and the second hidden decoding feature, a hidden mean squared error loss function corresponding to the second decoder is determined, wherein the divergence distribution loss function and the hidden mean squared error loss function constitute the decoder alignment loss function corresponding to the second decoder; The target loss function is determined based on the encoder alignment loss function, the divergence distribution loss function, the hidden mean square error loss function, and the initial task loss function. The lightweight student model is trained using the target loss function until the model converges, thus obtaining the target student model.

[0012] In one embodiment, determining the target loss function based on the encoder alignment loss function, the divergence distribution loss function, the hidden mean squared error loss function, and the initial task loss function includes: According to the formula Determine the target loss function.

[0013] in, This represents the initial task loss function. This represents the weight parameters corresponding to the initial task loss function. This represents the encoder alignment loss function. This represents the weight parameters corresponding to the encoder alignment loss function. Represents the divergence distribution loss function. This represents the weight parameters corresponding to the divergence distribution loss function. This represents the hidden mean squared error loss function. This represents the weight parameters corresponding to the hidden mean squared error loss function.

[0014] In one embodiment, the method further includes: Using the target large language teacher model and multiple non-standard texts, the training set of the non-standard text error correction training set is subjected to training set enhancement processing.

[0015] Secondly, embodiments of the present invention provide a training device for a non-standard text error correction model, the device comprising: The training set construction module is used to construct a non-standard text error correction training set, wherein the non-standard text error correction training set includes: multiple error correction training instructions and error correction training text pairs corresponding to each error correction training instruction, and the error correction training text pairs include: non-standard training text and standard training text; The target large language teacher model training module is used to fine-tune the large language classroom model using the non-standard text error correction training set to obtain the trained target large language teacher model. The target lightweight student model training module is used to train the lightweight student model using the knowledge distillation method based on the target large language teacher model and the non-standard text correction training set, until the model converges, and obtain the trained target lightweight student model. The target lightweight student model is used to obtain the standard text result corresponding to the non-standard text to be corrected.

[0016] Thirdly, embodiments of the present invention provide an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the non-standard text error correction model training method described in the first aspect.

[0017] The technical solution provided by the embodiments of the present invention has the following advantages compared with the prior art: This invention provides a method for training a non-standard text error correction model. The method involves constructing a non-standard text error correction training set, which includes multiple error correction training instructions and corresponding error correction training text pairs for each instruction. Each error correction training text pair includes both non-standard and standard training text. Using this training set, a large language classroom model is fine-tuned to obtain a trained target large language teacher model. Based on the target large language teacher model and the non-standard text error correction training set, a lightweight student model is trained using knowledge distillation until convergence, resulting in a trained target lightweight student model. This target lightweight student model is used to obtain the standard text results corresponding to the non-standard text to be corrected. In this way, by using knowledge distillation technology, the feature learning capabilities of a large target teacher model can be efficiently transferred to a lightweight student model, thereby improving the accuracy of the lightweight student model in obtaining the standard text result corresponding to the non-standard text to be corrected. This avoids the problems of huge computational resource consumption and high deployment costs of existing non-standard text correction methods based on large language models. Moreover, it can effectively realize the correction of non-standard text to be corrected in scenarios with high real-time requirements or limited resources, and has good versatility. Attached Figure Description

[0018] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this invention, illustrate exemplary embodiments of the invention and are used to explain the invention, but do not constitute an undue limitation of the invention. In the drawings: Figure 1 A flowchart illustrating a non-standard text error correction model training method provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of a non-standard text error correction model training device provided in an embodiment of the present invention. Detailed Implementation

[0019] To facilitate a clear description of the technical solutions in the embodiments of the present invention, the terms "first" and "second" are used to distinguish identical or similar items with essentially the same function and effect. For example, the first threshold and the second threshold are merely used to distinguish different thresholds and do not limit their order. Those skilled in the art will understand that the terms "first" and "second" do not limit the quantity or execution order, and that the terms "first" and "second" are not necessarily different.

[0020] It should be noted that in this invention, the terms "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in this invention should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.

[0021] In this invention, "at least one" means one or more, and "more than one" means two or more. "And / or" describes the relationship between the associated objects, indicating that three relationships can exist.

[0022] like Figure 1 As shown, Figure 1 This is a flowchart illustrating a non-standard text error correction model training method provided in an embodiment of the present invention, specifically including the following steps: S10: Construct a training set for non-standard text error correction.

[0023] The non-standard text correction training set includes: multiple correction training instructions and a corresponding correction training text pair for each instruction. The correction training text pair includes both non-standard training text and standard training text. The correction training instructions are used to instruct the acquisition of correction results for the standard training text corresponding to the non-standard training text. For example, the correction training instructions could be: "Please correct the following text," or "Please correct the errors in the following text to make it a standard Chinese sentence," but are not limited thereto. This invention is not specifically limited, and those skilled in the art can set the instructions according to the actual situation.

[0024] Non-standard training texts refer to texts with morphological errors, grammatical structure problems, and semantic and pragmatic non-standardization. Standard training texts refer to the standard texts corresponding to non-standard training texts. For example, for morphological errors, the non-standard training text could be "accounting room," while the standard training text corresponding to the non-standard training text "accounting room" is "meeting room." However, this invention is not limited to these specific examples, and those skilled in the art can set them according to the actual situation.

[0025] Optionally, based on the above embodiments, in some embodiments of the present invention, one implementation of S10 may be: S101: Obtain multiple standard texts and the corresponding error correction instructions for each standard text.

[0026] Specifically, this can be achieved by obtaining multiple standard texts from publicly available online corpora and determining corresponding error correction instructions for each standard text.

[0027] It should be noted that the error correction instructions corresponding to multiple standard texts can be the same instruction or different instructions. The specific determination of the error correction instructions corresponding to multiple standard texts is not limited in this invention, and those skilled in the art can set them according to the actual situation.

[0028] S102: Perform noise interference processing on multiple standard texts to obtain the non-standard texts corresponding to each standard text.

[0029] Specifically, noise interference is processed on the multiple standard texts obtained, thereby obtaining the non-standard texts corresponding to each standard text.

[0030] Optionally, based on the above embodiments, in some embodiments of the present invention, noise interference processing of multiple standard texts includes: noise interference processing of multiple standard texts by replacing homophones / near-homophones, noise interference processing of multiple standard texts by randomly inserting / deleting words, noise interference processing of multiple standard texts by replacing homophones / near-homophones, noise interference processing of multiple standard texts by simulating common grammatical error patterns, and noise interference processing of multiple standard texts by inserting internet slang. However, the present invention is not limited thereto, and those skilled in the art can set it according to the actual situation.

[0031] S103: Based on multiple standard texts and the non-standard texts corresponding to each standard text, determine multiple error correction training text pairs.

[0032] Specifically, after obtaining the non-standard texts corresponding to each standard text, each standard text and its corresponding non-standard text are identified as a pair of error correction training texts. Based on multiple standard texts and their corresponding non-standard texts, multiple error correction training text pairs are determined.

[0033] S104: Construct a non-standard text error correction training set based on multiple error correction instructions and multiple error correction training text pairs.

[0034] Specifically, after obtaining multiple error correction training text pairs corresponding to multiple standard texts, a non-standard text error correction training set is constructed based on the error correction instructions corresponding to the multiple standard texts and the multiple error correction training text pairs.

[0035] S11: Using a non-standard text error correction training set, fine-tune the large language classroom model to obtain a well-trained target large language teacher model.

[0036] The "large language classroom model" refers to a model that requires massive amounts of data for training and has a complex structure with numerous parameters. This large language classroom model could be, for example, Qwen-7B-Chat, but is not limited to this. This invention does not impose specific limitations, and those skilled in the art can configure it according to actual circumstances.

[0037] Fine-tuning training refers to the efficient fine-tuning of the large language classroom model using low-rank adaptation techniques. During training, only a subset of the injected low-rank parameters are trained, while the majority of the original model's parameters remain unchanged. This method reduces the computational load on the large language classroom model, saving resources.

[0038] Specifically, based on the constructed non-standard text error correction training set, the large language classroom model is fine-tuned and trained using low-rank adaptation technology to obtain the trained target large language teacher model.

[0039] S12: Based on the target large language teacher model and the non-standard text error correction training set, the lightweight student model is trained using the knowledge distillation method until the model converges, resulting in the trained target lightweight student model.

[0040] Among them, the knowledge distillation method refers to transferring the knowledge of a large "teacher model" to a small "student model", so that the small "student model" can learn the feature output of the large "teacher model".

[0041] The aforementioned lightweight student model is used to obtain the standard text result corresponding to the non-standard text to be corrected.

[0042] Specifically, using the pre-trained target large language teacher model and the non-standard text error correction training set, a lightweight student model is trained through knowledge distillation until the student model converges, thus obtaining the target student model.

[0043] Optionally, based on the above embodiments, in some embodiments of the present invention, S12 may be implemented as follows: S121: Input the non-standard text error correction training set into the target large language teacher model to obtain the first hidden coding feature, the first logits distribution feature, and the first hidden decoding feature.

[0044] Specifically, the constructed non-standard text error correction training set is input into the target large language teacher model, and the first hidden coding feature, the first logits distribution feature, and the first hidden decoding feature are obtained through the target large language teacher model.

[0045] Optionally, based on the above embodiments, the large language teacher model includes a hidden encoder and a first decoder, and the lightweight student model includes an encoder and a second decoder. It should be noted that the hidden encoder and the first decoder have a more complex network structure and more parameters compared to the encoder and the second decoder. The encoder in the lightweight student model consists of 6 encoding network layers, and the second decoder consists of 6 decoding network layers. Therefore, in some embodiments of the present invention, one implementation of S121 can be: S1211: Input the non-standard text error correction training set into the target large language teacher model, obtain the first hidden encoding feature according to the hidden encoder, obtain the first logits distribution feature and the first hidden decoding feature according to the first decoder.

[0046] Specifically, the constructed non-standard text correction training set is input into the target large language teacher model. The first hidden encoding feature is obtained based on the hidden encoder of the target large language teacher model, and the first logits distribution feature and the first hidden decoding feature are obtained based on the first decoder of the target large language teacher model.

[0047] S122: Input the non-standard text correction training set into the lightweight student model to obtain the second hidden coding feature aligned with the first hidden coding feature, the second logits distribution feature aligned with the first logits distribution feature, and the second hidden decoding feature aligned with the first hidden decoding feature.

[0048] Specifically, the constructed non-standard text correction training set is input into the lightweight student model, and the lightweight student model obtains the second hidden coding feature aligned with the first hidden coding feature, the second logits distribution feature aligned with the first logits distribution feature, and the second hidden decoding feature aligned with the first hidden decoding feature.

[0049] Optionally, based on the above embodiments, the lightweight student model includes: an encoder and a second decoder. Therefore, in some embodiments of the present invention, one implementation of S122 may be: S1221: Input the non-standard text correction training set into the lightweight student model, obtain the second hidden coding features based on the encoder, and obtain the second logits distribution features and the second hidden decoding features based on the second decoder.

[0050] Specifically, the constructed non-standard text correction training set is input into the lightweight student model. The encoder included in the lightweight student model obtains the second hidden coding feature aligned with the first hidden coding feature. The second decoder included in the lightweight student model obtains the second logits distribution feature aligned with the first logits distribution feature and the second hidden decoding feature aligned with the first hidden decoding feature.

[0051] S123: Train the lightweight student model based on the first hidden coding feature, the first logits distribution feature, the first hidden decoding feature, the second hidden coding feature, the second logits distribution feature, and the second hidden decoding feature until the model converges, and obtain the trained target lightweight student model.

[0052] Specifically, after obtaining the first hidden coding feature, the first logits distribution feature, the first hidden decoding feature, the second hidden coding feature aligned with the first hidden coding feature, the second logits distribution feature aligned with the first logits distribution feature, and the second hidden decoding feature aligned with the first hidden decoding feature, the lightweight student model is trained based on the first hidden coding feature, the first logits distribution feature, the first hidden decoding feature, the second hidden coding feature, the second logits distribution feature, and the second hidden decoding feature until the model converges, thus obtaining the trained target lightweight student model.

[0053] Optionally, based on the above embodiments, in some embodiments of the present invention, one implementation of S123 may be: S1231: Determine the encoder alignment loss function corresponding to the second encoder based on the first hidden coding feature and the second hidden coding feature.

[0054] The encoder alignment loss function is used to ensure that the second hidden coding feature acquired by the encoder can be aligned with the first hidden coding feature.

[0055] Specifically, based on the first hidden coding features obtained by the hidden encoder and the second hidden coding features obtained by the encoder, the encoder alignment loss function corresponding to the second encoder is determined.

[0056] Optionally, based on the above embodiments, in some embodiments of the present invention, the encoder alignment loss function may be defined by the following expression:

[0057] in, The cosine similarity between the first and second hidden coding features is represented by . This represents the attention mask for the i-th token in the b-th training batch. Let represent the total number of training batches, and b represent the b-th training batch. This indicates the total length of the token corresponding to the input object. This represents the i-th token corresponding to the input object. This represents the dimension of the hidden encoded features.

[0058] S1232: Determine the divergence distribution loss function corresponding to the second decoder based on the distribution characteristics of the first logits and the distribution characteristics of the second logits.

[0059] S1233: Determine the hidden mean squared error loss function corresponding to the second decoder based on the first hidden decoding feature and the second hidden decoding feature.

[0060] The divergence distribution loss function and the hidden mean squared error loss function constitute the decoder alignment loss function corresponding to the second decoder. The decoder alignment loss function is used to align the second logits distribution feature obtained by the second decoder with the first logits distribution feature, and to align the second hidden decoded feature with the first hidden decoded feature.

[0061] Specifically, based on the first logits distribution features obtained by the first decoder and the second logits distribution features obtained by the second decoder, the divergence distribution loss function corresponding to the second decoder is determined. Based on the first hidden decoding features obtained by the first decoder and the second hidden decoding features obtained by the second decoder, the hidden mean squared error loss function corresponding to the second decoder is determined.

[0062] Optionally, based on the above embodiments, in some embodiments of the present invention, the divergence distribution loss function may be defined by the following expression:

[0063] in, This represents the distribution characteristics of the first logits at step t. This represents the distribution characteristics of the second logits at step t. This represents the total number of training steps. This represents the temperature parameter.

[0064] Optionally, based on the above embodiments, in some embodiments of the present invention, the hidden mean squared error loss function may be defined by the following expression:

[0065] in, This represents the first hidden decoding feature at step t. This represents the second hidden decoding feature at step t. This represents a learnable linear projection matrix.

[0066] S1234: Determine the target loss function based on the encoder alignment loss function, divergence distribution loss function, hidden mean square error loss function, and initial task loss function.

[0067] The initial task loss function refers to the original loss function of the lightweight student model, which can be, for example, the standard cross-entropy loss function.

[0068] Specifically, the target loss function for training a lightweight student model is constructed by utilizing the encoder alignment loss function, divergence distribution loss function, hidden mean square error loss function, and initial task loss function.

[0069] Optionally, based on the above embodiments, in some embodiments of the present invention, according to the formula... Determine the target loss function.

[0070] in, This represents the initial task loss function. This represents the weight parameters corresponding to the initial task loss function. This represents the encoder alignment loss function. This represents the weight parameters corresponding to the encoder alignment loss function. Represents the divergence distribution loss function. This represents the weight parameters corresponding to the divergence distribution loss function. This represents the hidden mean squared error loss function. This represents the weight parameters corresponding to the hidden mean squared error loss function. For , , , The value of is not specifically limited in this invention, and those skilled in the art can set it according to the actual situation.

[0071] S1235: Train the lightweight student model using the target loss function until the model converges to obtain the target student model.

[0072] Specifically, the lightweight student model is trained based on the constructed target loss function, and the model parameters are adjusted until the model converges to obtain the target student model.

[0073] Thus, the non-standard text error correction model training method provided in this embodiment constructs a non-standard text error correction training set, which includes multiple error correction training instructions and corresponding error correction training text pairs for each instruction. These error correction training text pairs include both non-standard and standard training texts. Using this training set, a large language classroom model is fine-tuned to obtain a trained target large language teacher model. Based on the target large language teacher model and the non-standard text error correction training set, a lightweight student model is trained using knowledge distillation until convergence, resulting in a trained target lightweight student model. This target lightweight student model is used to obtain the standard text results corresponding to the non-standard text to be corrected. In this way, by using knowledge distillation technology, the feature learning capabilities of a large target teacher model can be efficiently transferred to a lightweight student model, thereby improving the accuracy of the lightweight student model in obtaining the standard text result corresponding to the non-standard text to be corrected. This avoids the problems of huge computational resource consumption and high deployment costs of existing non-standard text correction methods based on large language models. Moreover, it can effectively realize the correction of non-standard text to be corrected in scenarios with high real-time requirements or limited resources, and has good versatility.

[0074] Optionally, based on the above embodiments, in some embodiments of the present invention, the method further includes: Using a target large language teacher model and multiple non-standard texts, training set augmentation is performed on the non-standard text error correction training set.

[0075] Specifically, multiple non-standard texts are input into the target large language teacher model to obtain multiple standard texts corresponding to the multiple non-standard texts. Based on the multiple non-standard texts and the multiple standard texts, the training set of non-standard text error correction is subjected to training set enhancement processing.

[0076] Thus, this embodiment enhances the training set of non-standard text correction by performing training set augmentation, and then uses the enhanced non-standard text correction training set to train the lightweight student model, effectively improving the robustness of the lightweight student model and the accuracy of the standard text results corresponding to the non-standard text to be corrected.

[0077] It should be understood that, although Figure 1 The steps in the flowchart are shown sequentially as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order in which these steps are executed, and they can be performed in other orders. Figure 1At least some of the steps in the process may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but may be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but may be executed in turn or alternately with other steps or at least some of the sub-steps or stages of other steps.

[0078] In one embodiment, such as Figure 2 As shown, Figure 2 This is a schematic diagram of a non-standard text error correction model training device provided in an embodiment of the present invention, including: a training set construction module 10, a target large language teacher model training module 11, and a target lightweight student model training module 12.

[0079] The training set construction module 10 is used to construct a non-standard text error correction training set. The non-standard text error correction training set includes: multiple error correction training instructions and error correction training text pairs corresponding to each error correction training instruction. The error correction training text pairs include: non-standard training text and standard training text.

[0080] The target large language teacher model training module 11 is used to fine-tune the large language classroom model using a non-standard text error correction training set to obtain training module 12. The module 12 is used to train the lightweight student model using the knowledge distillation method based on the target large language teacher model and the non-standard text error correction training set until the model converges to obtain the trained target lightweight student model. The target lightweight student model is used to obtain the standard text result corresponding to the non-standard text to be corrected.

[0081] Thus, the non-standard text error correction model training device provided in this embodiment can construct a non-standard text error correction training set through a training set construction module. This training set includes multiple error correction training instructions and a pair of error correction training texts corresponding to each instruction. Each pair includes both non-standard training text and standard training text. The target large language teacher model training module uses the non-standard text error correction training set to fine-tune the large language classroom model, resulting in a trained target large language teacher model. The target lightweight student model training module uses the target large language teacher model and the non-standard text error correction training set to train the lightweight student model using a knowledge distillation method until the model converges, resulting in a trained target lightweight student model. This target lightweight student model is used to obtain the standard text result corresponding to the non-standard text to be corrected. In this way, by using knowledge distillation technology, the feature learning capabilities of a large target teacher model can be efficiently transferred to a lightweight student model, thereby improving the accuracy of the lightweight student model in obtaining the standard text result corresponding to the non-standard text to be corrected. This avoids the problems of huge computational resource consumption and high deployment costs of existing non-standard text correction methods based on large language models. Moreover, it can effectively realize the correction of non-standard text to be corrected in scenarios with high real-time requirements or limited resources, and has good versatility.

[0082] Specific limitations regarding the training device for non-standard text error correction models can be found in the limitations on the training method for non-standard text error correction models mentioned above, and will not be repeated here. Each module in the aforementioned server can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in the computer device in hardware form, or stored in the memory of the computer device in software form, so that the processor can call and execute the corresponding operations of each module.

[0083] This invention provides an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it can implement a non-canonical text error correction model training method provided in this invention. For example, when the processor executes the computer program, it can implement... Figure 1 The technical solutions of the method embodiments shown are similar in principle and in effect, and will not be described again here.

[0084] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the methods described above. Any references to memory, databases, or other media used in the embodiments provided by this invention can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical storage, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static random access memory (SRAM) and dynamic random access memory (DRAM), etc.

[0085] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0086] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these all fall within the protection scope of the present invention. Therefore, the protection scope of this invention patent should be determined by the appended claims.

Claims

1. A method for training a non-standard text error correction model, characterized in that, The method includes: Construct a non-standard text error correction training set, wherein the non-standard text error correction training set includes: multiple error correction training instructions and error correction training text pairs corresponding to each error correction training instruction, wherein the error correction training text pairs include: non-standard training text and standard training text; Using the aforementioned non-standard text error correction training set, the large language classroom model is fine-tuned to obtain a well-trained target large language teacher model. Based on the target large language teacher model and the non-standard text error correction training set, a lightweight student model is trained using the knowledge distillation method until the model converges, resulting in a trained target lightweight student model. The target lightweight student model is used to obtain the standard text result corresponding to the non-standard text to be corrected.

2. The method according to claim 1, characterized in that, The construction of the non-standard text error correction training set includes: Retrieve multiple specification texts and the corresponding error correction instructions for each specification text; Noise interference is processed on multiple standard texts to obtain the non-standard texts corresponding to each standard text. Based on the multiple standard texts and the non-standard texts corresponding to the multiple standard texts, multiple error correction training text pairs are determined; The non-standard text error correction training set is constructed based on multiple error correction instructions and multiple error correction training text pairs.

3. The method according to claim 2, characterized in that, The step of training a lightweight student model using knowledge distillation based on the target large language teacher model and the non-standard text error correction training set until the model converges, to obtain a trained target lightweight student model, includes: The non-standard text error correction training set is input into the target large language teacher model to obtain the first hidden coding feature, the first logits distribution feature, and the first hidden decoding feature; The non-standard text error correction training set is input into the lightweight student model to obtain a second hidden coding feature aligned with the first hidden coding feature, a second logits distribution feature aligned with the first logits distribution feature, and a second hidden decoding feature aligned with the first hidden decoding feature. The lightweight student model is trained based on the first hidden coding feature, the first logits distribution feature, the first hidden decoding feature, the second hidden coding feature, the second logits distribution feature, and the second hidden decoding feature until the model converges, thus obtaining the trained target lightweight student model.

4. The method according to claim 3, characterized in that, The large language teacher model includes a hidden encoder and a first decoder. The step of inputting the non-standard text correction training set into the target large language teacher model to obtain first hidden encoding features, first logits distribution features, and first hidden decoding features includes: The non-standard text error correction training set is input into the target large language teacher model. The first hidden encoding feature is obtained according to the hidden encoder, and the first logits distribution feature and the first hidden decoding feature are obtained according to the first decoder.

5. The method according to claim 4, characterized in that, The lightweight student model includes an encoder and a second decoder. The step of inputting the non-canonical text correction training set into the lightweight student model to obtain second hidden coding features aligned with the first hidden coding features, second logits distribution features aligned with the first logits distribution features, and second hidden decoding features aligned with the first hidden decoding features includes: The non-standard text error correction training set is input into the lightweight student model, and the second hidden encoding feature is obtained according to the encoder, and the second logits distribution feature and the second hidden decoding feature are obtained according to the second decoder.

6. The method according to claim 5, characterized in that, The step of training a lightweight student model based on the first hidden coding feature, the first logits distribution feature, the first hidden decoding feature, the second hidden coding feature, the second logits distribution feature, and the second hidden decoding feature until the model converges, to obtain a trained target lightweight student model, includes: Based on the first hidden coding feature and the second hidden coding feature, determine the encoder alignment loss function corresponding to the second encoder; Based on the first logits distribution characteristics and the second logits distribution characteristics, determine the divergence distribution loss function corresponding to the second decoder; Based on the first hidden decoding feature and the second hidden decoding feature, a hidden mean squared error loss function corresponding to the second decoder is determined, wherein the divergence distribution loss function and the hidden mean squared error loss function constitute the decoder alignment loss function corresponding to the second decoder; The target loss function is determined based on the encoder alignment loss function, the divergence distribution loss function, the hidden mean square error loss function, and the initial task loss function. The lightweight student model is trained using the target loss function until the model converges, thus obtaining the target student model.

7. The method according to claim 6, characterized in that, Based on the encoder alignment loss function, the divergence distribution loss function, the hidden mean squared error loss function, and the initial task loss function, the target loss function is determined, including: According to the formula Determine the target loss function; in, This represents the initial task loss function. This represents the weight parameters corresponding to the initial task loss function. This represents the encoder alignment loss function. This represents the weight parameters corresponding to the encoder alignment loss function. Represents the divergence distribution loss function. This represents the weight parameters corresponding to the divergence distribution loss function. This represents the hidden mean squared error loss function. This represents the weight parameters corresponding to the hidden mean squared error loss function.

8. The method according to claim 1, characterized in that, The method further includes: Using the target large language teacher model and multiple non-standard texts, the training set of the non-standard text error correction training set is subjected to training set enhancement processing.

9. A training device for a non-standard text error correction model, characterized in that, The device includes: The training set construction module is used to construct a non-standard text error correction training set, wherein the non-standard text error correction training set includes: multiple error correction training instructions and error correction training text pairs corresponding to each error correction training instruction, and the error correction training text pairs include: non-standard training text and standard training text; The target large language teacher model training module is used to fine-tune the large language classroom model using the non-standard text error correction training set to obtain the trained target large language teacher model. The target lightweight student model training module is used to train the lightweight student model using the knowledge distillation method based on the target large language teacher model and the non-standard text correction training set, until the model converges, and obtain the trained target lightweight student model. The target lightweight student model is used to obtain the standard text result corresponding to the non-standard text to be corrected.

10. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the non-standard text error correction model training method according to any one of claims 1 to 8.