Model generation method, text generation method, device, medium, equipment and product
By optimizing the error description generation model, detailed error description information is provided, which solves the problem of insufficient error correction explanation in the existing technology and improves the long-term user experience and language capabilities.
Patent Information
- Application Number
- CN202610450627.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-07
- Publication Date
- 2026-07-03
Smart Images

Figure CN122334490A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, and more specifically, to a model generation method, a text generation method, an apparatus, a medium, a device, and a product. Background Technology
[0002] In existing technologies, when correcting erroneous text input by users, the erroneous segments can usually be accurately located, and corresponding correction suggestions can be provided. This approach achieves surface-level text correction and can assist learners in revising text in the short term. However, because it only provides the correction results without explaining the reasons for the errors, learners find it difficult to understand "why the error occurred" and "how to avoid it," thus limiting its long-term impact on language proficiency improvement. Summary of the Invention
[0003] The purpose of this disclosure is to provide a model generation method, a text generation method, an apparatus, a medium, a device, and a product.
[0004] To achieve the above objectives, according to a first aspect of this disclosure, a model generation method is provided, the method comprising: Obtain first training data, which includes multiple first erroneous texts, a first correct text corresponding to each first erroneous text, and a first modified segment, wherein the modified segment includes erroneous segments in the erroneous texts and the correct segments corresponding to the erroneous segments in the correct texts; The first training data is used as input to a pre-trained candidate error description generation model to obtain the first error description information corresponding to the first modified segment output by the candidate error description generation model. The error description information is used to describe the error information in the error segment. Based on the first error description information, the candidate error description generation model is optimized to obtain the optimized target error description generation model. The target error description generation model is used to output the target error description information corresponding to the target modified fragments in the target error text and the target correct text based on the input target error text and target correct text.
[0005] Optionally, optimizing the candidate error description generation model based on the first error description information to obtain the optimized target error description generation model includes: From the plurality of first modification fragments, determine a plurality of pending modification fragments selected by the user, and from the first error description information, determine the pending error description information corresponding to each pending modification fragment; Obtain the target correct description information and target incorrect description information corresponding to each of the proposed modification fragments, wherein the target correct description information and target incorrect description information are determined by the user based on the proposed incorrect description information of the proposed modification fragments; Based on the target correct description information, the target incorrect description information, and the candidate error description information, the candidate error description generation model is optimized to obtain the optimized target error description generation model; the candidate error description information is the first error description information other than the pending error description information among the plurality of first error description information.
[0006] Optionally, optimizing the candidate error description generation model based on the target correct description information, the target incorrect description information, and the candidate error description information to obtain an optimized target error description generation model includes: Filter the target error description information in the candidate error description information to identify the pseudo-correct description information in the candidate error description information; Based on the first training data, the candidate error description generation model is optimized with the goal of increasing the probability of generating correct target description information and pseudo-correct description information, and decreasing the probability of generating incorrect target description information, so as to obtain the target candidate error description generation model.
[0007] Optionally, the modified segment can be determined in the following way: Align the erroneous text with the correct text and divide it into multiple fragment pairs, each fragment pair including a first fragment in the erroneous text and a second fragment in the correct text corresponding to the first fragment; The segment pair that is inconsistent between the first segment and the second segment is taken as the modified segment.
[0008] Optionally, the candidate error description generation model can be trained in the following way: Obtain second training data, which includes multiple second error texts, a corresponding second correct text for each second error text, and second error description information corresponding to the second modified segment. The preset model is trained based on the second training data to obtain a candidate error description generation model.
[0009] Optionally, training a preset model based on the second training data to obtain a candidate error description generation model includes: The second erroneous text, the second correct text, and the second modified fragment are used as inputs to the preset model, and the second error description information is used as the output of the preset model. The preset model is then trained to obtain the candidate error description generation model.
[0010] According to a second aspect of this disclosure, a text generation method is provided, the method comprising: Receive target text input by the user, wherein the target text includes target error text and target correct text; The target text is used as input to a pre-generated target error description generation model to obtain target error description information corresponding to the target text output by the target error description generation model; the target explanation generation model is generated according to the model generation method described in the first aspect.
[0011] According to a third aspect of this disclosure, a model generation apparatus is provided, the apparatus comprising: The acquisition module is configured to acquire first training data, which includes a plurality of first erroneous texts, a first correct text corresponding to each first erroneous text, and a first modified fragment, wherein the modified fragment includes an erroneous fragment in the erroneous text and a correct fragment corresponding to the erroneous fragment in the correct text; The first training module is configured to take the first training data as input to a pre-trained candidate error description generation model to obtain first error description information corresponding to the first modified segment output by the candidate error description generation model, wherein the error description information is used to describe the error information in the error segment. The second training module is configured to optimize the candidate error description generation model based on the first error description information to obtain the optimized target error description generation model. The target error description generation model is used to output the target error description information corresponding to the target modified fragments in the target error text and the target correct text based on the input target error text and target correct text.
[0012] Optionally, the second training module is further configured as follows: From the plurality of first modification fragments, determine a plurality of pending modification fragments selected by the user, and from the first error description information, determine the pending error description information corresponding to each pending modification fragment; Obtain the target correct description information and target incorrect description information corresponding to each of the proposed modification fragments, wherein the target correct description information and target incorrect description information are determined by the user based on the proposed incorrect description information of the proposed modification fragments; Based on the target correct description information, the target incorrect description information, and the candidate error description information, the candidate error description generation model is optimized to obtain the optimized target error description generation model; the candidate error description information is the first error description information other than the pending error description information among the plurality of first error description information.
[0013] Optionally, the second training module is further configured as follows: Filter the target error description information in the candidate error description information to identify the pseudo-correct description information in the candidate error description information; Based on the first training data, the candidate error description generation model is optimized with the goal of increasing the probability of generating correct target description information and pseudo-correct description information, and decreasing the probability of generating incorrect target description information, so as to obtain the target candidate error description generation model.
[0014] Optionally, the modified segment can be determined in the following way: Align the erroneous text with the correct text and divide it into multiple fragment pairs, each fragment pair including a first fragment in the erroneous text and a second fragment in the correct text corresponding to the first fragment; The segment pair that is inconsistent between the first segment and the second segment is taken as the modified segment.
[0015] Optionally, the candidate error description generation model can be trained in the following way: Obtain second training data, which includes multiple second error texts, a corresponding second correct text for each second error text, and second error description information corresponding to the second modified segment. The preset model is trained based on the second training data to obtain a candidate error description generation model.
[0016] Optionally, training a preset model based on the second training data to obtain a candidate error description generation model includes: The second erroneous text, the second correct text, and the second modified fragment are used as inputs to the preset model, and the second error description information is used as the output of the preset model. The preset model is then trained to obtain the candidate error description generation model.
[0017] According to a fourth aspect of this disclosure, a text generation apparatus is provided, the apparatus comprising: The receiving module is configured to receive target text input by the user, the target text including target error text and target correct text; The output module is configured to take the target text as input to a pre-generated target error description generation model, and obtain target error description information corresponding to the target text output by the target error description generation model; the target explanation generation model is generated according to the model generation method described in the first aspect.
[0018] According to a fifth aspect of this disclosure, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the steps of the method described in the first or second aspect.
[0019] According to a sixth aspect of this disclosure, an electronic device is provided, comprising: A memory on which computer programs are stored; A processor for executing the computer program in the memory to implement the steps of the method described in the first or second aspect.
[0020] According to a seventh aspect of this disclosure, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the steps of the method described in the first or second aspect.
[0021] The above technical solution uses the first training data as input to a pre-trained candidate error description generation model to obtain the first error description information corresponding to the first modified segment output by the candidate error description generation model. Based on the first error description information, the candidate error description generation model is optimized through supervised learning to obtain an optimized target error description generation model. This can effectively improve the interpretability and accuracy of the target error description information generated by the target error description generation model and the target error description information corresponding to the target modified segment in the target correct text, thereby effectively improving the user experience and satisfaction.
[0022] Other features and advantages of this disclosure will be described in detail in the following detailed description section. Attached Figure Description
[0023] The accompanying drawings are provided to further illustrate the present disclosure and form part of the specification. They are used together with the following detailed description to explain the present disclosure, but do not constitute a limitation thereof. In the drawings: Figure 1 This is a flowchart illustrating a model generation method according to an exemplary embodiment; Figure 2 This is a schematic diagram illustrating the determination of a modified segment according to an exemplary embodiment; Figure 3 It is based on Figure 1 The illustrated embodiment shows a flowchart for determining the error description information corresponding to the modified segment; Figure 4 It is based on Figure 1 The illustrated embodiment presents a schematic diagram of error type classification in error description information; Figure 5 It is based on Figure 1 The illustrated embodiment shows a flowchart of a model generation method; Figure 6 This is a flowchart illustrating a model generation method according to an exemplary embodiment; Figure 7 This is a flowchart illustrating a text generation method according to an exemplary embodiment; Figure 8 This is a block diagram illustrating a model generation apparatus according to an exemplary embodiment; Figure 9 This is a block diagram illustrating a text generation apparatus according to an exemplary embodiment; Figure 10 This is a block diagram illustrating an electronic device according to an exemplary embodiment. Detailed Implementation
[0024] The specific embodiments of this disclosure will be described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are for illustration and explanation only and are not intended to limit this disclosure.
[0025] Before detailing the specific implementation methods of this disclosure, the application scenarios of this disclosure are first explained below. This disclosure can be applied to application scenarios where error information is generated in erroneous segments of modified text. In the prior art, when correcting erroneous text input by a user, it is usually possible to accurately locate the erroneous segments in the text and provide corresponding correction suggestions. This processing method achieves surface-level text correction and can assist learners in completing text modification in a short period of time. However, since only the correction results are provided without explaining the cause of the error in the erroneous segment, learners find it difficult to understand "why the error occurred" and "how to avoid it," thus limiting its long-term effect on improving language skills. Although existing technical solutions have attempted to generate explanations of the causes of errors in erroneous segments, the generated explanations still suffer from poor interpretability and accuracy.
[0026] In existing solutions, common methods for generating error explanations primarily rely on pre-defined rules to simply categorize and describe error information within erroneous segments. For example, prompts might include "insert a word," "delete a word," or "replace a word with another word." These explanations typically remain at the surface operational level, lacking deeper semantic or linguistic basis. To provide richer error explanations, it's possible to combine natural language processing (NLP) tools to analyze linguistic features such as word part-of-speech. For instance, in research projects like MuCGEC, error types are divided into two levels: the first level is operational (at the character or word level), including four categories: M (Missing): missing character or word needs to be added; R (Redundant): redundant character or word needs to be deleted; S (Substitute): incorrect character or word needs to be corrected; W (Word-order): word order error needs to be adjusted. The second level is linguistic (applicable only to word-level errors), introducing part-of-speech tags on top of operational types to further refine the error categories. Besides spelling errors (SPELL) and word order errors (W), other errors can be tagged with 14 parts of speech. For example, an "adjective redundancy error" can be represented as "R:ADJ". This level helps to describe errors from a grammatical perspective, but its coverage is limited. However, whether using operation categories alone or combining them with part-of-speech information, existing explanations are still relatively weak. For example, prompts such as deleting the verb "run", inserting the punctuation ".", or replacing "Ya'an" with the noun "security guard" indicate the action and object of the modification in the erroneous segment, but fail to explain why the operation is necessary, and lack explanation of the contextual semantics, grammatical structure, or expression logic. Therefore, existing schemes still suffer from weak overall explanatory power when generating explanations for the causes of errors.
[0027] Currently, some related businesses are attempting to generate error explanations for error segments by designing and debugging "suitable" prompts and combining them with Large Language Model (LLM) APIs. However, this approach suffers from significant instability. First, different designed prompts may result in different error explanations for the same error segment. Second, even using the same prompts, different LLM APIs may generate error explanations with variations in style, detail, and even semantics. Furthermore, even with a fixed model and prompts, multiple calls to the same LLM API to explain the same error segment often produce inconsistent error explanations, exhibiting a degree of randomness. Therefore, while this method of relying on large models to generate explanations improves the naturalness and richness of expression, its results suffer from poor repeatability and consistency, making it difficult to meet the requirements of accurate, stable, and reliable feedback in educational scenarios.
[0028] To address the aforementioned issues, the first training data is used as input to a pre-trained candidate error description generation model to obtain first error description information corresponding to the first modified segment output by the candidate error description generation model. Based on the first error description information, the candidate error description generation model is optimized through supervised learning to obtain an optimized target error description generation model. This effectively improves the interpretability and accuracy of the target error description information generated by the target error generation model and the target error description information corresponding to the target modified segment in the target correct text, thereby effectively improving user experience and satisfaction.
[0029] Figure 1 This is a flowchart illustrating a model generation method according to an exemplary embodiment, such as... Figure 1 As shown, the method may include the following steps.
[0030] Step 101: Obtain the first training data.
[0031] The first training data includes multiple first erroneous texts, a first correct text corresponding to each first erroneous text, and a first modified segment. The modified segment includes erroneous segments within the erroneous texts and their corresponding correct segments within the correct texts.
[0032] In this step, the modified segment can be determined by aligning the erroneous text with the correct text and dividing it into multiple segment pairs, each segment pair including a first segment in the erroneous text and a second segment in the correct text corresponding to the first segment; the segment pairs where the first segment and the second segment are inconsistent are taken as the modified segment.
[0033] For example, the incorrect text could be: I have a little friend, for example, a little monkey, a little dog, a little pig...; the correct text could be: I have many little friends, for example, a little monkey, a little dog, a little pig.... The incorrect text and the correct text can be aligned and divided into multiple fragment pairs. The first fragment pair could be [I have] -> [I have]; the second fragment pair could be [one] -> [many]; the third fragment pair could be [little friend] -> [little friend]; the fourth fragment pair could be [] -> [,]; the fifth fragment pair could be [for example] -> [for example]; the sixth fragment pair could be [say] -> []; the seventh fragment pair could be [little monkey, little dog, little pig...] -> [little monkey, little dog, little pig...]. The second fragment pair where the first fragment is inconsistent with the second fragment, the fourth fragment pair where the first fragment is inconsistent with the second fragment, and the sixth fragment pair where the first fragment is inconsistent with the second fragment can be used as the modified fragments.
[0034] Another example, Figure 2 This is a schematic diagram illustrating the determination of a modified segment according to an exemplary embodiment, such as... Figure 2 As shown, the incorrect text could be: "Nowadays the technologies were improved a lot compared for the last century." The correct text could be: "Nowadays technologies have improved a lot compared to the last century." A text alignment algorithm can be used to output a sequence of segments with consistent boundaries between two strings. Segments with identical top and bottom boundaries are grouped into a single segment pair. The first segment pair could be [Nowdays]->[Nowdays]; the second segment pair could be [the technologles were]->[technologles have]; the third segment pair could be [improved a lot compared]->[improved a lot compared]; the fourth segment pair could be [for]->[to]; and the fifth segment pair could be [the last century]->[the last century]. The second segment pair where the first and second segments are inconsistent, and the fourth segment pair where the first and second segments are inconsistent, can be considered as the modified segments.
[0035] The above technical solution, by aligning the erroneous text with the correct text and dividing it into multiple fragment pairs, and taking the fragment pairs where the first fragment and the second fragment are inconsistent as the modified fragments, can accurately determine the modified fragments of the erroneous text and the correct text, thereby providing a basis for subsequently determining the error description information of the modified fragments.
[0036] Step 102: Use the first training data as input to the pre-trained candidate error description generation model to obtain the first error description information corresponding to the first modified segment output by the candidate error description generation model.
[0037] The error description information is used to describe the error information in the error segment. The error description information includes the error type and the cause of the error corresponding to the error segment. The generation of the error cause must meet three principles: First, the language must be fluent and coherent. Second, it must conform to human cognition and be reasonable. Third, it must be specific to the error correction content and be rigorous. Natural language can be used to represent the description information. To ensure the completeness and richness of the error description information, the provided error description information needs to include the error part, the correction part, linguistic knowledge, the error cause, and modification suggestions.
[0038] Exemplarily, Figure 3 is a flowchart for determining the error description information corresponding to the modified fragment according to the Figure 1 illustrated embodiment, as shown in Figure 3 In step 301, the error text and the correct text are input. The error text can be, Src: I have a little friend, such as a little monkey, a little dog, a little pig... The correct text can be, Tgt: I have many little friends, such as a little monkey, a little dog, a little pig... In step 302, the modified fragments Edits in the error text and the correct text are determined. The first modified fragment Edit1: [a] -> [many]; the second modified fragment Edit2: [] -> [,]; the third modified fragment Edit3: [say] -> []. In step 303, the correct text, the error text, and the modified fragments Edits are input into a pre-trained candidate error description generation model. In step 304, the candidate error description generation model outputs the error description information corresponding to the modified fragments Edits. Among them, the error description information corresponding to the first modified fragment can be that the error type is word misuse, and the error reason is that [a] in the correct text is usually used to represent the singularity of quantity, but it is used improperly here because [a] is not sufficient to express the meaning of "many" little friends that the author wants to express. [a] should be replaced with [many] to accurately express the quantity. Here, "[a]" can be the error part in the error description information, "[many]" can be the correction part in the error description information, "[a] is usually used to represent the singularity of quantity, but it is used improperly here" can be the linguistic knowledge part in the error description information, "because [a] is not sufficient to express the meaning of "many" little friends that the author wants to express" can be the error reason part in the error description information, and "should replace [a] with [many] to accurately express the quantity" can be the modification suggestion part in the error description information. The error description information corresponding to the second modified fragment can be that the error type is punctuation missing, and the error reason is that when listing multiple examples, a comma is usually needed to separate different examples to improve the readability of the sentence. Therefore, [,] should be added between [little friends] and [a little monkey]. The error description information corresponding to the third modified fragment can be that the error type is word redundancy, and the error reason is that [say] in the original sentence is redundant here because [such as] is already sufficient to express the meaning of listing and there is no need to use [say] anymore. [say] should be deleted to improve the fluency of the sentence.
[0039] It should be noted that the error text, the correct text, and a modified fragment are sent into the candidate error description generation model at the same time. If there are multiple modified fragments, the candidate error description generation model is called separately multiple times. For example Figure 3 There are 3 modified fragments, so the candidate error description generation model needs to be called 3 times.
[0040] Figure 4 It is based on Figure 1 The illustrated embodiment presents a schematic diagram of the classification of error types in error description information, such as... Figure 4 As shown, error types can be divided into 5 coarse-grained error types and 16 fine-grained error types. The 5 coarse-grained error types can include punctuation-level errors, spelling-level errors, word-level errors, sentence-level errors, and other errors. The 16 fine-grained error types can include punctuation redundancy (unnecessary insertion of punctuation marks), missing punctuation (omission of necessary punctuation marks in the middle or at the end of a sentence), punctuation misuse (incorrect use of punctuation marks, including but not limited to confusion of punctuation function, hierarchical errors, and symbol form errors), phonetic confusion errors (errors caused by similar or identical pinyin pronunciations of Chinese characters), character shape confusion errors (misspellings caused by similar shapes or strokes between Chinese characters), internal character misplacement errors (incorrect word order in multiple words), named entity spelling errors (Chinese has a large number of named entity words, including personal names, organization names, place names, and other entities referred to by professional terms), word redundancy (sentences containing words with the same or similar meanings), and missing words. (Modern Chinese sentences typically contain six components: subject, predicate, object, attributive, adverbial, and complement. While not every sentence needs to contain all components, essential components that convey complete meaning must be retained. Omission of essential components constitutes a word omission error.) Other issues include: misuse of words (inappropriate word usage in sentences), incorrect word order, illogical sentences (sentences that conform to grammatical rules but violate logical relationships, such as incorrect logical order, confused cause-and-effect relationships, or subject-object reversal), mixed sentence structures (forcibly combining two sentence structures or semantically similar sentences), errors in reference (inappropriate use of referential relationships between words), ambiguity (words or sentences that can have multiple interpretations), and inconsistent tone (inconsistent tone or style between sentences, often due to inappropriate stylistic shifts).
[0041] The above technical solution, by refining error types into 5 to 16 fine-grained error types, enables the target error interpretation generation model to more accurately identify error types, avoiding "general error correction" (such as only prompting "There is an error" or "Please correct"). Furthermore, each error type can be associated with a predefined interpretation template or generation rule, effectively improving the interpretability of the target error description information generated by the target error description generation model, corresponding to the target error text and the target modified fragments in the target correct text. In addition, based on a clear error type system, the feedback generated by the target error description generation model for the target error text and the target error description information corresponding to the target modified fragments in the target correct text is more consistent, reducing the randomness of the large model output, thereby effectively improving the accuracy of the target error description information generated by the target error description generation model.
[0042] Step 103: Optimize the candidate error description generation model based on the first error description information to obtain the optimized target error description generation model.
[0043] The target error description generation model is used to output target error description information corresponding to the target modification fragments in the target error text and the target correct text based on the input target error text and target correct text.
[0044] The above technical solution uses the first training data as input to a pre-trained candidate error description generation model to obtain the first error description information corresponding to the first modified segment output by the candidate error description generation model. Based on the first error description information, the candidate error description generation model is optimized through supervised learning to obtain an optimized target error description generation model. This can effectively improve the interpretability and accuracy of the target error description information generated by the target error description generation model and the target error description information corresponding to the target modified segment in the target correct text, thereby effectively improving the user experience and satisfaction.
[0045] Figure 5 It is based on Figure 1 The illustrated embodiment shows a flowchart of a model generation method. Figure 1 Step 103, which involves optimizing the candidate error description generation model based on the first error description information to obtain the optimized target error description generation model, may include: Step 501: Determine multiple pending modification segments selected by the user from the multiple first modification segments, and determine the pending error description information corresponding to each pending modification segment from the first error description information.
[0046] Step 502: Obtain the target correct description information and target incorrect description information corresponding to each of the proposed modification fragments.
[0047] The target correct description information and the target incorrect description information are determined by the user based on the pending incorrect description information of the pending modification fragment.
[0048] In this step, the target correct description information includes first target correct description information and second target correct description information, and the target error description information includes first target error description information and second target error description information. The target correct description information and target error description information corresponding to each pending modification fragment are determined by the user. For each pending error description information corresponding to a pending modification fragment, if the pending error description information is determined to be correct, it is used as the first target correct description information, and rules are used to construct the first target error description information corresponding to the pending error description information. If the pending error description information is determined to be incorrect, it is used as the second target error description information, and the pending error description information is modified to become the correct second target correct description information.
[0049] Step 503: Optimize the candidate error description generation model based on the target correct description information, the target error description information, and the candidate error description information to obtain the optimized target error description generation model.
[0050] The candidate error description information is the first error description information other than the pending error description information among the plurality of first error description information.
[0051] In this step, the target error description information in the candidate error description information can be filtered to identify pseudo-correct description information. For the first training data, the candidate error description generation model is optimized with the goal of increasing the probability of generating both the target correct description information and the pseudo-correct description information, and decreasing the probability of generating the target error description information, to obtain the target candidate error description generation model.
[0052] For example, the first erroneous text input to the candidate error description generation model is: "The teacher immediately gave Lin Nan a sick note after hearing this." The first correct text is: "The teacher immediately gave Lin Nan a sick note after hearing this." The first modified fragment is []->[,]. The first error description information output by the candidate error description generation model is: Error type: punctuation redundancy, Error description: A comma is missing between "The teacher immediately gave Lin Nan a sick note after hearing this," which makes the pause in the sentence unclear and affects the fluency of the sentence. A comma should be added after "The teacher immediately gave this" to indicate a pause. Obviously, "punctuation redundancy" is an incorrect error type and is the target error description information. Because the first modified fragment shows that a comma is inserted from the first erroneous text to the first correct text, the error type in the first erroneous text should be "missing punctuation." The error type in the first erroneous text can be adjusted to "missing punctuation." The data modified in this way may not be correct, but it is more reasonable than "punctuation redundancy." Such a sample is "pseudo-correct description information." Alternatively, candidate error descriptions similar to the target error description can be filtered to identify pseudo-correct descriptions among the candidate error descriptions.
[0053] In another example, for the first training data, triplet data containing the prompt word "prompt", the preferred answer "chosen", and the rejection answer "reject" are determined. The prompt word includes a first incorrect text, a first correct text, and a first modified fragment. The preferred answer includes target correct description information and pseudo-correct description information. The rejection answer includes target incorrect description information. prompt: First incorrect text: The teacher immediately gave Lin Nan a leave note. First correct text: The teacher immediately gave Lin Nan a leave note. First modified fragment []->[,]. chosen y1: Error type: missing punctuation, error description: A comma is missing between "The teacher immediately gave Lin Nan a leave note", which makes the pause in the sentence unclear and affects the fluency of the sentence. A comma should be added after "The teacher immediately gave Lin Nan a leave note" to indicate a pause. reject y2: Error type: redundant punctuation, error description: A comma is missing between "The teacher immediately gave Lin Nan a leave note", which makes the pause in the sentence unclear and affects the fluency of the sentence. A comma should be added after "[The teacher listened]" to indicate a pause. Note: This is the DPO training data format.<prompt x, chosen y1, reject y2> For a prompt, there are only two answers, and then preference is ranked between these two answers. The candidate error description generation model is trained using the DPO algorithm to generate more texts that prioritize the answer "chosen y1" and fewer texts that reject the answer "reject y2," in order to fine-tune and optimize the model parameters in the candidate error description generation model to obtain the target candidate error description generation model.
[0054] Another example, Figure 6 This is a flowchart illustrating a model generation method according to an exemplary embodiment, such as... Figure 6 As shown, the method may include the following steps.
[0055] Step 601, Start the process.
[0056] Step 602: Obtain a small amount of labeled data.
[0057] The labeled data is the second training data, which includes multiple second error texts, the second correct text corresponding to each second error text, and the second error description information corresponding to the second modified segment.
[0058] Step 603: Based on the labeled data Data1, the preset model is fine-tuned using the LoRa method to obtain the candidate error description generation model Model1.
[0059] In this step, the second erroneous text, the second correct text, and the second modified fragment can be used as inputs to the preset model, and the second error description information can be used as outputs to the preset model. The preset model is trained using LoRa to obtain the candidate error description generation model Model1.
[0060] Step 604: Generate unlabeled data Data2 (first error description information) for Model1 inference based on the second training data and candidate error descriptions.
[0061] Step 605: Obtain a large amount of triplet labeled data (Data3) by manually adding rule annotations.
[0062] The triplet-annotated data Data3 includes the prompt word, the chosen answer, and the rejected answer. The prompt word includes the first incorrect text, the first correct text, and the first modified fragment. The chosen answer includes both correct and pseudo-correct descriptions of the target (correct interpretations of errors). The rejected answer includes incorrect descriptions of the target (erroneous interpretations of errors).
[0063] Step 606: Fine-tune the candidate error description generation model Model1 using the DPO algorithm to obtain the optimized target error description generation model Model2.
[0064] Step 607, end the process.
[0065] The above technical solution trains a candidate error description generation model using a small amount of labeled data and fine-tuning it with LoRa. However, the generalization ability of the candidate error description generation model is poor, mainly due to the limited training data, requiring expansion of the training data. The candidate error description generation model is then used to infer a large number of error explanations from unlabeled data, generating triple data. Based on the triples and the DPO algorithm, the candidate error description generation model is fine-tuned to obtain the target error description generation model. This effectively expands the training data, alleviating the problem of high manual annotation costs. Furthermore, through a preference learning mechanism, the target error description generation model learns to distinguish between high-quality and low-quality explanations, thereby improving the interpretability (e.g., semantic clarity and logical rationality) and stability (reducing output fluctuations during repeated calls) of the target error description information generated by the model, and solving the problems of poor interpretability and stability of target error description information generated by existing solutions.
[0066] Figure 7 This is a flowchart illustrating a text generation method according to an exemplary embodiment, such as... Figure 7 As shown, the method may include the following steps.
[0067] Step 701: Receive target text input by the user, wherein the target text includes target error text and target correct text.
[0068] Step 702: Use the target text as input to the pre-generated target error description generation model to obtain the target error description information corresponding to the target text output by the target error description generation model.
[0069] The target explanation generation model is based on Figures 1-6 The model was generated by the aforementioned model generation method.
[0070] The above technical solution uses the target text as input to a pre-generated target error description generation model to obtain the target error description information corresponding to the target text output by the target error description generation model, which can effectively improve the user experience and satisfaction.
[0071] Figure 8 This is a block diagram illustrating a model generation apparatus according to an exemplary embodiment, such as... Figure 8 As shown, the model generation device may include: The acquisition module 801 is configured to acquire first training data, the first training data including a plurality of first erroneous texts, a first correct text corresponding to each first erroneous text, and a first modified segment, wherein the modified segment includes an erroneous segment in the erroneous text and a correct segment corresponding to the erroneous segment in the correct text; The first training module 802 is configured to take the first training data as input to a pre-trained candidate error description generation model to obtain first error description information corresponding to the first modified segment output by the candidate error description generation model, wherein the error description information is used to describe the error information in the error segment. The second training module 803 is configured to optimize the candidate error description generation model based on the first error description information to obtain an optimized target error description generation model. The target error description generation model is used to output target error description information corresponding to the target modified segments in the target error text and the target correct text based on the input target error text and target correct text.
[0072] Optionally, the second training module 803 is further configured to: From the plurality of first modification fragments, determine a plurality of pending modification fragments selected by the user, and from the first error description information, determine the pending error description information corresponding to each pending modification fragment; Obtain the target correct description information and target incorrect description information corresponding to each of the proposed modification fragments, wherein the target correct description information and target incorrect description information are determined by the user based on the proposed incorrect description information of the proposed modification fragments; Based on the target correct description information, the target incorrect description information, and the candidate error description information, the candidate error description generation model is optimized to obtain the optimized target error description generation model; the candidate error description information is the first error description information other than the pending error description information among the plurality of first error description information.
[0073] Optionally, the second training module 803 is further configured to: Filter the target error description information in the candidate error description information to identify the pseudo-correct description information in the candidate error description information; Based on the first training data, the candidate error description generation model is optimized with the goal of increasing the probability of generating correct target description information and pseudo-correct description information, and decreasing the probability of generating incorrect target description information, so as to obtain the target candidate error description generation model.
[0074] Optionally, the modified segment can be determined in the following way: Align the erroneous text with the correct text and divide it into multiple fragment pairs, each fragment pair including a first fragment in the erroneous text and a second fragment in the correct text corresponding to the first fragment; The segment pair that is inconsistent between the first segment and the second segment is taken as the modified segment.
[0075] Optionally, the candidate error description generation model can be trained in the following way: Obtain second training data, which includes multiple second error texts, a corresponding second correct text for each second error text, and second error description information corresponding to the second modified segment. The preset model is trained based on the second training data to obtain a candidate error description generation model.
[0076] Optionally, training a preset model based on the second training data to obtain a candidate error description generation model includes: The second erroneous text, the second correct text, and the second modified fragment are used as inputs to the preset model, and the second error description information is used as the output of the preset model. The preset model is then trained to obtain the candidate error description generation model.
[0077] Figure 9 This is a block diagram illustrating a text generation apparatus according to an exemplary embodiment, such as... Figure 8 As shown, the text generation device may include: The receiving module 901 is configured to receive target text input by a user, the target text including target error text and target correct text; The output module 902 is configured to take the target text as input to a pre-generated target error description generation model to obtain target error description information corresponding to the target text output by the target error description generation model; the target explanation generation model is generated according to the model generation method described in the first aspect.
[0078] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.
[0079] Figure 10 This is a block diagram illustrating an electronic device 1000 according to an exemplary embodiment. For example... Figure 10 As shown, the electronic device 1000 may include: a processor 1001 and a memory 1002. The electronic device 1000 may also include one or more of a multimedia component 1003, an input / output (I / O) interface 1004, and a communication component 1005.
[0080] The processor 1001 controls the overall operation of the electronic device 1000 to complete all or part of the steps in the model generation method and / or text generation method described above. The memory 1002 stores various types of data to support the operation of the electronic device 1000. This data may include, for example, instructions for any application or method operating on the electronic device 1000, and application-related data such as contact data, sent and received messages, images, audio, video, etc. The memory 1002 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. Multimedia component 1003 may include a screen and an audio component. The screen may be, for example, a touchscreen, and the audio component is used to output and / or input audio signals. For example, the audio component may include a microphone for receiving external audio signals. The received audio signals may be further stored in memory 1002 or transmitted via communication component 1005. The audio component also includes at least one speaker for outputting audio signals. I / O interface 1004 provides an interface between processor 1001 and other interface modules, such as a keyboard, mouse, buttons, etc. These buttons may be virtual or physical buttons. Communication component 1005 is used for wired or wireless communication between the electronic device 1000 and other devices. Wireless communication may include Wi-Fi, Bluetooth, Near Field Communication (NFC), 2G, 3G, or 4G, or a combination of these. Therefore, the corresponding communication component 1005 may include a Wi-Fi module, a Bluetooth module, or an NFC module.
[0081] In an exemplary embodiment, the electronic device 1000 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the model generation method and / or text generation method described above.
[0082] In another exemplary embodiment, a computer-readable storage medium including program instructions is also provided, which, when executed by a processor, implement the steps of the model generation method and / or text generation method described above. For example, the computer-readable storage medium may be the memory 1002 including the program instructions described above, which may be executed by the processor 1001 of the electronic device 1000 to complete the model generation method and / or text generation method described above.
[0083] In another exemplary embodiment, a computer program product is also provided, which includes a computer program executable by a processor, which, when executed by the processor, implements the steps of the model generation method and / or text generation method described above.
[0084] In another exemplary embodiment, a computer program product is also provided, which includes a computer program executable by a processor, which, when executed by the processor, implements the steps of the model generation method and / or text generation method described above.
[0085] The preferred embodiments of this disclosure have been described in detail above with reference to the accompanying drawings. However, this disclosure is not limited to the specific details of the above embodiments. Within the scope of the technical concept of this disclosure, various simple modifications can be made to the technical solutions of this disclosure, and these simple modifications all fall within the protection scope of this disclosure.
[0086] It should also be noted that the various specific technical features described in the above specific embodiments can be combined in any suitable manner without contradiction. In order to avoid unnecessary repetition, this disclosure will not describe the various possible combinations separately.
[0087] Furthermore, various different embodiments of this disclosure can be combined in any way, as long as they do not violate the spirit of this disclosure, they should also be regarded as the content disclosed in this disclosure.
Claims
1. A model generation method characterized by comprising: The method includes: Obtain first training data, which includes multiple first erroneous texts, a first correct text corresponding to each first erroneous text, and a first modified segment, wherein the modified segment includes erroneous segments in the erroneous texts and the correct segments corresponding to the erroneous segments in the correct texts; The first training data is used as input to a pre-trained candidate error description generation model to obtain the first error description information corresponding to the first modified segment output by the candidate error description generation model. The error description information is used to describe the error information in the error segment. Based on the first error description information, the candidate error description generation model is optimized to obtain the optimized target error description generation model. The target error description generation model is used to output the target error description information corresponding to the target modified fragments in the target error text and the target correct text based on the input target error text and target correct text.
2. The model generation method according to claim 1, wherein The step of optimizing the candidate error description generation model based on the first error description information to obtain the optimized target error description generation model includes: From the plurality of first modification fragments, determine a plurality of pending modification fragments selected by the user, and from the first error description information, determine the pending error description information corresponding to each pending modification fragment; Obtain the target correct description information and target incorrect description information corresponding to each of the proposed modification fragments, wherein the target correct description information and target incorrect description information are determined by the user based on the proposed incorrect description information of the proposed modification fragments; Based on the target correct description information, the target incorrect description information, and the candidate error description information, the candidate error description generation model is optimized to obtain the optimized target error description generation model; the candidate error description information is the first error description information other than the pending error description information among the plurality of first error description information.
3. The model generation method according to claim 2, characterized by, The step of optimizing the candidate error description generation model based on the target correct description information, the target incorrect description information, and the candidate error description information to obtain an optimized target error description generation model includes: Filter the target error description information in the candidate error description information to identify the pseudo-correct description information in the candidate error description information; Based on the first training data, the candidate error description generation model is optimized with the goal of increasing the probability of generating correct target description information and pseudo-correct description information, and decreasing the probability of generating incorrect target description information, so as to obtain the target candidate error description generation model.
4. The model generation method according to claim 1, characterized by, The modified fragment can be determined in the following ways: Align the erroneous text with the correct text and divide it into multiple fragment pairs, each fragment pair including a first fragment in the erroneous text and a second fragment in the correct text corresponding to the first fragment; The segment pair that is inconsistent between the first segment and the second segment is taken as the modified segment.
5. The model generation method according to claim 1, wherein The candidate error description generation model can be trained in the following way: Obtain second training data, which includes multiple second error texts, a corresponding second correct text for each second error text, and second error description information corresponding to the second modified segment. The preset model is trained based on the second training data to obtain a candidate error description generation model.
6. The model generation method according to claim 5, characterized in that, The step of training a preset model based on the second training data to obtain a candidate error description generation model includes: The second erroneous text, the second correct text, and the second modified fragment are used as inputs to the preset model, and the second error description information is used as the output of the preset model. The preset model is then trained to obtain the candidate error description generation model.
7. A text generation method, characterized in that, The method includes: Receive target text input by the user, wherein the target text includes target error text and target correct text; The target text is used as input to a pre-generated target error description generation model to obtain target error description information corresponding to the target text output by the target error description generation model; the target interpretation generation model is generated by the model generation method according to any one of claims 1 to 6.
8. A model generation apparatus, characterized in that, The device includes: The acquisition module is configured to acquire first training data, which includes a plurality of first erroneous texts, a first correct text corresponding to each first erroneous text, and a first modified fragment, wherein the modified fragment includes an erroneous fragment in the erroneous text and a correct fragment corresponding to the erroneous fragment in the correct text; The first training module is configured to take the first training data as input to a pre-trained candidate error description generation model to obtain first error description information corresponding to the first modified segment output by the candidate error description generation model, wherein the error description information is used to describe the error information in the error segment. The second training module is configured to optimize the candidate error description generation model based on the first error description information to obtain the optimized target error description generation model. The target error description generation model is used to output the target error description information corresponding to the target modified fragments in the target error text and the target correct text based on the input target error text and target correct text.
9. A text generation device, characterized in that, The device includes: The receiving module is configured to receive target text input by the user, the target text including target error text and target correct text; The output module is configured to take the target text as input to a pre-generated target error description generation model, and obtain target error description information corresponding to the target text output by the target error description generation model; the target interpretation generation model is generated by the model generation method according to any one of claims 1 to 6.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the program implements the steps of the method described in any one of claims 1-7.
11. An electronic device, characterized in that, include: A memory on which computer programs are stored; A processor for executing the computer program in the memory to implement the steps of the method according to any one of claims 1-7.
12. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the steps of the method described in any one of claims 1-7.