Corpus generation model training method and device
By using corpus generation prompt words and fine-tuning corpus in the training of the corpus generation model, and instructing the fine-tuning large language model to add reference marks, the problem of text hallucination in the corpus generation model is solved, and the accuracy and controllability of the generated corpus are improved.
Patent Information
- Application Number
- CN202510648905.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-19
- Publication Date
- 2025-09-19
AI Technical Summary
Existing corpus generation models are prone to text hallucination when generating corpus, which affects user experience and may lead to the spread of false information.
By obtaining multiple fine-tuning samples, including corpus generation prompt words and corresponding fine-tuning corpora, the instruction fine-tunes the large language model to add reference identifiers in the generated corpus, enhance the attention to reference information, and reduce text hallucination phenomena.
During the model training process, the model is guided to pay more attention to reference information, which can reduce the text hallucination phenomenon during the use of the corpus generation model and improve the accuracy and controllability of the generated corpus.
Smart Images

Figure CN120671736A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of natural language processing, and more specifically, to a method and device for training a corpus generation model. Background Art
[0002] With the continuous advancement of artificial intelligence technology, the field of natural language processing has made significant progress. Large language models (LLMs), as one of the key achievements in this field, have received widespread attention and application in recent years. Large language models are AI models based on deep learning technology, specifically designed for processing and generating natural language. Trained on large amounts of text data, they can understand and generate high-quality natural language content and complete a variety of complex language tasks, such as question-answering, translation, writing, and conversation. As a common application, large language models are often further trained as corpus generation models for generating corpora. However, corpus generation models trained using existing corpus generation model training methods are often prone to text hallucination during use. This is when generating corpora based on provided reference content. The model may generate seemingly reasonable corpora that are actually inconsistent with the reference content, lack logic, or are unverifiable. This not only affects user experience but can also lead to the spread of misinformation. Summary of the Invention
[0003] In view of this, an embodiment of the present invention provides a corpus generation model training method and device to guide the model to enhance its attention to reference information during the model training process, thereby reducing the text hallucination phenomenon that occurs during the use of the corpus generation model.
[0004] In a first aspect, an embodiment of the present invention is to provide a method for training a corpus generation model, the method comprising:
[0005] Acquiring multiple fine-tuning samples, wherein the fine-tuning samples include a corpus generation prompt word and a corresponding fine-tuning corpus, the corpus generation prompt word includes a reference information set and a corpus generation instruction, the reference information set includes multiple reference information, the corpus generation instruction is used to instruct a large language model to generate corresponding corpus based on the reference information set and add a reference identifier to the generated corpus, the reference identifier is used to mark the reference relationship between each sentence in the corpus and each reference information in the reference information set, and the reference identifier is added to the fine-tuning corpus;
[0006] Performing instruction fine-tuning on the large language model according to the multiple fine-tuning samples;
[0007] The corpus generation model is obtained by fine-tuning the large language model according to the instructions.
[0008] In a second aspect, an embodiment of the present invention is to provide a corpus generation model training device, the device comprising:
[0009] an acquisition unit configured to acquire a plurality of fine-tuning samples, wherein the fine-tuning samples include a corpus generation prompt word and a corresponding fine-tuning corpus, the corpus generation prompt word includes a reference information set and a corpus generation instruction, the reference information set includes a plurality of reference information, the corpus generation instruction is configured to instruct a large language model to generate corresponding corpus based on the reference information set and to add a reference identifier to the generated corpus, the reference identifier is configured to mark a reference relationship between each sentence in the corpus and each reference information in the reference information set, and the reference identifier is added to the fine-tuning corpus;
[0010] A fine-tuning unit, configured to perform instruction fine-tuning on the large language model according to the plurality of fine-tuning samples;
[0011] The training unit is used to obtain a corpus generation model based on the large language model fine-tuned according to the instructions.
[0012] In a third aspect, an embodiment of the present invention aims to provide a computer-readable storage medium storing computer program instructions, which implement the method described in the first aspect when executed by a processor.
[0013] In a fourth aspect, an embodiment of the present invention is directed to providing an electronic device, comprising:
[0014] a memory for storing one or more computer program instructions;
[0015] A processor, wherein the one or more computer program instructions are executed by the processor to implement the method as described in the first aspect.
[0016] In a fifth aspect, an embodiment of the present invention aims to provide a computer program product, which, when executed on a computer, enables the computer to execute the method as described in the first aspect.
[0017] The embodiment of the present invention obtains multiple fine-tuning samples, and fine-tunes the large language model according to the multiple fine-tuning samples, and then obtains a corpus generation model according to the large language model fine-tuned by the instructions. The fine-tuning samples include corpus generation prompt words and corresponding fine-tuning corpus, the corpus generation prompt words include a reference information set and a corpus generation instruction, the reference information set includes multiple reference information, the corpus generation instruction is used to instruct the large language model to generate corresponding corpus according to the reference information set and add a reference identifier to the generated corpus, the reference identifier is used to mark the reference relationship between the sentence in the corpus and the reference information in the reference information set, and the reference identifier is added to the fine-tuning corpus. Therefore, the embodiment of the present invention can guide the model to enhance its attention to reference information during the model training process, thereby reducing the text hallucination phenomenon that occurs during the use of the corpus generation model. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The above and other objects, features and advantages of the present invention will become more apparent through the following description of the embodiments of the present invention with reference to the accompanying drawings, in which:
[0019] Figure 1 This is a flowchart of a corpus generation model training method according to an embodiment of the present invention;
[0020] Figure 2 is a flow chart of a preference alignment method according to an embodiment of the present invention;
[0021] Figure 3 This is a flow chart of a method for generating preference corpus according to an embodiment of the present invention;
[0022] Figure 4 This is a flowchart of a method for generating positive preference corpus according to an embodiment of the present invention;
[0023] Figure 5 This is a flowchart of a node score updating method according to an embodiment of the present invention;
[0024] Figure 6 This is a flowchart of a method for generating negative preference corpus according to an embodiment of the present invention;
[0025] Figure 7 This is a flowchart of a node score updating method according to an embodiment of the present invention;
[0026] Figure 8 Schematic diagram of a corpus generation model training device according to an embodiment of the present invention;
[0027] Figure 9 FIG. 4 is a schematic diagram of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0028] The present application is described below based on the following embodiments, but the present application is not limited to these embodiments. In the detailed description of the present application below, certain specific details are described in detail. Those skilled in the art can fully understand the present application without the description of these details. To avoid obscuring the essence of the present application, well-known methods, processes, procedures, components, and circuits are not described in detail.
[0029] Furthermore, persons of ordinary skill in the art will appreciate that the figures provided herein are for illustration purposes only and are not necessarily drawn to scale.
[0030] Unless the context clearly requires otherwise, words like “include”, “comprising” and the like throughout this application should be interpreted as including rather than exclusive or exhaustive; that is, as meaning “including but not limited to”.
[0031] In the description of this application, it should be understood that the terms "first", "second", etc. are used for descriptive purposes only and should not be understood to indicate or imply relative importance. In addition, in the description of this application, unless otherwise specified, "plurality" means two or more.
[0032] Where the solutions described in this specification and in the examples involve the processing of personal information, such processing will be conducted with a legitimate basis (e.g., with the consent of the personal information subject or as necessary for the performance of a contract) and only within the prescribed or agreed scope. A user's refusal to process personal information other than that required for basic functions will not affect the user's use of these basic functions.
[0033] In the following description, the solution in the embodiment of the present invention will be applied in an e-commerce scenario, and will be used as an example to illustrate the application of the solution in the embodiment of the present invention in an e-commerce scenario, and to train a corpus generation model with the ability to generate e-commerce corpus. However, it should be understood that the solution in the embodiment of the present invention can also be applied in other scenarios where there is a demand for corpus generation, to train a corpus generation model with the ability to generate the corresponding type of corpus, and this application does not limit this. For example, the solution in the embodiment of the present invention can be applied in a live broadcast scenario to train a corpus generation model with the ability to generate live broadcast corpus. For another example, the solution in the embodiment of the present invention can be applied in a teaching scenario to train a corpus generation model with the ability to generate teaching corpus, which will not be elaborated on here.
[0034] At the same time, it is hoped that it will be noted that, in the embodiments of the present invention, e-commerce corpus can refer to text data of various types and formats related to many aspects of e-commerce scenarios, such as goods, users, transactions, and service provision, which may include product descriptions, advertising copy, product reviews, search queries, product FAQs, and customer service conversations, etc., and this application does not impose any restrictions on this. It should be understood that, in the embodiments of the present invention, the e-commerce corpus generated by the corpus generation model can be used as an important resource for training and optimizing natural language processing models, especially when building large language models, intelligent recommendation systems, search engines, and customer service robots in e-commerce scenarios.
[0035] Also, it is hoped that it will be noted that, in an embodiment of the present invention, the large language model used as a training object may be a model that has been pre-trained (i.e., a model that has been trained on a large-scale data set). A pre-trained large language model may have the ability to understand and execute general natural language instructions, but if it is directly applied to a corpus generation task in a specific scenario, it will be difficult to generate corpus that meets the requirements. The solution in an embodiment of the present invention may be to further perform specialized training on the large language model that has been pre-trained, so that it can be better applied to corpus generation tasks in specific scenarios.
[0036] Figure 1 This is a flow chart of the corpus generation model training method according to an embodiment of the present invention. It is to be noted that the execution subject of the corpus generation model training method can be a tablet computer, a laptop computer, a desktop computer or any other type of data processing device, and this application does not limit this. Figure 1 As shown, the corpus generation model training method may specifically include the following steps:
[0037] Step S100: Acquire multiple fine-tuning samples.
[0038] Specifically, the data processing device can obtain multiple fine-tuning samples. Among them, the fine-tuning samples can be understood as samples used to perform instruction fine-tuning training on the large language model. Instruction fine-tuning can be a fine-tuning method for a pre-trained model, which aims to train the large language model more finely through a specific instruction data set, so that it can better understand and execute natural language instructions given by humans. The core goal of instruction fine-tuning is to make the model more versatile and task adaptable, so that it can show higher accuracy and controllability when facing a variety of tasks.
[0039] In an embodiment of the present invention, as a method for setting a fine-tuning sample, the fine-tuning sample can be configured to include a corpus generation prompt and the corresponding fine-tuning corpus. The corpus generation prompt can be a piece of text input content that can be used to guide the large language model to generate corpus that meets the requirements. Furthermore, the corpus generation prompt can include a reference information set and corpus generation instructions.
[0040] The reference information set may include multiple reference information. The reference information can be understood as information provided by the model and used as a basis for corpus generation or background information to assist the model in corpus generation. It should be understood that when different types of e-commerce corpora need to be generated, the content of the reference information provided may be different. For example, when the e-commerce corpora to be generated are product promotion copy, the reference information content provided may be the product attributes of the product. When the e-commerce corpora to be generated are product reviews, the reference information content provided may be product evaluations on various aspects of the product. When the e-commerce corpora to be generated are customer service conversations, the reference information content provided may be question-and-answer pairs about frequently asked product questions.
[0041] The corpus generation instruction may be an instruction for specifying the tasks and requirements that the large language model needs to complete. In an embodiment of the present invention, the corpus generation instruction may specifically be used to instruct the large language model to generate corresponding corpus based on the reference information set and add a reference identifier to the generated corpus. The reference identifier may be used to mark the reference relationship between a sentence in the corpus (a sentence may refer to a complete sentence, a complex sentence composed of multiple clauses, or a paragraph) and the reference information in the reference information set.
[0042] Furthermore, in embodiments of the present invention, the specific format of the reference identifier and the manner in which it is added to the corpus can be configured by relevant personnel based on actual needs, and this application does not impose any restrictions on this. Optionally, as an implementation method, the reference identifier can be configured to include a sentence-defining portion and a reference-marking portion. The sentence-defining portion can be used to define the sentence in the corpus that references the reference information. Specifically, the sentence-defining portion can include two qualifiers, which can be set at the beginning and end of the sentence that references the reference information, respectively, to confine the sentence between the two qualifiers. It should be understood that the qualifiers can be words or mathematical symbols, and their specific design can be determined by relevant personnel based on actual needs. Furthermore, the qualifiers at the beginning and end of the sentence can be the same or different, and this application does not impose any restrictions on this. The reference-marking portion can be used to annotate the source of the reference information cited in the sentence. Specifically, the reference-marking portion can be the index number of the reference information cited in the sentence, which can be set within either of the two qualifiers, or within both qualifiers. It should be understood that to facilitate the determination of the index number of each reference information, each reference information in the reference information set can be pre-assigned a corresponding independent index number. It should be understood that in order to make the large language model clear about the specific format of the reference identifier and the method of adding it to the corpus, the corpus generation instruction can also be used to indicate the specific format of the reference identifier and the method of adding it to the corpus.
[0043] The fine-tuning corpus can be high-quality real corpus collected during the actual operation of the e-commerce platform, which can be used as a reference standard for corpus generation for the large language model to fine-tune its own parameters. It should be understood that in order to ensure that the fine-tuning corpus can be used as a reference standard for corpus generation, the fine-tuning corpus needs to meet the various requirements expressed in the corpus generation instructions, that is, the fine-tuning corpus needs to be added with a reference identifier. Optionally, the fine-tuning corpus can be obtained by collecting data from relevant web pages and adding a reference identifier to the collected data. The collection and reference identifier adding operations can be implemented manually or by the data collection and annotation model obtained through training, and this application does not limit this.
[0044] Indicatively, when the e-commerce corpus to be generated is a product promotion copy, the reference information set in the corpus generation prompt can be: [1. Brand: ABC, 2. Category: Trolley Case, 3. Model: Walk Trolley Case, 4. Color: Lake Blue, 5. Interior Design: Built-in Storage System; Layered Arrangement; Multi-layer Storage]. The corpus generation instruction in the corpus generation prompt can be: [Please write an e-commerce text based on the input reference information set. Several requirements are as follows: 1. The form of the e-commerce text needs to be consistent with the article type of the product usage experience; 2. The target group of the e-commerce text is users who pursue fashionable appearance and travel quality; 3. The reference information set gives several product attributes, and the name and attribute value of each product attribute are marked with an index number; 4. Ensure that all text content in the e-commerce text related to product information must come from the reference information set, and cannot produce additional associations and illusions. In order to implement this principle, if the generated sentence is related to product information during the process, it is necessary to reference the related product attributes at the same time. The specific format is "<Reference[{ID1};…]> Reference information statement". Among them, [{ID1};...] indicates which product attributes in the reference information set are referenced, and the specific index number needs to be filled in. If there are multiple related items, the numbers are separated by ",".
[0045] The fine-tuning corpus corresponding to the prompt word generated by this corpus can be:<Reference[1,3]> What's a good suitcase brand? How about the ABC Wander trolley case? Many people love to travel. Today's fast-paced lives demand a change of scenery to relax and unwind, and a suitcase is essential for travel. Just like the suitcase I'm trying out this time, its design theme is "Wander," accompanying you as you explore the meaning of travel.<Reference[4]> The first impression of this luggage is its stylish design. Breaking away from the usual color schemes, it uses a striking lake blue as the base, accented with contrasting 3D patterns, creating a modernist aesthetic. While the colors and patterns are bold, the overall design is minimalist, creating a sense of simplicity rather than fuss.<Reference[5]> The box is lightweight and well-designed, perfectly storing clothes and gliding smoothly and comfortably. The store also has several other models of this design, each with its own unique theme, using different colors and patterns, creating a beautiful effect.
[0046] It should be understood that the corpus generation prompt words and the corresponding fine-tuning corpus given above are only for illustration. In actual application, the specific contents of the corpus generation prompt words and the corresponding fine-tuning corpus are not limited to these.
[0047] Step S200: fine-tune the large language model based on the multiple fine-tuning samples.
[0048] Specifically, after obtaining a plurality of fine-tuning samples, the data processing device may perform instruction fine-tuning on the large language model according to the plurality of fine-tuning samples.
[0049] It should be understood that in step S200, the model training method adopted by the data processing device may be supervised fine-tuning (SFT). Supervised fine-tuning is a deep learning technique that can be used to further optimize model performance based on a pre-trained model using a task-specific dataset. In embodiments of the present invention, supervised fine-tuning can use multiple fine-tuning samples as a dataset to further optimize the performance of a large language model.
[0050] Step S300: Obtain a corpus generation model based on the large language model fine-tuned according to the instructions.
[0051] Specifically, after obtaining the fine-tuned large language model, the data processing device can continue to execute subsequent corresponding training content on the fine-tuned large language model to obtain a corpus generation model based on the fine-tuned large language model. It should be understood that in some embodiments, the data processing device can also directly use the fine-tuned large language model as the corpus generation model.
[0052] It is hoped that it will be explained that in the fine-tuning sample settings of the related art, the corpus generation instructions are usually only used to instruct the large language model to generate the corresponding corpus based on the reference information set, and no reference identifiers are added to the fine-tuning corpus. The large language model trained by the related art often suffers from more serious text hallucinations during use. Compared with the related art, the embodiment of the present invention adjusts the content and structure of the fine-tuning samples (that is, the corpus generation instructions are also used to instruct the large language model to add reference identifiers to the generated corpus, and to add reference identifiers to the fine-tuning corpus). This can guide the model to enhance its attention to reference information during the model training process, thereby reducing the text hallucination phenomenon that occurs during the use of the corpus generation model.
[0053] Optionally, in step S300, the subsequent training content performed by the data processing device on the large language model after fine-tuning the instructions may specifically include preference alignment. Preference alignment is a technology or method designed to align the output of an artificial intelligence model with human preferences. Its core goal is to make the content generated by the model more in line with human values, expectations and needs through training and optimization, thereby improving the actual application effect of the model and user experience. In an embodiment of the present invention, the data processing device can use the reduction of text synthesis hallucinations as a human preference to perform preference alignment on the large language model, so that the large language model is more inclined to generate corpus without text hallucination phenomena.
[0054] Figure 2FIG. 1 is a flow chart of a preference alignment method according to an embodiment of the present invention. It should be understood that by executing Figure 2 In the preference alignment method shown, the data processing device can use the reduction of text synthesis hallucination as a human preference to perform preference alignment on the large language model, thereby obtaining a corpus generation model that meets the requirements. Figure 2 As shown, the preference alignment method may specifically include the following steps:
[0055] Step S310: Obtain a plurality of the corpus to generate prompt words.
[0056] Specifically, the data processing device may obtain multiple corpus generation prompt words. It should be understood that the corpus generation prompt words obtained by the data processing device may be the same as the corpus generation prompt words in the fine-tuning sample, and may include a reference information set and corpus generation instructions. The reference information set may include multiple reference information. The corpus generation instructions may be used to instruct the large language model to generate corresponding corpus based on the reference information set and add a reference identifier to the generated corpus.
[0057] Optionally, in step S310, the corpus generation prompt words obtained by the data processing device may be corpus generation prompt words in the fine-tuning sample used in the instruction fine-tuning stage, or may be corpus generation prompt words that have not been used, and this application does not impose any restrictions on this.
[0058] Step S320: For each of the corpus generation prompt words, generate a preferred corpus according to the corpus generation prompt words to obtain a preferred sample.
[0059] Specifically, after obtaining multiple corpus generation prompt words, for each corpus generation prompt word, the data processing device can generate a preferred corpus based on the corpus generation prompt word to obtain a preferred sample. The preferred sample can be understood as a sample used for preference alignment training of the large language model.
[0060] In an embodiment of the present invention, as a method for setting preference samples, the preference samples may be configured to include a corpus generation prompt word, a corresponding positive preference corpus, and a corresponding negative preference corpus. The positive preference corpus may be a corpus in which each sentence is consistent with the reference information in the reference information set. The negative preference corpus may be a corpus in which sentences are inconsistent with the reference information in the reference information set.
[0061] Figure 3 FIG. 1 is a flow chart of a method for generating preference corpus according to an embodiment of the present invention. It should be understood that by executing Figure 3 The preferred corpus generation method shown in FIG, the data processing device can generate positive preference corpus and negative preference corpus corresponding to the corpus generation prompt word. Figure 3As shown, the preferred corpus generation method may specifically include the following steps:
[0062] Step S321: generate positive preference corpus according to the corpus generation prompt word to obtain positive preference corpus corresponding to the corpus generation prompt word.
[0063] Specifically, the data processing device may generate positive preference corpus according to the corpus generation prompt word to obtain positive preference corpus corresponding to the corpus generation prompt word.
[0064] Figure 4 FIG is a flow chart of a method for generating positive preference corpus according to an embodiment of the present invention. It should be understood that by executing Figure 4 The positive preference corpus generation method shown in FIG. 1 can generate positive preference corpus corresponding to the corpus generation prompt word, that is, to implement the above step S321. Figure 4 As shown, the positive preference corpus generation method may specifically include the following steps:
[0065] Step S3211: Generate an initial node based on the prompt words generated by the corpus.
[0066] Specifically, the data processing device can generate an initial node based on the corpus generation prompt word. It should be understood that in the embodiment of the present invention, a node (that is, including the initial node and the subnodes mentioned later) can be understood as a sentence with a corresponding reference identifier. The sentence can be a complete sentence or a complex sentence or paragraph composed of multiple clauses. Schematically, the node can be: [<Reference[1,3]> What's a good suitcase brand? How about the ABC Wander luggage?
[0067] Optionally, in an embodiment of the present invention, the initial node and subsequent sub-nodes can be generated by inputting the corpus generation prompt word into the corresponding large language model. Further optionally, here, the large language model used to generate the node generation can be the large language model that is currently undergoing preference alignment training, or it can be another large language model, and this application does not limit this. Further optionally, the large language model can implement the generation of a single node by inputting the already generated node, corpus generation prompt word and additional natural language instructions (the instructions are used to instruct the large language model to generate a single node based on the existing content) into the large language model. It should be understood that in some embodiments, the implementation of the large language model to generate a single node can also be implemented by fine-tuning the code layer of the large language model, and this application does not limit this.
[0068] Step S3212: Select the current node from the incompletely expanded nodes.
[0069] Specifically, the data processing device may select the current node from the incompletely expanded nodes, wherein the incompletely expanded node may refer to a node where statement continuation is possible.
[0070] In an embodiment of the present invention, each node may have a corresponding node score (when a node is generated, the node may be assigned an initial node score, and in subsequent iterations, the specific value of the node score may be continuously updated according to the iteration situation). Optionally, in step S3212, the data processing device may select a current node from a plurality of incompletely expanded nodes based on the node scores of the current incompletely expanded nodes and a confidence interval upper bound algorithm (e.g., UCB1). The confidence interval upper bound algorithm is an algorithm for determining a better choice among multiple choices. In an embodiment of the present invention, the confidence interval upper bound algorithm may be used to select a better (i.e., the one with the highest value) incompletely expanded node as the current node from a plurality of incompletely expanded nodes based on relevant parameters such as the node score of each incompletely expanded node itself, the number of times each incompletely expanded node itself has been selected, and the number of times the corresponding previous node (also understood as the parent node) of each incompletely expanded node has been selected.
[0071] Optionally, the number of current nodes selected by the data processing device during each iteration can be set by relevant personnel based on actual needs, and this application does not impose any restrictions on this. Illustratively, the number of current nodes selected by the data processing device during each iteration can be three. It should be understood that when selecting the current node, if the number of incompletely expanded nodes is less than the required number of current nodes to be selected, the data processing device can directly select all incompletely expanded nodes as the current nodes.
[0072] Step S3213: For each current node, update multiple child nodes corresponding to the current node.
[0073] Specifically, after selecting the current node, the data processing device may update the multiple child nodes corresponding to the current node for each current node. It should be understood that, as described above, the child nodes may be updated by the data processing device inputting the current node, the previous node corresponding to the current node, a corpus generation prompt, and an additional natural language instruction (the instruction is used to instruct the large language model to generate a single node based on the existing content) into the large language model, so as to generate the multiple child nodes corresponding to the current node through the large language model.
[0074] Optionally, the number of child nodes updated by the data processing device for each current node in each round of iteration can be set by relevant personnel according to actual needs, and this application does not limit this. Schematically, the number of child nodes updated by the data processing device for each current node in each round of iteration can be 5.
[0075] Step S3214: Generate corresponding corpus for each of the sub-nodes.
[0076] Specifically, after generating multiple child nodes corresponding to each current node relationship, the data processing device can generate corresponding corpora for each child node. It should be understood that the corresponding corpora for the child nodes generated by the data processing device here can be specifically understood as the complete corpus obtained by continuing the child node and the corresponding previous node of the child node. The complete corpus can include the child node, the previous node of the child node, and the continued content (the continued content is also generated according to the corpus generation prompt word).
[0077] Optionally, in step S3214, for each child node, the corresponding corpus of the child node may be generated by a data processing device inputting the child node, the corresponding previous node of the child node, and the corpus generation prompt word into a large language model to generate the corresponding corpus of the child node through the large language model.
[0078] It should be understood that when generating node sub-nodes, in order to ensure the smooth flow of sentences between the current node and the sub-nodes, additional cohesive sentences without corresponding reference identifiers may be generated between the current node and the sub-nodes (the cohesive sentences are used to connect sentences between the current node and the sub-nodes). In an embodiment of the present invention, these cohesive sentences can be stored separately and used to generate corresponding corpora for the sub-nodes or to generate the final preferred corpus. That is, in addition to containing multiple nodes, the generated corresponding corpus or the final preferred corpus can also contain cohesive sentences that may exist between multiple nodes.
[0079] Step S3215: For each of the child nodes, update the node scores of the child node and the corresponding previous node according to the fluency of each sentence in the corresponding corpus and the consistency between each sentence in the corresponding corpus and the reference information in the reference information set.
[0080] Specifically, after generating the corresponding corpus for each child node, for each child node, the data processing device can update the node score of the child node and the corresponding previous node of the child node based on the fluency of each sentence in the corresponding corpus of the child node and the degree of consistency between each sentence in the corresponding corpus and the reference information in the reference information set.
[0081] Figure 5 FIG. 1 is a flow chart of a node score updating method according to an embodiment of the present invention. It should be understood that by executing Figure 5 The node score updating method shown in FIG. 1 can update the node score according to the current iteration situation, that is, implement the above step S3215. Figure 5 As shown, the node score updating method may specifically include the following steps:
[0082] Step S32151: Determine the fluency score of the corresponding corpus based on the fluency of each sentence in the corresponding corpus.
[0083] Specifically, the data processing device may determine the fluency score of the corresponding corpus according to the fluency of each sentence in the corresponding corpus.
[0084] Optionally, in step S32151, the fluency score can be determined by a data processing device inputting the corresponding corpus into a fluency score evaluation model to determine the fluency score of the corresponding corpus using the fluency score evaluation model. The fluency score evaluation model can be any type of score evaluation model, and this application does not limit this. Furthermore, as a training method, the fluency score evaluation model can be obtained by training corpus samples labeled with fluency scores, and this application does not limit the method for obtaining the model.
[0085] Step S32152: Determine the consistency score of the corresponding corpus according to the consistency between each sentence in the corresponding corpus and the reference information in the reference information set.
[0086] Specifically, the data processing device may determine a consistency score for the corresponding corpus according to a consistency degree between each sentence in the corresponding corpus and the reference information in the reference information set.
[0087] Optionally, similar to the method for determining the fluency score, in step S32152, the consistency score can be determined by the data processing device inputting the corresponding corpus and the reference information set into the consistency score evaluation model to determine the consistency score of the corresponding corpus through the consistency score evaluation model. The consistency score evaluation model can be any type of score evaluation model, and this application does not limit this. In addition, as a training method, the consistency score evaluation model can be obtained by training corpus samples annotated with consistency scores and the corresponding reference information set. This application does not limit the method for obtaining the model.
[0088] Step S32153: Update the node scores of the child node and the corresponding previous node according to the fluency score and the consistency score.
[0089] Specifically, after determining the fluency score and the consistency score of the corresponding corpus, the data processing device may update the node scores of the child node and the corresponding previous node of the child node according to the fluency score and the consistency score.
[0090] Optionally, in step S32153, as an implementation method, the data processing device may perform a weighted sum of the fluency score and the consistency score, and adjust the node scores of the child node and its corresponding predecessor node based on the weighted sum result. The specific score adjustment rules may be designed by relevant personnel based on actual circumstances, and this application does not impose any restrictions thereon.
[0091] Step S3216: Determine whether the preset iteration end condition is met.
[0092] Specifically, after updating the node scores, the data processing device can determine whether a preset iteration termination condition is met. If so, the data processing device can terminate the iteration and determine a positive preference corpus based on the node scores of each node. If not, the data processing device can return and re-execute step S3212 to begin the next round of iteration. The positive preference corpus can be the node string with the highest sum of node scores when the preset iteration termination condition is met.
[0093] It should be understood that in the embodiment of the present invention, for each fully expanded node, the fully expanded node, the initial node, and at least one intermediate node connected therebetween can constitute a node string. Each node string can be regarded as a complete corpus.
[0094] Optionally, in an embodiment of the present invention, the preset iteration end condition can be designed by relevant personnel according to actual conditions, and this application does not limit this. Schematically, as a setting method, the preset iteration end condition is satisfied and can be set to include the current iteration round number meeting the preset iteration round number threshold, there is no incompletely expanded node at present, there is a fully expanded node at present and the sum of the node scores of the fully expanded node and the corresponding previous node is greater than or equal to the first preset score threshold, and / or there is a fully expanded node at present and the sum of the node scores of each incompletely expanded node and the corresponding previous node is less than the second preset score threshold. Among them, the specific values of the preset iteration round number threshold, the first preset score threshold and the second preset score threshold can be set and adjusted by relevant personnel according to actual needs, and this application does not limit this.
[0095] Step S322: generating negative preference corpus according to the corpus generation prompt word to obtain negative preference corpus corresponding to the corpus generation prompt word.
[0096] Specifically, the data processing device may generate negative preference corpus according to the corpus generation prompt word to obtain negative preference corpus corresponding to the corpus generation prompt word.
[0097] Figure 6 FIG. 1 is a flow chart of a method for generating negative preference corpus according to an embodiment of the present invention. It should be understood that by executing Figure 6 The negative preference corpus generation method shown in FIG. 1 can generate negative preference corpus corresponding to the corpus generation prompt word, that is, implement the above step S322. Figure 6 As shown, the negative preference corpus generation method may specifically include the following steps:
[0098] Step S3221: Generate an initial node based on the prompt words generated by the corpus.
[0099] Specifically, the data processing device can generate the initial node according to the corpus generation prompt word. It should be understood that the implementation of step S3221 is the same as that of step S3211. The details can be referred to the above description and will not be repeated here.
[0100] Step S3222: Select the current node from the nodes that are not fully expanded.
[0101] Specifically, the data processing device may select the current node from the incompletely expanded nodes. The incompletely expanded node may refer to a node where statement continuation is possible. It should be understood that the implementation of step S3222 is the same as that of step S3212. The details can be found in the above description and will not be repeated here.
[0102] Step S3223: For each current node, update multiple child nodes corresponding to the current node.
[0103] Specifically, after selecting the current node, for each current node, the data processing device may update the multiple child nodes corresponding to the current node. It should be understood that, compared to step S3213, in step S3223, the data processing device may introduce random perturbations (e.g., replacing related words with antonyms or similar words of the word) when updating the multiple child nodes corresponding to any current node, so that the multiple child nodes updated for the current node include at least one perturbed node. The perturbed node may be inconsistent with the reference information in the reference information set.
[0104] For example, let's assume the reference information set is: [1. Brand: ABC, 2. Category: Luggage, 3. Model: Walk-in Luggage, 4. Color: Lake Blue, 5. Interior Design: Built-in Storage System; Tiered Arrangement; Multi-layer Storage]. The current node is: [<Reference[1,3]> What brand of suitcase is good? How about the ABC stroller suitcase? After introducing random perturbations, the data processing device can update the content of the current node as follows:<Reference[4]> Especially the fiery red design, which gives people a refreshing feeling. It should be understood that the perturbation nodes given above are only for illustration. In actual application, the generated perturbation node content is not limited to this.
[0105] Step S3224: Generate corresponding corpus for each of the sub-nodes.
[0106] Specifically, after generating multiple child nodes corresponding to each current node relationship, the data processing device can generate corresponding corpus for each child node. It should be understood that the implementation of step S3224 is the same as that of step S3214. For details, please refer to the above description and will not be repeated here.
[0107] Step S3225: For each of the child nodes, update the node scores of the child node and the corresponding previous node according to the fluency of each sentence in the corresponding corpus and the consistency between other sentences in the corresponding corpus except the disturbance node and the reference information in the reference information set.
[0108] Specifically, after generating the corresponding corpus for each child node, for each child node, the data processing device can update the node score of the child node and the corresponding previous node based on the fluency of each sentence in the corresponding corpus and the degree of consistency between other sentences in the corresponding corpus except the disturbance node and the reference information in the reference information set.
[0109] Figure 7 FIG. 1 is a flow chart of a node score updating method according to an embodiment of the present invention. It should be understood that by executing Figure 7 The node score updating method shown in FIG. 1 can update the node score according to the current iteration situation, that is, implement the above step S3225. Figure 7 As shown, the node score updating method may specifically include the following steps:
[0110] Step S32251: Determine the fluency score of the corresponding corpus according to the fluency of each sentence in the corresponding corpus.
[0111] Specifically, the data processing device can determine the fluency score of the corresponding corpus according to the fluency of each sentence in the corresponding corpus. It should be understood that the implementation of step S32251 is the same as that of step S32151. The details can be referred to the above description and will not be repeated here.
[0112] Step S32252: Remove at least one disturbance node in the corresponding corpus.
[0113] Specifically, before determining the consistency score of the corresponding corpus, the data processing device may first remove at least one disturbance node in the corresponding corpus.
[0114] Step S32253: Determine a consistency score of the corresponding corpus according to the consistency between each sentence in the corresponding corpus after the disturbance node is removed and the reference information in the reference information set.
[0115] Specifically, after removing at least one perturbation node from the corresponding corpus, the data processing device may determine a consistency score for the corresponding corpus based on the degree of consistency between each sentence in the corresponding corpus after the perturbation node is removed and the reference information in the reference information set. It should be understood that the implementation of step S32253 is the same as that of step S32152. For details, reference may be made to the above description and will not be repeated here.
[0116] Step S32254: Update the node scores of the child node and the corresponding previous node according to the consistency score and the fluency score.
[0117] Specifically, after determining the fluency score and consistency score of the corresponding corpus, the data processing device can update the node scores of the child node and the corresponding predecessor node of the child node based on the fluency score and consistency score. It should be understood that the implementation method of step S32254 is the same as the implementation method of step S32153. The details can be referred to the above description and will not be repeated here. At the same time, it is hoped that the consistency score obtained here can be specifically understood as the consistency score determined after ignoring the impact of the perturbation node in the corresponding corpus.
[0118] Step S3226: Determine whether the preset iteration end condition is met.
[0119] Specifically, after updating the node scores, the data processing device can determine whether a preset iteration termination condition is met. If so, the data processing device can terminate the iteration and determine a negative preference corpus based on the node scores of each node. If not, the data processing device can return and re-execute step S3222 to begin the next iteration. The negative preference corpus can be the node string with the highest sum of node scores when the preset iteration termination condition is met.
[0120] Optionally, in an embodiment of the present invention, in order to ensure that the negative preference corpus ultimately selected is a node string containing a perturbation node, the data processing device may, during the node update process, selectively select the perturbation node addition object so that each node string includes the perturbation node. Alternatively, the data processing device may also, during the negative preference corpus selection process, directly select the node string with the highest sum of node scores and containing the perturbation node as the negative preference corpus. Alternatively, the data processing device may also use other methods to ensure that the negative preference corpus ultimately selected is a node string containing a perturbation node, and this application does not impose any restrictions on this.
[0121] Step S330: performing preference alignment on the large language model after instruction fine-tuning according to the plurality of preference samples to obtain the corpus generation model.
[0122] Specifically, after obtaining multiple preference samples, the data processing device can perform preference alignment on the large language model that has been fine-tuned based on the multiple preference samples to obtain a corpus generation model. The corpus generation model is the large language model that has been trained with preference alignment.
[0123] It should be understood that in step S330, the model training method adopted by the data processing device may be direct preference optimization (DPO). Direct preference optimization is a model training method for aligning preferences for large language models, which aims to make the model output more consistent with human preferences. In an embodiment of the present invention, direct preference optimization can use positive preference corpora in multiple preference samples as answers preferred by humans, and negative preference corpora as answers not preferred by humans, thereby achieving training of a large language model. The large language model that has completed training will be more inclined to generate answers preferred by humans, that is, to generate corpora in which each sentence is consistent with the reference information in the reference information set. Thus, the embodiment of the present invention can alleviate the text hallucination phenomenon that occurs during the use of the corpus generation model.
[0124] The embodiment of the present invention obtains multiple fine-tuning samples, and fine-tunes the large language model according to the multiple fine-tuning samples, and then obtains a corpus generation model according to the large language model fine-tuned by the instructions. The fine-tuning samples include corpus generation prompt words and corresponding fine-tuning corpus, the corpus generation prompt words include a reference information set and a corpus generation instruction, the reference information set includes multiple reference information, the corpus generation instruction is used to instruct the large language model to generate corresponding corpus according to the reference information set and add a reference identifier to the generated corpus, the reference identifier is used to mark the reference relationship between the sentence in the corpus and the reference information in the reference information set, and the reference identifier is added to the fine-tuning corpus. Therefore, the embodiment of the present invention can guide the model to enhance its attention to reference information during the model training process, thereby reducing the text hallucination phenomenon that occurs during the use of the corpus generation model.
[0125] Figure 8 Schematic diagram of the corpus generation model training device according to an embodiment of the present invention. Figure 8 As shown, the corpus generation model training device of the embodiment of the present invention includes an acquisition unit 81, a fine-tuning unit 82 and a training unit 83.
[0126] Specifically, the acquisition unit 81 is used to acquire multiple fine-tuning samples, wherein the fine-tuning samples include a corpus generation prompt word and a corresponding fine-tuning corpus, the corpus generation prompt word includes a reference information set and a corpus generation instruction, the reference information set includes multiple reference information, the corpus generation instruction is used to instruct the large language model to generate corresponding corpus based on the reference information set and add a reference identifier to the generated corpus, the reference identifier is used to mark the reference relationship between each sentence in the corpus and each reference information in the reference information set, and the reference identifier is added to the fine-tuning corpus;
[0127] The fine-tuning unit 82 is used to perform instruction fine-tuning on the large language model according to the multiple fine-tuning samples;
[0128] The training unit 83 is used to obtain a corpus generation model based on the large language model fine-tuned by the instruction.
[0129] The embodiment of the present invention obtains multiple fine-tuning samples, and fine-tunes the large language model according to the multiple fine-tuning samples, and then obtains a corpus generation model according to the large language model fine-tuned by the instructions. The fine-tuning samples include corpus generation prompt words and corresponding fine-tuning corpus, the corpus generation prompt words include a reference information set and a corpus generation instruction, the reference information set includes multiple reference information, the corpus generation instruction is used to instruct the large language model to generate corresponding corpus according to the reference information set and add a reference identifier to the generated corpus, the reference identifier is used to mark the reference relationship between the sentence in the corpus and the reference information in the reference information set, and the reference identifier is added to the fine-tuning corpus. Therefore, the embodiment of the present invention can guide the model to enhance its attention to reference information during the model training process, thereby reducing the text hallucination phenomenon that occurs during the use of the corpus generation model.
[0130] Figure 9 Schematic diagram of an electronic device according to an embodiment of the present invention. It should be understood that Figure 9 The electronic device shown can specifically be a data processing device for performing model training in the above embodiment. Figure 9 As shown, the electronic device includes: at least one processor 91; a memory 92 communicatively connected to the at least one processor 91; and a communication component 93 communicatively connected to the scanning device, the communication component 93 receiving and sending data under the control of the processor 91; wherein the memory 92 stores instructions that can be executed by the at least one processor 91, and the instructions are executed by the at least one processor 91 to implement the above-mentioned corpus generation model training method.
[0131] Specifically, the electronic device includes: one or more processors 91 and a memory 92, Figure 9 A processor 91 is used as an example. The processor 91 and the memory 92 may be connected via a bus or other means. Figure 9In the example, a bus connection is used. Memory 92, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer executable programs, and modules. Processor 91 executes the non-volatile software programs, instructions, and modules stored in memory 92 to perform various functional applications and data processing of the device, thereby implementing the above-mentioned corpus generation model training method.
[0132] The memory 92 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and application programs required for at least one function; the data storage area may store a list of options, etc. In addition, the memory 92 may include a high-speed random access memory and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some embodiments, the memory 92 may optionally include a memory remotely located relative to the processor 91, and these remote memories may be connected to an external device via a network. Examples of the aforementioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0133] One or more modules are stored in the memory 92 and, when executed by one or more processors 91 , perform the corpus generation model training method in any of the above method embodiments.
[0134] The above-mentioned product can execute the method provided in the embodiment of this application, and has the functional modules and beneficial effects corresponding to the execution method. For technical details not fully described in this embodiment, please refer to the method provided in the embodiment of this application.
[0135] The embodiment of the present invention obtains multiple fine-tuning samples, and fine-tunes the large language model according to the multiple fine-tuning samples, and then obtains a corpus generation model according to the large language model fine-tuned by the instructions. The fine-tuning samples include corpus generation prompt words and corresponding fine-tuning corpus, the corpus generation prompt words include a reference information set and a corpus generation instruction, the reference information set includes multiple reference information, the corpus generation instruction is used to instruct the large language model to generate corresponding corpus according to the reference information set and add a reference identifier to the generated corpus, the reference identifier is used to mark the reference relationship between the sentence in the corpus and the reference information in the reference information set, and the reference identifier is added to the fine-tuning corpus. Therefore, the embodiment of the present invention can guide the model to enhance its attention to reference information during the model training process, thereby reducing the text hallucination phenomenon that occurs during the use of the corpus generation model.
[0136] Another embodiment of the present invention relates to a non-volatile storage medium for storing a computer-readable program, wherein the computer-readable program is used to enable a computer to execute part or all of the above method embodiments.
[0137] That is, those skilled in the art will understand that all or part of the steps in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a program, which is stored in a storage medium and includes a number of instructions for causing a device (which may be a single-chip microcomputer, chip, etc.) or a processor to execute all or part of the steps of the methods described in the embodiments of the present application. The aforementioned storage medium includes: a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc., various media that can store program code.
[0138] The foregoing is merely a preferred embodiment of the present application and is not intended to limit the present application. Persons skilled in the art will readily appreciate that various modifications and variations are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present application are intended to be within the scope of protection of the present application.
Claims
1. A corpus generation model training method, characterized in that: The method comprises: Acquire multiple fine-tuning samples, wherein the fine-tuning samples include a corpus generation prompt word and a corresponding fine-tuning corpus, the corpus generation prompt word includes a reference information set and a corpus generation instruction, the reference information set includes multiple reference information, the corpus generation instruction is used to instruct a large language model to generate corresponding corpus based on the reference information set and add a reference identifier to the generated corpus, the reference identifier is used to mark the reference relationship between a sentence in the corpus and the reference information in the reference information set, and the reference identifier is added to the fine-tuning corpus; Performing instruction fine-tuning on the large language model according to the multiple fine-tuning samples; The corpus generation model is obtained by fine-tuning the large language model according to the instructions.
2. The method according to claim 1, characterized in that The acquisition of a corpus generation model based on the large language model fine-tuned by the instruction includes: Obtaining a plurality of the corpus to generate prompt words; For each of the corpus generation prompt words, generating a preference corpus according to the corpus generation prompt word to obtain a preference sample, wherein the preference sample includes the corpus generation prompt word, a corresponding positive preference corpus, and a corresponding negative preference corpus, wherein the positive preference corpus is a corpus in which each sentence is consistent with the reference information in the reference information set, and the negative preference corpus is a corpus in which there are sentences that are inconsistent with the reference information in the reference information set; The large language model after instruction fine-tuning is subjected to preference alignment according to the plurality of preference samples to obtain the corpus generation model.
3. The method according to claim 2, characterized in that The generating of the preference corpus according to the corpus generating prompt words includes: generating positive preference corpus according to the corpus generation prompt word to obtain positive preference corpus corresponding to the corpus generation prompt word; Negative preference corpus generation is performed according to the corpus generation prompt word to obtain negative preference corpus corresponding to the corpus generation prompt word.
4. The method according to claim 3, characterized in that The generating of positive preference corpus according to the corpus generation prompt word to obtain positive preference corpus corresponding to the corpus generation prompt word includes: Generate an initial node based on the prompt word generated by the corpus, and iteratively perform the following steps until the positive preference corpus is obtained: Select the current node among the nodes that are not fully expanded; For each current node, updating multiple child nodes corresponding to the current node; Generate corresponding corpus for each of the sub-nodes; For each of the child nodes, the node scores of the child node and the corresponding previous node are updated according to the fluency of each sentence in the corresponding corpus and the consistency between each sentence in the corresponding corpus and the reference information in the reference information set, wherein the node is a sentence with a corresponding reference identifier, and the positive preference corpus is a node string with the highest sum of node scores when the preset iteration end condition is met.
5. The method according to claim 4, characterized in that The updating of the node scores of the child node and the corresponding previous node according to the fluency of each sentence in the corresponding corpus and the consistency between each sentence in the corresponding corpus and the reference information in the reference information set includes: Determine the fluency score of the corresponding corpus according to the fluency of each sentence in the corresponding corpus; Determining a consistency score for the corresponding corpus based on the consistency between each sentence in the corresponding corpus and the reference information in the reference information set; The node scores of the child node and the corresponding previous node are updated according to the fluency score and the consistency score.
6. The method according to claim 3, characterized in that The generating of negative preference corpus according to the corpus generation prompt word to obtain negative preference corpus corresponding to the corpus generation prompt word includes: Generate an initial node based on the prompt word generated by the corpus, and iteratively perform the following steps until the negative preference corpus is obtained: Select the current node among the nodes that are not fully expanded; For each of the current nodes, updating a plurality of child nodes corresponding to the current node, wherein the plurality of child nodes include at least one disturbance node, and the disturbance node is inconsistent with reference information in the reference information set; Generate corresponding corpus for each of the sub-nodes; For each of the child nodes, the node scores of the child node and the corresponding previous node are updated according to the fluency of each sentence in the corresponding corpus and the consistency between the other sentences in the corresponding corpus except the disturbance node and the reference information in the reference information set, wherein the node is a sentence with a corresponding reference identifier, and the negative preference corpus is a node string with the highest sum of node scores when the preset iteration end condition is met.
7. The method according to claim 6, characterized in that The updating of the node scores of the child node and the corresponding previous node according to the fluency of each sentence in the corresponding corpus and the consistency between other sentences in the corresponding corpus except the disturbance node and the reference information in the reference information set includes: Determine the fluency score of the corresponding corpus according to the fluency of each sentence in the corresponding corpus; Removing at least one of the disturbance nodes in the corresponding corpus; Determining a consistency score of the corresponding corpus according to the consistency between each sentence in the corresponding corpus after removing the disturbance node and the reference information in the reference information set; The node scores of the child node and the corresponding previous node are updated according to the consistency score and the fluency score.
8. The method according to claim 4 or 6, characterized in that The selecting the current node from the incompletely expanded nodes includes: The current node is selected from the incompletely expanded nodes according to the node score and the confidence interval upper bound algorithm of each incompletely expanded node.
9. The method according to claim 4 or 6, characterized in that The preset iteration end conditions are met, including that the current iteration round number meets the preset iteration round number threshold, there is currently no incompletely expanded node, there is currently a fully expanded node and the sum of the node scores of the fully expanded node and the corresponding previous node is greater than or equal to the first preset score threshold, and / or there is currently a fully expanded node and the sum of the node scores of each incompletely expanded node and the corresponding previous node is less than the second preset score threshold.
10. The method according to claim 4 or 6, characterized in that Generating the corresponding corpus of each of the sub-nodes includes: For each of the sub-nodes, corresponding corpus is generated according to the sub-node and the corresponding previous node.
11. A corpus generation model training device, characterized in that: The device comprises: an acquisition unit configured to acquire a plurality of fine-tuning samples, wherein the fine-tuning samples include a corpus generation prompt word and a corresponding fine-tuning corpus, the corpus generation prompt word includes a reference information set and a corpus generation instruction, the reference information set includes a plurality of reference information, the corpus generation instruction is configured to instruct a large language model to generate corresponding corpus based on the reference information set and to add a reference identifier to the generated corpus, the reference identifier is configured to mark a reference relationship between each sentence in the corpus and each reference information in the reference information set, and the reference identifier is added to the fine-tuning corpus; A fine-tuning unit, configured to perform instruction fine-tuning on the large language model according to the plurality of fine-tuning samples; The training unit is used to obtain a corpus generation model based on the large language model fine-tuned according to the instructions.
12. A computer-readable storage medium storing computer program instructions, characterized in that: The computer program instructions implement the method according to any one of claims 1 to 10 when executed by a processor.
13. An electronic device, characterized in that: The device comprises: a memory for storing one or more computer program instructions; A processor, wherein the one or more computer program instructions are executed by the processor to implement the method according to any one of claims 1 to 10.
14. A computer program product, characterized in that When the computer program product is run on a computer, the computer is caused to perform the method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Big language model illusion relieving scheme based on citation correction
CN117910449A
Fine-grained large language model illusion detection method and device, and storage medium
CN118364065A
Cited By
Data generation method and device based on intelligent model, equipment and product
CN121765122A