Retrieval enhancement-based generative language model training method, dialogue generation method and apparatus
By introducing a generative language model training method based on retrieval enhancement in RAG technology, difficult problems in the data preparation process are solved, the accuracy and training efficiency of dialogue generation are improved, and the needs of vertical fields are met.
Patent Information
- Application Number
- CN202510101063.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-06-03
AI Technical Summary
The existing RAG technology has difficulties in the data preparation process, especially in the construction of training data, which leads to the high cost of obtaining high-quality dialogue data and is unable to meet the actual needs of vertical fields.
A training method based on retrieval enhancement generative language model is proposed, including data cleaning, knowledge extraction, retrieval model training, scoring model training, and hybrid training data construction and training of generative language models. Through this method, high-quality training data can be constructed and the generative language model can be trained in combination with the search model and the database.
Improves the accuracy of RAG replies, reduces training costs, and meets the actual needs of vertical fields.
Smart Images

Figure CN120087475A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of dialogue generation, and in particular to a training method, a dialogue generation method and a device for a generative language model based on retrieval enhancement. Background Art
[0002] As one of the current mainstream application solutions for large language models (LLMs), the main principle of RAG (Retrieval Augmented Generation) is to obtain relevant knowledge through retrieval and integrate it into the Prompt, enabling the large language model to refer to the corresponding knowledge and give reasonable answers. Therefore, the core of RAG can be understood as "retrieval + generation". The former mainly utilizes the efficient storage and retrieval capabilities of the vector database to recall the target knowledge; the latter utilizes the large language model and Prompt engineering to reasonably utilize the recalled knowledge to generate the target answer.
[0003] Currently, most existing RAGs only involve the optimization of a certain module, such as the optimization of the retriever or the chain of thought. There is still no technology that takes the LLM as the core to complete the data preparation process of RAG, including data cleaning, knowledge extraction, retrieval model, scoring model, and training data construction of the generative language model, etc.; the difficulty of training data construction is relatively large, and the cost of obtaining high-quality dialogue data is relatively high.
[0004] Existing RAG technologies still cannot meet the actual needs of vertical fields. For example, the retrieval model trained in an unsupervised manner cannot accurately retrieve the required knowledge, and the scoring model often selects sub-optimal responses, resulting in inaccurate responses generated by GPT. Summary of the Invention
[0005] The purpose of this application is to propose a training method, a dialogue generation method and a device for a generative language model based on retrieval enhancement for the above-mentioned technical problems.
[0006] In a first aspect, the present invention provides a training method for a generative language model based on retrieval enhancement, including the following steps:
[0007] Obtain a plurality of original historical dialogues and perform data cleaning to obtain a plurality of cleaned historical dialogues;
[0008] Construct a pre-trained first large language model and a second large language model. Input each cleaned historical dialogue into the first large language model to generate corresponding question description statements and response statements. Input the question description statements and response statements into the second large language model to generate corresponding single-piece first knowledge. Perform data screening on a number of cleaned historical dialogues and their corresponding first knowledge and response statements to obtain the screened data, which includes high-quality historical dialogues and their corresponding first knowledge and response statements.
[0009] Construct a retrieval model and train it to obtain a trained retrieval model. Construct a database that stores the dialogue vectors obtained by passing each high-quality historical dialogue in the screened data through the trained retrieval model and the corresponding single-piece first knowledge.
[0010] Select each high-quality historical dialogue in a part of the screened data and input it into the trained retrieval model to obtain a dialogue vector. Retrieve n pieces of second knowledge in the database using the dialogue vector. Combine the retrieved n pieces of second knowledge with the single-piece first knowledge to form n + 1 pieces of knowledge corresponding to each high-quality historical dialogue. Construct first training data based on each high-quality historical dialogue and its corresponding knowledge and response statements, and construct second training data based on each high-quality historical dialogue and its corresponding response statements in another part of the screened data. Mix the first training data, second training data, and the public dataset to obtain mixed training data.
[0011] Construct a generative language model, and use the mixed training data in combination with the trained retrieval model and the database to train the generative language model to obtain a trained generative language model.
[0012] Preferably, the retrieval model includes a BERT model, and the generative language model includes a GPT model.
[0013] Preferably, the training process of the retrieval model is as follows:
[0014] Input the high-quality historical dialogue and the response statement into the retrieval model respectively, and output the corresponding dialogue vector and response vector. Construct similar sample pairs and dissimilar sample pairs.
[0015] Construct a contrastive learning loss function. The value of the contrastive learning loss function is directly proportional to the similarity of the similar sample pairs and inversely proportional to the similarity of the dissimilar sample pairs. Train the retrieval model based on the contrastive learning loss function to obtain a trained retrieval model.
[0016] Preferably, during data screening, a pre-trained third large language model is used to score the cleaned historical conversations and their corresponding first knowledge, or the cleaned historical conversations and their corresponding reply sentences, to obtain quality scores. The quality scores are compared with a threshold, and the cleaned historical conversations and their corresponding first knowledge and reply sentences with quality scores higher than the threshold are screened out to obtain the screened data.
[0017] Preferably, the training process of the generative language model is as follows:
[0018] If there is corresponding knowledge for the high-quality historical conversations in the mixed training data, the high-quality historical conversations are concatenated with the knowledge and then input into the trained retrieval model to output reply sentences; if there is no corresponding knowledge for the high-quality historical conversations in the mixed training data, the high-quality historical conversations are input into the generative language model to output reply sentences. A cross-entropy loss function is constructed based on the reply sentences output by the generative language model and the reply sentences corresponding to the high-quality historical conversations in the mixed training data, and the generative language model is supervised and trained based on the cross-entropy loss function to obtain the trained generative language model.
[0019] In a second aspect, the present invention provides a training device for a generative language model based on retrieval enhancement, including:
[0020] A data processing module configured to obtain a plurality of original historical conversations and perform data cleaning to obtain a plurality of cleaned historical conversations;
[0021] A knowledge acquisition module configured to construct a pre-trained first large language model and a second large language model, input each cleaned historical conversation into the first large language model to generate corresponding question description sentences and reply sentences; input the question description sentences and reply sentences into the second large language model to generate corresponding single-piece first knowledge, and perform data screening on a plurality of cleaned historical conversations and their corresponding first knowledge and reply sentences to obtain the screened data, and the screened data includes high-quality historical conversations and their corresponding first knowledge and reply sentences;
[0022] A database construction module configured to construct and train a retrieval model to obtain a trained retrieval model; construct a database, and store the dialogue vectors obtained by each high-quality historical conversation in the screened data passing through the trained retrieval model and the corresponding single-piece first knowledge;
[0023] The training data construction module is configured to select each high-quality historical dialogue from a part of the filtered data and input it into the trained retrieval model to obtain a dialogue vector, retrieve n pieces of second knowledge in the database based on the dialogue vector, form n + 1 pieces of knowledge corresponding to each high-quality historical dialogue by combining the retrieved n pieces of second knowledge with a single piece of first knowledge, construct first training data based on each high-quality historical dialogue and its corresponding knowledge and response sentences, construct second training data based on each high-quality historical dialogue and its corresponding response sentences in another part of the filtered data, and mix the first training data, the second training data, and the public data set to obtain mixed training data;
[0024] The model training module is configured to construct a generative language model, and use the mixed training data in combination with the trained retrieval model and the database to train the generative language model to obtain a trained generative language model.
[0025] In a third aspect, the present invention provides a retrieval-enhanced dialogue generation method, which uses the trained generative language model and the trained retrieval model obtained by training using the method described in any implementation manner in the first aspect, and includes the following steps:
[0026] Construct and train a scoring model to obtain a trained scoring model;
[0027] Obtain the historical dialogue to be replied;
[0028] Input the historical dialogue to be replied into the trained retrieval model to obtain the dialogue vector corresponding to the historical dialogue to be replied;
[0029] Retrieve the dialogue vector corresponding to the historical dialogue to be replied in the database. If knowledge is retrieved, splice the historical dialogue to be replied and the knowledge and input it into the trained generative language model; if no knowledge is retrieved, input the historical dialogue to be replied into the trained generative language model to obtain m candidate response sentences, use the trained scoring model to score the historical dialogue to be replied and its corresponding knowledge or the historical dialogue to be replied and each of its corresponding candidate response sentences to obtain quality scores, and select the candidate response sentence corresponding to the maximum quality score as the response sentence output by the trained generative language model.
[0030] Preferably, the scoring model includes a BERT model and a multi-layer perceptron connected in sequence, and the training process of the scoring model is as follows:
[0031] Use the quality scores corresponding to the high-quality historical dialogues and the high-quality historical dialogues and their corresponding knowledge or high-quality historical dialogues and their response sentences to form scoring training data, and use the scoring training data to train the scoring model to obtain a trained scoring model.
[0032] Fourth aspect, the present invention provides an electronic device, including one or more processors; a storage device for storing one or more programs, which when executed by the one or more processors, cause the one or more processors to implement the method described in any implementation manner of the first aspect.
[0033] Fifth aspect, the present invention provides a computer-readable storage medium, having stored thereon a computer program, which when executed by a processor, implements the method described in any implementation manner of the first aspect.
[0034] Compared with the prior art, the present invention has the following beneficial effects:
[0035] (1) The training method of the retrieval-enhanced generative language model proposed by the present invention, by leveraging the capabilities of the SOTA LLM, performs fine cleaning of the training data and screening of high-quality data, solves the shortcoming of poor model performance after unsupervised fine-tuning, and greatly improves the accuracy of RAG responses.
[0036] (2) The training method of the retrieval-enhanced generative language model proposed by the present invention realizes and optimizes the training processes of the retrieval model, the scoring model, and the generative language model, shares the training data among different models, and greatly reduces the training cost. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.
[0038] Figure 1 It is a schematic flowchart of the training method of the retrieval-enhanced generative language model according to the embodiment of the present application;
[0039] Figure 2 It is a schematic diagram of the training device of the retrieval-enhanced generative language model according to the embodiment of the present application;
[0040] Figure 3 It is a schematic diagram of the hardware structure of the electronic device provided by the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0041] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. Apparently, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0042] Figure 1 There is shown a training method for a retrieval-enhanced generative language model provided by an embodiment of the present application, including the following steps:
[0043] S1. Obtain a plurality of original historical conversations and perform data cleaning to obtain a plurality of cleaned historical conversations.
[0044] Specifically, first perform data cleaning on the collected original historical conversations to ensure the accuracy and consistency of the data.
[0045] S2. Construct a pre-trained first large language model and a second large language model. Input each cleaned historical conversation into the first large language model to generate corresponding question description statements and reply statements; input the question description statements and reply statements into the second large language model to generate corresponding single-piece first knowledge. Perform data screening on the plurality of cleaned historical conversations and their corresponding first knowledge and reply statements to obtain screened data. The screened data includes high-quality historical conversations and their corresponding first knowledge and reply statements.
[0046] In a specific embodiment, during the data screening process, a pre-trained third large language model is used to score the cleaned historical conversations and their corresponding first knowledge or the cleaned historical conversations and their corresponding reply statements to obtain quality scores. Compare the quality scores with a threshold value, and screen out the cleaned historical conversations and their corresponding first knowledge and reply statements with quality scores higher than the threshold value to obtain the screened data.
[0047] Specifically, in the embodiment of the present application, a pre-trained first large language model is used to extract question description statements from the cleaned historical conversations and give corresponding reply statements, and then input the question description statements and the corresponding reply statements into the pre-trained second large language model to obtain single-piece first knowledge. An example is as follows:
[0048] Cleaned historical conversation:
[0049] A: I want to consult information related to the thyroid.
[0050] B: May I ask which aspects you would like to know about?
[0051] A: The treatment methods and causes of hyperthyroidism.
[0052] B: It is mainly caused by excessive thyroid hormones. There are two treatment methods: medication and surgery.
[0053] A: Mhm.
[0054] B: Our hospital has cured many patients by using medications such as propylthiouracil and methylthiouracil.
[0055] Problem description statement: Treatment methods and causes of hyperthyroidism.
[0056] Reply statement: It is mainly caused by excessive thyroid hormones. Our hospital uses treatment methods with medications such as propylthiouracil and methylthiouracil.
[0057] First knowledge: The treatment method of hyperthyroidism is to use medications such as propylthiouracil and methylthiouracil, and the cause is excessive thyroid hormones.
[0058] Furthermore, perform data screening on the cleaned historical conversations and their corresponding first knowledge and reply statements. Use the third large language model to score the cleaned historical conversations and their corresponding knowledge or the cleaned historical conversations and their corresponding reply statements to obtain quality scores. Then, perform screening based on the quality scores to obtain high-quality historical conversations and their corresponding first knowledge and reply statements, further ensuring the effectiveness of the data. The quality scores obtained at this time can be combined with the high-quality historical conversations and their corresponding first knowledge and reply statements to form the scoring training data required in the scoring model training process.
[0059] It should be noted that the first large language model, the second large language model, and the third large language model use the Qianwen model or the LLaMA model. The first large language model, the second large language model, and the third large language model can use the same model or different models, and are pre-trained according to specific tasks to enable them to meet their respective task requirements.
[0060] S3. Construct a retrieval model and train it to obtain a trained retrieval model; construct a database, in which each high-quality historical conversation in the screened data, the corresponding conversation vector obtained through the trained retrieval model, and the corresponding single piece of first knowledge are stored.
[0061] In a specific embodiment, the retrieval model includes a BERT model, and the generative language model includes a GPT model.
[0062] In a specific embodiment, the training process of the retrieval model is as follows:
[0063] Input the high-quality historical conversations and reply statements into the retrieval model respectively, and output the corresponding conversation vectors and reply vectors, and construct similar sample pairs and dissimilar sample pairs;
[0064] Construct a contrastive learning loss function. The value of the contrastive learning loss function is directly proportional to the similarity of similar sample pairs and inversely proportional to the similarity of dissimilar sample pairs. Train the retrieval model based on the contrastive learning loss function to obtain a trained retrieval model.
[0065] Specifically, construct a BERT-based retrieval model. This retrieval model uses a BERT model in the Duel encoder manner. During the training process of the retrieval model, input high-quality historical conversations and response sentences into the retrieval model respectively to obtain corresponding conversation vectors and response vectors. The conversation vectors and response vectors obtained from a high-quality historical conversation and its corresponding response sentence form a similar sample pair, and the conversation vectors and response vectors obtained from a high-quality historical conversation and its non-corresponding response sentence form a dissimilar sample pair. Calculate the similarity of similar sample pairs and the similarity of dissimilar sample pairs respectively, and calculate the contrastive learning loss function based on the similarity of similar sample pairs and the similarity of dissimilar sample pairs. The value of this contrastive learning loss function is directly proportional to the similarity of similar sample pairs and inversely proportional to the similarity of dissimilar sample pairs. Train the retrieval model based on this contrastive learning loss function to obtain a trained retrieval model.
[0066] S4. Select each high-quality historical conversation in some of the filtered data and input it into the trained retrieval model to obtain a conversation vector. Retrieve n pieces of second knowledge in the database with the conversation vector. Combine the retrieved n pieces of second knowledge with a single piece of first knowledge to form n + 1 pieces of knowledge corresponding to each high-quality historical conversation. Construct first training data based on each high-quality historical conversation, its corresponding knowledge, and response sentence. Construct second training data based on each high-quality historical conversation and its corresponding response sentence in another part of the filtered data. Mix the first training data, the second training data, and the public dataset to obtain mixed training data.
[0067] Specifically, obtain the conversation vector corresponding to each high-quality historical conversation through the trained retrieval model. Retrieve in the database the conversation vector corresponding to each high-quality historical conversation to retrieve n pieces of second knowledge that are semantically similar to the conversation vector. Combine the retrieved n pieces of second knowledge with the first knowledge corresponding to this high-quality historical conversation to form n + 1 pieces of knowledge corresponding to this high-quality historical conversation. Thus, the relevant data of this high-quality historical conversation can be expanded from 1 piece to n + 1 pieces, achieving the effect of data augmentation. At the same time, construct second training data through high-quality historical conversations and their corresponding response sentences, and add the public dataset to avoid catastrophic forgetting. Mix the first training data, the second training data, and the public dataset to obtain mixed training data.
[0068] S5. Construct a generative language model, and train the generative language model by using mixed training data and combining a trained retrieval model and a database to obtain a trained generative language model.
[0069] In a specific embodiment, the training process of the generative language model is as follows:
[0070] If there is corresponding knowledge in the high-quality historical conversations in the mixed training data, splice the high-quality historical conversations with the knowledge and input them into the trained retrieval model to output a response statement; if there is no corresponding knowledge in the high-quality historical conversations in the mixed training data, input the high-quality historical conversations into the generative language model to output a response statement, construct a cross-entropy loss function based on the response statement output by the generative language model and the response statement corresponding to the high-quality historical conversations in the mixed training data, and perform supervised training on the generative language model based on the cross-entropy loss function to obtain a trained generative language model.
[0071] Specifically, the generative language model can select the GPT structure or other generative language model structures, which is not limited here. Use this mixed training data to train the generative language model so that it forms a retrieval-enhanced generative language model with the trained retrieval model, database, etc., solving the disadvantage of poor model performance after unsupervised fine-tuning and greatly improving the accuracy of the responses of the retrieval-enhanced generative language model.
[0072] Further referring to Figure 2 , as an implementation of the methods shown in the above figures, an embodiment of a training device for a retrieval-enhanced generative language model is provided in the present application. This device embodiment corresponds to Figure 1 the method embodiment shown, and this device can be specifically applied to various electronic devices.
[0073] An embodiment of a training device for a retrieval-enhanced generative language model is provided in an embodiment of the present application, including:
[0074] A data processing module 1, configured to obtain a number of original historical conversations and perform data cleaning to obtain a number of cleaned historical conversations;
[0075] A knowledge acquisition module 2, configured to construct a pre-trained first large language model and a second large language model, input each cleaned historical conversation into the first large language model to generate corresponding question description statements and response statements; input the question description statements and response statements into the second large language model to generate corresponding single-piece first knowledge, and perform data screening on a number of cleaned historical conversations and their corresponding first knowledge and response statements to obtain screened data, and the screened data includes high-quality historical conversations and their corresponding first knowledge and response statements;
[0076] A database construction module 3, configured to construct a retrieval model and train it to obtain a trained retrieval model; construct a database, where each high-quality historical conversation in the filtered data is stored with the corresponding single piece of first knowledge obtained by passing through the trained retrieval model;
[0077] A training data construction module 4, configured to select each high-quality historical conversation in a part of the filtered data and input it into the trained retrieval model to obtain a conversation vector, retrieve n pieces of second knowledge for the conversation vector in the database, and form n + 1 pieces of knowledge corresponding to each high-quality historical conversation by combining the retrieved n pieces of second knowledge with the single piece of first knowledge, construct first training data based on each high-quality historical conversation and its corresponding knowledge and response statements, construct second training data based on each high-quality historical conversation and its corresponding response statements in another part of the filtered data, and mix the first training data, the second training data, and a public dataset to obtain mixed training data;
[0078] A model training module 5, configured to construct a generative language model, and use the mixed training data in combination with the trained retrieval model and the database to train the generative language model to obtain a trained generative language model.
[0079] An embodiment of the present application also proposes a retrieval-enhanced dialogue generation method, which uses the trained generative language model and the trained retrieval model obtained by training with the above-mentioned training method of the retrieval-enhanced generative language model, and includes the following steps:
[0080] Construct a scoring model and train it to obtain a trained scoring model;
[0081] Obtain the historical conversation to be replied;
[0082] Input the historical conversation to be replied into the trained retrieval model to obtain the conversation vector corresponding to the historical conversation to be replied;
[0083] Retrieve the conversation vector corresponding to the historical conversation to be replied in the database. If knowledge is retrieved, splice the historical conversation to be replied and the knowledge and input it into the trained generative language model; if no knowledge is retrieved, input the historical conversation to be replied into the trained generative language model to obtain m candidate response statements, use the trained scoring model to score the historical conversation to be replied and its corresponding knowledge or the historical conversation to be replied and each corresponding candidate response statement to obtain a quality score, and select the candidate response statement corresponding to the maximum quality score as the response statement output by the trained generative language model.
[0084] In a specific embodiment, the scoring model includes a BERT model and a multi-layer perceptron connected in sequence. The training process of the scoring model is as follows:
[0085] The quality scores corresponding to high-quality historical conversations, the high-quality historical conversations and their corresponding knowledge, or the high-quality historical conversations and their reply sentences are used to form scoring training data. The scoring training data is used to train the scoring model to obtain a trained scoring model.
[0086] Specifically, the embodiment of the present application also combines the quality scores corresponding to high-quality historical conversations obtained by scoring with a third large language model with the high-quality historical conversations and their corresponding knowledge, or the high-quality historical conversations and their reply sentences to form scoring training data. Further, the constructed scoring model is trained using the scoring training data, which can realize sharing of training data and greatly reduce the training cost. Finally, the trained retrieval model, the trained generative language model, and the trained scoring model obtained by separate training are combined with the database to form a retrieval-enhanced generative language model (RAG), and the reply sentence corresponding to the input historical conversation is generated using the retrieval-enhanced generative language model.
[0087] Figure 3 This is a schematic diagram of the hardware structure of the electronic device provided by the embodiment of the present invention. As Figure 3 shown, the electronic device of this embodiment includes: a processor 301 and a memory 302; wherein the memory 302 is used to store computer execution instructions; the processor 301 is used to execute the computer execution instructions stored in the memory to implement the various steps executed by the electronic device in the above embodiment. Specifically, reference can be made to the relevant descriptions in the foregoing method embodiments.
[0088] Optionally, the memory 302 can be either independent or integrated with the processor 301.
[0089] When the memory 302 is independently provided, the electronic device further includes a bus 303 for connecting the memory 302 and the processor 301.
[0090] The embodiment of the present invention also provides a computer storage medium, in which computer execution instructions are stored. When the processor 301 executes the computer execution instructions, the above method is implemented.
[0091] The embodiment of the present invention also provides a computer program product, including a computer program. When the computer program is executed by the processor 301, the above method is implemented.
[0092] In the embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the couplings or direct couplings or communication connections shown or discussed with each other can be indirect couplings or communication connections through some interfaces, devices or modules, and can be in electrical, mechanical or other forms.
[0093] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to implement the solution of this embodiment.
[0094] In addition, in each embodiment of the present invention, the functional modules can be integrated in a processing unit, or each module can exist physically alone, or two or more modules can be integrated in one unit. The unit formed by the above modules can be implemented in the form of hardware, or in the form of a combination of hardware and software functional units.
[0095] The integrated modules implemented in the form of software functional modules can be stored in a computer-readable storage medium. The above software functional modules are stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor 301 to execute some steps of the methods in various embodiments of the present application.
[0096] It should be understood that the above processor 301 can be a central processing unit (Central Processing Unit, abbreviated as CPU), and can also be other general-purpose processors, digital signal processors (Digital Signal Processor, abbreviated as DSP), application-specific integrated circuits (Application Specific Integrated Circuit, abbreviated as ASIC), etc. The general-purpose processor can be a microprocessor or the processor 301 can also be any conventional processor 301, etc. The steps of the method disclosed in combination with the invention can be directly embodied as being executed by the hardware processor 301, or can be executed by a combination of hardware and software modules in the processor 301.
[0097] The memory 302 may include high-speed RAM memory and may also include non-volatile storage NVM, such as at least one disk memory, and may also be a USB flash drive, a portable hard drive, a read-only memory, a magnetic disk, or an optical disc, etc.
[0098] The bus 303 may be an Industry Standard Architecture (ISA), a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus 303 can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience in representation, the bus 303 in the drawings of this application is not limited to only one bus 303 or one type of bus 303.
[0099] The above storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disc. The storage medium can be any available medium that can be accessed by a general-purpose or special-purpose computer.
[0100] An exemplary storage medium is coupled to the processor 301, so that the processor 301 can read information from the storage medium and can write information to the storage medium. Of course, the storage medium can also be a component of the processor 301. The processor 301 and the storage medium can be located in an Application Specific Integrated Circuits (ASIC). Of course, the processor 301 and the storage medium can also exist as discrete components in an electronic device or a master control device.
[0101] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments; and the foregoing storage medium includes: ROM, RAM, magnetic disks, or optical discs and other media that can store program codes.
[0102] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A training method for a generative language model based on retrieval enhancement, characterized in that: The following steps are involved: Obtain several original historical conversations and perform data cleaning to obtain several cleaned historical conversations; Constructing a pre-trained first large language model and a second large language model, inputting each cleaned historical dialogue into the first large language model to generate a corresponding question description sentence and a reply sentence; inputting the question description sentence and the reply sentence into the second large language model to generate a corresponding single first knowledge, performing data screening on a plurality of the cleaned historical dialogues and their corresponding first knowledge and reply sentences to obtain screened data, wherein the screened data includes high-quality historical dialogues and their corresponding first knowledge and reply sentences; A retrieval model is constructed and trained to obtain a trained retrieval model; a database is constructed to store a conversation vector and a corresponding single piece of first knowledge obtained by the trained retrieval model for each high-quality historical conversation in the screened data; Select each high-quality historical dialogue in part of the screened data and input it into the trained retrieval model to obtain a dialogue vector, retrieve n second knowledge from the database using the dialogue vector, and use the retrieved n second knowledge and a single first knowledge to form n+1 knowledge corresponding to each high-quality historical dialogue, construct first training data based on each high-quality historical dialogue and its corresponding knowledge and reply sentence, construct second training data based on each high-quality historical dialogue and its corresponding reply sentence in another part of the screened data, and mix the first training data, the second training data and the public data set to obtain mixed training data; A generative language model is constructed, and the mixed training data is used to train the generative language model in combination with the trained retrieval model and the database to obtain a trained generative language model.
2. The method for training a generative language model based on retrieval enhancement according to claim 1, characterized in that: The retrieval model includes a BERT model, and the generative language model includes a GPT model.
3. The method for training a retrieval-enhanced generative language model according to claim 1, characterized in that: The training process of the retrieval model is as follows: Inputting the high-quality historical dialogue and the reply sentence into the retrieval model respectively, outputting corresponding dialogue vectors and reply vectors, and constructing similar sample pairs and dissimilar sample pairs; A contrastive learning loss function is constructed, wherein the value of the contrastive learning loss function is proportional to the similarity of the similar sample pairs and inversely proportional to the similarity of the non-similar sample pairs; the retrieval model is trained based on the contrastive learning loss function to obtain a trained retrieval model.
4. The method for training a generative language model based on retrieval enhancement according to claim 1, characterized in that: During the data screening process, a pre-trained third language model is used to score the cleaned historical conversations and their corresponding first knowledge or the cleaned historical conversations and their corresponding reply sentences to obtain a quality score, and the quality score is compared with a threshold to screen out the cleaned historical conversations and their corresponding first knowledge and reply sentences with quality scores higher than the threshold to obtain screened data.
5. The method for training a generative language model based on retrieval enhancement according to claim 1, characterized in that: The training process of the generative language model is as follows: If there is corresponding knowledge in the high-quality historical dialogue in the mixed training data, the high-quality historical dialogue is spliced with the knowledge and input into the trained retrieval model to output a reply sentence; if there is no corresponding knowledge in the high-quality historical dialogue in the mixed training data, the high-quality historical dialogue is input into the generative language model, a reply sentence is output, a cross-entropy loss function is constructed based on the reply sentence output by the generative language model and the reply sentence corresponding to the high-quality historical dialogue in the mixed training data, and the generative language model is supervisedly trained based on the cross-entropy loss function to obtain a trained generative language model.
6. A training device for a generative language model based on retrieval enhancement, characterized in that: include: A data processing module is configured to obtain a number of original historical conversations and perform data cleaning to obtain a number of cleaned historical conversations; The knowledge acquisition module is configured to construct a pre-trained first large language model and a second large language model, input each cleaned historical dialogue into the first large language model to generate a corresponding question description sentence and a reply sentence; input the question description sentence and the reply sentence into the second large language model to generate a corresponding single first knowledge, and perform data screening on a plurality of the cleaned historical dialogues and their corresponding first knowledge and reply sentences to obtain screened data, wherein the screened data includes high-quality historical dialogues and their corresponding first knowledge and reply sentences; A database construction module is configured to construct and train a retrieval model to obtain a trained retrieval model; construct a database, wherein the database stores a conversation vector and a corresponding single piece of first knowledge obtained by the trained retrieval model for each high-quality historical conversation in the screened data; a training data construction module configured to select each high-quality historical dialogue in a part of the screened data and input it into the trained retrieval model to obtain a dialogue vector, retrieve n second knowledge from the database using the dialogue vector, and use the retrieved n second knowledge and a single first knowledge to form n+1 knowledge corresponding to each high-quality historical dialogue, construct first training data based on each high-quality historical dialogue and its corresponding knowledge and reply sentence, construct second training data based on each high-quality historical dialogue and its corresponding reply sentence in another part of the screened data, and mix the first training data, the second training data and the public data set to obtain mixed training data; The model training module is configured to construct a generative language model, and train the generative language model using the mixed training data and in combination with the trained retrieval model and database to obtain a trained generative language model.
7. A dialogue generation method based on retrieval enhancement, characterized in that: The trained generative language model and the trained retrieval model obtained by training the training method of the generative language model based on retrieval enhancement according to any one of claims 1 to 6 include the following steps: Construct a scoring model and train it to obtain a trained scoring model; Get the historical conversations to be replied; Inputting the historical conversation to be replied to into the trained retrieval model to obtain a conversation vector corresponding to the historical conversation to be replied; The dialogue vector corresponding to the historical dialogue to be replied is searched in the database. If knowledge is retrieved, the historical dialogue to be replied and the knowledge are spliced and input into the trained generative language model. If knowledge is not retrieved, the historical dialogue to be replied is input into the trained generative language model to obtain m candidate reply sentences. The trained scoring model is used to score the historical dialogue to be replied and its corresponding knowledge or the historical dialogue to be replied and each corresponding candidate reply sentence to obtain a quality score, and the candidate reply sentence corresponding to the maximum quality score is selected as the reply sentence output by the trained generative language model.
8. The method for generating dialogue based on retrieval enhancement according to claim 7, characterized in that: The scoring model includes a BERT model and a multi-layer perceptron connected in sequence, and the training process of the scoring model is as follows: The quality scores corresponding to the high-quality historical conversations and the high-quality historical conversations and their corresponding knowledge or the high-quality historical conversations and their reply statements constitute scoring training data, and the scoring model is trained using the scoring training data to obtain a trained scoring model.
9. An electronic device, comprising: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 5.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.
Citation Information
Cited By
Joint decision model-based undertaking word acquisition method and system
CN120509495A
Dialogue generation method and device based on improved DPO training and readable medium
CN121503645A
A dialogue generation method and device based on improved DPO training and readable medium
CN121503645B