Generative dialogue method and device based on action tag retrieval and readable medium
By introducing a generative dialogue method based on action tag retrieval in the dialogue generation system, the shortcomings of the existing system in dialogue generation direction control, action tag generation and coverage are solved, and more efficient and higher quality dialogue generation is achieved.
Patent Information
- Application Number
- CN202411877051.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-19
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2044-12-19
AI Technical Summary
The existing GPT-based dialogue generation system has shortcomings in dialogue generation direction control, action tag generation, coverage and efficiency, resulting in reduced dialogue effect and high labor costs.
A generative dialogue method based on action tag retrieval is adopted. By constructing coding models and soft tag generation models, a generative dialogue model is trained, and a vector library is used to search action tags to guide dialogue generation.
It significantly improves the efficiency and quality of dialogue generation, reduces the cost of manual labeling, enhances the controllability and diversity of dialogue, and can better cover diverse dialogue scenarios.
Smart Images

Figure CN119961391A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of dialogue generation, and in particular to a generative dialogue method, device and readable medium based on action tag retrieval. Background Art
[0002] At present, the representative model in the field of generative dialogue technology is GPT (Generative Pre-trained Transformer). GPT adopts an autoregressive model architecture, and its core is based on the decoder module of Transformer. Through training by next token prediction, GPT can use large-scale corpus and its huge number of parameters to store a large amount of professional field information, showing the strong potential of general AI intelligence. The emergence of GPT has had a subversive impact on the basic architecture of traditional dialogue systems, showing extremely high application value and broad prospects.
[0003] The current GPT-based dialogue generation system has the following shortcomings:
[0004] (1) It is difficult to control the direction of dialogue generation: GPT generates text in the form of next token, and the generation result is highly dependent on its decoding strategy. However, whether it is a deterministic top-k algorithm or a sampling-based top-p algorithm, it is difficult to ensure that the generation direction of the dialogue content is completely in line with expectations. Therefore, how to guide the model to generate target content through clear action labels plays a key role in improving the relevance of the dialogue and the user retention rate. Due to the limitations of its own architectural design and decoding strategy, it is obviously not a very good strategy to let the same GPT generate its own action labels.
[0005] (2) Limitations of the model’s self-generated action labels: Due to the limitations of the GPT model’s architecture design and decoding strategy, relying solely on GPT to generate action labels is obviously not an ideal solution. This approach not only makes it difficult to optimize conversation quality, but may also affect the consistency and diversity of generated content.
[0006] (3) Limited coverage: In real online conversation environments, the user’s expression methods, the complexity of questions, and the diversity of scenarios make it difficult for the action labeling strategy specified by experts to cover all possible conversation scenarios. This means that in some cases, the system may not be able to cope with sudden or atypical conversation situations, resulting in a decline in conversation effectiveness.
[0007] (4) High labor cost and low efficiency: Relying on the experience of experts to specify action labels is very time-consuming and laborious. The number of experts in each field is limited, and a lot of time is required to design, adjust, and optimize. This not only reduces efficiency, but also makes the scalability and rapid iteration of the model difficult, especially when facing new dialogue scenarios or requirements.
[0008] In addition, different customers have different conversation habits with customer service robots. How to learn the distribution of action labels through original conversations to simulate human conversation habits and achieve command following is also one of the difficulties. Summary of the invention
[0009] The purpose of this application is to propose a generative dialogue method, device and readable medium based on action tag retrieval in response to the above-mentioned technical problems.
[0010] In a first aspect, the present invention provides a generative dialogue method based on action tag retrieval, comprising the following steps:
[0011] Constructing and training an encoding model to obtain a trained encoding model, inputting candidate historical dialogues into the trained encoding model respectively to obtain candidate historical dialogue vectors, and storing the candidate historical dialogue vectors and their corresponding answers in a vector library, wherein the answers include reply sentences and action labels corresponding to the reply sentences;
[0012] Acquire training data, where each sample in the training data includes a historical dialogue and its corresponding answer; construct a soft label generation model and a generative dialogue model, input the answer of each sample in the training data into the soft label generation model, guide the soft label generation model through prompt engineering to modify the action label in the answer of each sample into an instruction sentence according to the action label and reply sentence in the answer of each sample, and output soft label data, where the soft label generation model includes an instruction sentence and a reply sentence;
[0013] The generative dialogue model is trained using the historical dialogue and soft-label data of each sample in the training data to obtain a trained generative dialogue model;
[0014] The historical conversation to be replied is obtained and input into the trained encoding model to obtain the historical conversation vector to be replied, the top N most similar candidate historical conversation vectors are retrieved from the vector library for the historical conversation vector to be replied, action labels are extracted from the answers corresponding to the top N most similar candidate historical conversation vectors, the historical conversation to be replied and the action labels are input into the trained generative dialogue model to generate the corresponding reply sentence.
[0015] Preferably, the encoding model includes a pre-trained Bert module or a pre-trained Roberta-wwm-base model.
[0016] Preferably, the training process of the encoding model is as follows:
[0017] Construct historical dialogues and their corresponding positive and negative samples;
[0018] Input the historical conversations and their corresponding positive and negative samples into the encoding model to obtain the historical conversation vector, positive sample vector, and negative sample vector;
[0019] Contrastive learning loss function is constructed based on historical dialogue vectors, positive sample vectors, and negative sample vectors, as shown in the following formula:
[0020]
[0021] Among them, L represents the contrastive learning loss function, h represents the historical dialogue vector, p represents the positive sample vector, and n i represents the i-th negative sample vector, i∈{1,2,...,N}, N is the number of negative samples; sim represents cosine similarity, τ represents temperature coefficient;
[0022] The encoding model is trained based on the contrastive learning loss function to obtain a trained encoding model.
[0023] Preferably, the soft label generation model is guided by the prompt engineering to modify the action label in the answer of each sample into an instruction sentence according to the action label and reply sentence in the answer of each sample, and output the soft label data, specifically including:
[0024] Construct prompt words, which require the user to determine the separator between the action label and the reply statement in the answer according to the given answer, and modify the content in the action label to make it diverse and at the same time be able to appropriately describe the content of the reply statement;
[0025] The soft label generation model generates soft label data based on the prompt words and answers.
[0026] Preferably, the generative dialogue model is trained using the historical dialogue and soft label data of each sample in the training data to obtain a trained generative dialogue model, specifically including:
[0027] The historical dialogue of each sample in the training data and the instruction sentences in the soft label data are input into the generative dialogue model to obtain the generated reply sentences, a loss function is constructed according to the generated reply sentences and the reply sentences in the soft label data, and the generative dialogue model is trained based on the loss function to obtain a trained generative dialogue model.
[0028] Preferably, the soft label generation model includes the qwen2-72b model, and the generative dialogue model includes the qwen2-7b model.
[0029] In a second aspect, the present invention provides a generative dialogue device based on action tag retrieval, comprising:
[0030] A vector library construction module is configured to construct and train an encoding model to obtain a trained encoding model, input candidate historical dialogues into the trained encoding model respectively to obtain candidate historical dialogue vectors, and store the candidate historical dialogue vectors and their corresponding answers in a vector library, wherein the answers include reply sentences and action labels corresponding to the reply sentences;
[0031] The data enhancement module is configured to obtain training data, each sample in the training data includes a historical dialogue and its corresponding answer; construct a soft label generation model and a generative dialogue model, input the answer of each sample in the training data into the soft label generation model, guide the soft label generation model through prompt engineering to modify the action label in the answer of each sample into an instruction sentence according to the action label and reply sentence in the answer of each sample, and output soft label data, the soft label generation model includes the instruction sentence and the reply sentence;
[0032] A model building module is configured to train the generative dialogue model using the historical dialogue and soft label data of each sample in the training data to obtain a trained generative dialogue model;
[0033] The retrieval generation module is configured to obtain the historical dialogue to be replied and input it into the trained encoding model to obtain the historical dialogue vector to be replied, retrieve the top N most similar candidate historical dialogue vectors from the vector library for the historical dialogue vector to be replied, extract the action label from the answers corresponding to the top N most similar candidate historical dialogue vectors, input the historical dialogue to be replied and the action label into the trained generative dialogue model, and generate the corresponding reply sentence.
[0034] In a third aspect, the present invention provides an electronic device comprising one or more processors; a storage device for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation manner in the first aspect.
[0035] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described in any implementation manner in the first aspect.
[0036] In a fifth aspect, the present invention provides a computer program product, comprising a computer program, which, when executed by a processor, implements the method described in any implementation manner in the first aspect.
[0037] Compared with the prior art, the present invention has the following beneficial effects:
[0038] (1) The generative dialogue method based on action label retrieval proposed in the present invention generates soft label data based on action labels and reply sentences through a soft label generation model, thereby achieving data enhancement effects, reducing the cost of manually annotating a large number of action labels, and significantly improving efficiency.
[0039] (2) The generative dialogue method based on action label retrieval proposed in the present invention generates soft label data through a soft label generation model, avoiding the limitations of self-generated action labels, and obtaining instruction sentences that are closer to reply sentences and ensure integrity and diversity. The historical dialogue and soft label data are used to guide the training of the generative dialogue model, so that the generative dialogue model can implement instruction following according to the action labels and generate more controllable reply sentences.
[0040] (3) The generative dialogue method based on action label retrieval proposed in the present invention can simulate human dialogue habits by learning action label distribution through historical dialogues and realize command following. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0042] Figure 1 A schematic diagram of a flow chart of a generative dialogue method based on action tag retrieval according to an embodiment of the present application;
[0043] Figure 2 A schematic diagram of a comparative learning process of an encoding model of a generative dialogue method based on action label retrieval according to an embodiment of the present application;
[0044] Figure 3 A schematic diagram of a generative dialogue device based on action tag retrieval according to an embodiment of the present application;
[0045] Figure 4 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0046] In order to make the purpose, technical scheme and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0047] Figure 1 A generative dialogue method based on action tag retrieval provided by an embodiment of the present application is shown, comprising the following steps:
[0048] S1, construct and train the encoding model to obtain the trained encoding model, input the candidate historical dialogues into the trained encoding model respectively, obtain the candidate historical dialogue vectors, store the candidate historical dialogue vectors and their corresponding answers in the vector library, and the answers include the reply sentences and the action labels corresponding to the reply sentences.
[0049] In a specific embodiment, the encoding model includes a pre-trained Bert module or a pre-trained Roberta-wwm-base model.
[0050] In a specific embodiment, the training process of the encoding model is as follows:
[0051] Construct historical dialogues and their corresponding positive and negative samples;
[0052] Input the historical conversations and their corresponding positive and negative samples into the encoding model to obtain the historical conversation vector, positive sample vector, and negative sample vector;
[0053] Contrastive learning loss function is constructed based on historical dialogue vectors, positive sample vectors, and negative sample vectors, as shown in the following formula:
[0054]
[0055] Among them, L represents the contrastive learning loss function, h represents the historical dialogue vector, p represents the positive sample vector, and n i represents the i-th negative sample vector, i∈{1,2,...,N}, N is the number of negative samples; sim represents cosine similarity, τ represents temperature coefficient;
[0056] The encoding model is trained based on the contrastive learning loss function to obtain a trained encoding model.
[0057] Specifically, refer to Figure 2, obtain the context dialogue, divide the context dialogue into historical dialogue and answer, take the answer with strong correlation with the historical dialogue as the positive sample, and take the answer with weak correlation with the dialogue as the negative sample. The encoding model uses the pre-trained Bert model or the pre-trained Roberta-wwm-base model, and trains the encoding model by contrastive learning. The historical dialogue and its corresponding positive and negative samples are input into the encoding model to obtain the historical dialogue vector, positive sample vector and negative sample vector, and respectively calculate the cosine similarity between the historical dialogue vector and the positive sample vector and the cosine similarity between the historical dialogue vector and the negative sample vector, and divide them by a smaller temperature coefficient respectively to make the encoding model more sensitive to the positive and negative samples, and calculate the contrastive learning loss function. The encoding model is trained based on the contrastive learning loss function to obtain the trained encoding model.
[0058] Furthermore, in the embodiment of the present application, N of the contrastive learning loss function is equal to 2, that is, the negative samples respectively select the reply statements of other rounds of the current historical dialogue and any reply statements in other dialogues as negative samples. The temperature coefficient is used to control the sharpness of the distribution. A lower temperature coefficient makes the branch sharper, while a higher temperature coefficient makes it smoother.
[0059] Using the trained encoding model to build a vector library is to input each candidate historical dialogue into the trained encoding model to obtain the corresponding candidate historical dialogue vector. Several candidate historical dialogue vectors and their corresponding answers are stored in the vector library. Each answer is composed of a reply sentence and an action label (action). A separator is set between the reply sentence and the action label. The action label is a small number of categories of labels predefined by expert experience and obtained through manual annotation. The action label includes consultation, Q&A, and set, and the secondary labels can be further refined on this basis, such as Q&A-examination introduction, consultation-disease-related consultation, etc. However, the sample size of the action label is too small, and further data enhancement is needed to obtain a larger sample size.
[0060] S2, obtain training data, each sample in the training data includes historical dialogues and their corresponding answers; build a soft label generation model and a generative dialogue model, input the answer of each sample in the training data into the soft label generation model, guide the soft label generation model through prompt engineering to modify the action label in the answer of each sample into an instruction sentence according to the action label and reply sentence in the answer of each sample, and output soft label data, the soft label generation model includes instruction sentences and reply sentences.
[0061] In a specific embodiment, the soft label generation model includes the qwen2-72b model, and the generative dialogue model includes the qwen2-7b model.
[0062] In a specific embodiment, the soft label generation model is guided by the prompt engineering to modify the action label in the answer of each sample into an instruction sentence according to the action label and reply sentence in the answer of each sample, and output the soft label data, specifically including:
[0063] Construct prompt words, which require the user to determine the separator between the action label and the reply statement in the answer according to the given answer, and modify the content in the action label to make it diverse and at the same time be able to appropriately describe the content of the reply statement;
[0064] The soft label generation model generates soft label data based on the prompt words and answers.
[0065] Specifically, the embodiment of the present application generates soft label data by constructing a prompt word (prompt), and the soft label generation model uses a qwen2-72b model with a larger parameter amount to meet the requirements of the soft label data generation task. The answer in the training data is spliced with an action label and a reply sentence, and the action label and the reply sentence are separated by a separator. The soft label data is spliced with an instruction sentence and a reply sentence, and the action label and the reply sentence are separated by a separator. The generative dialogue model is trained through historical dialogues and soft label data, and the generative dialogue model uses a qwen2-7b model with a smaller amount of calculation.
[0066] In order to enhance the generative dialogue model's understanding of instruction statements, the embodiment of the present application constructs soft label data through the following pipeline.
[0067] (1) One answer was collected: "Q&A - Introduction to examination, medical consultation - disease-related medical consultation\t||\tIf you usually have poor sleep quality, you can also do liver function and blood routine tests. What do you think?"
[0068] (2) Construct a prompt word: "Please modify the content of the action label based on the given answer, using \t||\t as the separator between the action label and the reply statement, to make it more diverse and more accurately describe the content of the reply statement." Input the answer and prompt word into the soft label generation model to generate soft label data.
[0069] (3) Output soft label data: "Reply to the visitor on how to check the symptoms through examination, and ask the visitor whether he is willing to continue the treatment.\t||\tIf you usually have poor sleep quality, you can also do liver function and blood routine tests. What do you think?". This part of "Reply to the visitor on how to check the symptoms through examination, and ask the visitor whether he is willing to continue the treatment" is a command sentence modified by action labels and reply sentences. Different command sentences can be generated by combining action labels with reply sentences, thereby increasing the sample size and strengthening the command following ability.
[0070] (4) The historical dialogue of each sample in the training data and the generated soft-label data are used together to train the generative dialogue model, and the trained generative dialogue model is used to conduct generative dialogue.
[0071] S3, using the historical dialogue and soft label data of each sample in the training data to train the generative dialogue model to obtain a trained generative dialogue model.
[0072] In a specific embodiment, step S3 specifically includes:
[0073] The historical dialogue of each sample in the training data and the instruction sentences in the soft label data are input into the generative dialogue model to obtain the generated reply sentences, a loss function is constructed according to the generated reply sentences and the reply sentences in the soft label data, and the generative dialogue model is trained based on the loss function to obtain a trained generative dialogue model.
[0074] Specifically, in the training process of the generative dialogue model, the historical dialogue of each sample in the training data and the instruction sentence in the soft label data need to be input into the generative dialogue model to obtain the response sentence generated by the generative dialogue model, and the loss function is constructed using the generated response sentence and the response sentence in the soft label data. The generative dialogue model is trained on the basis of the loss function to obtain a trained generative dialogue model. The trained generative dialogue model can enhance the understanding of action labels, improve the relevance and consistency of the generated content, reduce the reliance on expert experience, reduce labor costs, and improve the scalability and flexibility of the system.
[0075] S4, obtain the historical conversation to be replied and input it into the trained encoding model to obtain the historical conversation vector to be replied, retrieve the top N most similar candidate historical conversation vectors from the vector library for the historical conversation vector to be replied, extract the action label from the answers corresponding to the top N most similar candidate historical conversation vectors, input the historical conversation to be replied and the action label into the trained generative dialogue model, and generate the corresponding reply sentence.
[0076] Specifically, the trained generative dialogue model is deployed, and in the inference stage, the historical dialogue to be replied is input into the trained encoding model to generate the historical dialogue vector to be replied. The first N candidate historical dialogue vectors that are most similar to the historical dialogue vector to be replied are retrieved from the vector library, and the cosine similarity between the historical dialogue vector to be replied and the candidate historical dialogue vector can be calculated. In one embodiment, N=5, so the 5 most similar candidate historical dialogue vectors are retrieved from the vector library, and the action labels are extracted from the corresponding answers. The historical dialogue to be replied and the extracted action labels are input into the trained generative dialogue model to generate the corresponding reply sentence.
[0077] Further references Figure 3 As an implementation of the methods shown in the above figures, the present application provides an embodiment of a generative dialogue device based on action tag retrieval, and the device embodiment is similar to Figure 1 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.
[0078] The embodiment of the present application provides a generative dialogue device based on action tag retrieval, including:
[0079] The vector library construction module 1 is configured to construct and train a coding model to obtain a trained coding model, input candidate historical dialogues into the trained coding model respectively to obtain candidate historical dialogue vectors, and store the candidate historical dialogue vectors and their corresponding answers in a vector library, wherein the answers include reply sentences and action labels corresponding to the reply sentences;
[0080] The data enhancement module 2 is configured to obtain training data, each sample in the training data includes a historical dialogue and its corresponding answer; construct a soft label generation model and a generative dialogue model, input the answer of each sample in the training data into the soft label generation model, guide the soft label generation model through prompt engineering to modify the action label in the answer of each sample into an instruction sentence according to the action label and reply sentence in the answer of each sample, and output soft label data, the soft label generation model includes an instruction sentence and a reply sentence;
[0081] A model building module 3 is configured to train the generative dialogue model using the historical dialogue and soft label data of each sample in the training data to obtain a trained generative dialogue model;
[0082] The retrieval generation module 4 is configured to obtain the historical conversation to be replied and input it into the trained encoding model to obtain the historical conversation vector to be replied, retrieve the top N most similar candidate historical conversation vectors from the vector library for the historical conversation vector to be replied, extract the action label from the answers corresponding to the top N most similar candidate historical conversation vectors, input the historical conversation to be replied and the action label into the trained generative dialogue model, and generate the corresponding reply sentence.
[0083] Figure 4 Schematic diagram of the hardware structure of the electronic device provided by the embodiment of the present invention. Figure 4 As shown, the electronic device of this embodiment includes: a processor 401 and a memory 402; wherein the memory 402 is used to store computer-executable instructions; the processor 401 is used to execute the computer-executable instructions stored in the memory to implement the various steps performed by the electronic device in the above embodiment. For details, please refer to the relevant description in the above method embodiment.
[0084] Optionally, the memory 402 may be independent or integrated with the processor 401 .
[0085] When the memory 402 is independently provided, the electronic device further includes a bus 403 for connecting the memory 402 and the processor 401 .
[0086] The embodiment of the present invention further provides a computer storage medium, in which computer execution instructions are stored. When the processor 401 executes the computer execution instructions, the above method is implemented.
[0087] The embodiment of the present invention further provides a computer program product, including a computer program. When the computer program is executed by the processor 401, the above method is implemented.
[0088] In the embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic, for example, the division of modules is only a logical function division, and there may be other division methods in actual implementation, such as multiple modules can be combined or integrated into another system, or some features can be ignored or not executed. Another point, the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or modules, which can be electrical, mechanical or other forms.
[0089] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the modules may be selected according to actual needs to implement the solution of this embodiment.
[0090] In addition, each functional module in each embodiment of the present invention may be integrated into one processing unit, each module may exist physically separately, or two or more modules may be integrated into one unit. The unit formed by the above modules may be implemented in the form of hardware or in the form of hardware plus software functional units.
[0091] The above-mentioned integrated module implemented in the form of a software function module can be stored in a computer-readable storage medium. The above-mentioned software function module is stored in a storage medium, including a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor 401 to perform some steps of the methods of various embodiments of the present application.
[0092] It should be understood that the processor 401 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), etc. A general-purpose processor may be a microprocessor or the processor 401 may be any conventional processor 401, etc. The steps of the method disclosed in the invention may be directly embodied in the hardware processor 401 for execution, or may be executed by a combination of hardware and software modules in the processor 401.
[0093] The memory 402 may include a high-speed RAM memory, and may also include a non-volatile storage NVM, such as at least one disk memory, and may also be a USB flash drive, a mobile hard disk, a read-only memory, a magnetic disk, or an optical disk.
[0094] The bus 403 may be an Industry Standard Architecture (ISA), a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus 403 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, the bus 403 in the drawings of the present application is not limited to only one bus 403 or one type of bus 403.
[0095] The above storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk. The storage medium can be any available medium that can be accessed by a general or special purpose computer.
[0096] An exemplary storage medium is coupled to the processor 401, so that the processor 401 can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor 401. The processor 401 and the storage medium can be located in an application specific integrated circuit (ASIC). Of course, the processor 401 and the storage medium can also exist as discrete components in an electronic device or a main control device.
[0097] Those skilled in the art can understand that all or part of the steps of implementing the above-mentioned method embodiments can be completed by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, the steps of the above-mentioned method embodiments are executed; and the aforementioned storage medium includes: ROM, RAM, disk or optical disk and other media that can store program codes.
[0098] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A generative dialogue method based on action label retrieval, characterized in that: The following steps are involved: Constructing and training an encoding model to obtain a trained encoding model, inputting candidate historical dialogues into the trained encoding model respectively to obtain candidate historical dialogue vectors, and storing the candidate historical dialogue vectors and their corresponding answers in a vector library, wherein the answers include reply statements and action labels corresponding to the reply statements; Acquire training data, wherein each sample in the training data includes a historical dialogue and its corresponding answer; construct a soft label generation model and a generative dialogue model, input the answer of each sample in the training data into the soft label generation model, guide the soft label generation model through prompt engineering to modify the action label in the answer of each sample into an instruction sentence according to the action label and the reply sentence in the answer of each sample, and output soft label data, wherein the soft label generation model includes the instruction sentence and the reply sentence; Using the historical dialogue and soft label data of each sample in the training data to train the generative dialogue model to obtain a trained generative dialogue model; The historical conversation to be replied is obtained and input into the trained encoding model to obtain the historical conversation vector to be replied, the top N most similar candidate historical conversation vectors are retrieved from the vector library for the historical conversation vector to be replied, action labels are extracted from the answers corresponding to the top N most similar candidate historical conversation vectors, the historical conversation to be replied and the action labels are input into the trained generative dialogue model to generate a corresponding reply sentence.
2. The generative dialogue method based on action tag retrieval according to claim 1 is characterized in that: The encoding model includes a pre-trained Bert module or a pre-trained Roberta-wwm-base model.
3. The generative dialogue method based on action tag retrieval according to claim 1 is characterized in that: The training process of the encoding model is as follows: Construct historical dialogues and their corresponding positive and negative samples; Inputting the historical conversation and its corresponding positive samples and negative samples into the encoding model respectively to obtain a historical conversation vector, a positive sample vector and a negative sample vector; A contrastive learning loss function is constructed based on the historical dialogue vector, the positive sample vector, and the negative sample vector, as shown in the following formula: Among them, L represents the contrastive learning loss function, h represents the historical dialogue vector, p represents the positive sample vector, and n i represents the i-th negative sample vector, i∈{1,2,...,N}, N is the number of negative samples; sim represents cosine similarity, τ represents temperature coefficient; The encoding model is trained based on the contrastive learning loss function to obtain a trained encoding model.
4. The generative dialogue method based on action tag retrieval according to claim 1 is characterized in that: The soft label generation model is guided by the prompt engineering to modify the action label in the answer of each sample into an instruction sentence according to the action label and reply sentence in the answer of each sample, and output soft label data, specifically including: Constructing prompt words, wherein the prompt words require determining the separator between the action label and the reply statement in the answer according to the given answer, and modifying the content in the action label to make it diverse and at the same time to appropriately describe the content of the reply statement; The soft label generation model generates soft label data according to the prompt word and the answer.
5. The generative dialogue method based on action tag retrieval according to claim 1 is characterized in that: The generative dialogue model is trained using the historical dialogue and soft label data of each sample in the training data to obtain a trained generative dialogue model, specifically including: The historical dialogue of each sample in the training data and the instruction sentences in the soft label data are input into the generative dialogue model to obtain a generated reply sentence, a loss function is constructed according to the generated reply sentence and the reply sentence in the soft label data, and the generative dialogue model is trained based on the loss function to obtain a trained generative dialogue model.
6. The generative dialogue method based on action tag retrieval according to claim 1 is characterized in that: The soft label generation model includes the qwen2-72b model, and the generative dialogue model includes the qwen2-7b model.
7. A generative dialogue device based on action tag retrieval, characterized in that: include: A vector library construction module is configured to construct and train a coding model to obtain a trained coding model, input candidate historical dialogues into the trained coding model respectively to obtain candidate historical dialogue vectors, and store the candidate historical dialogue vectors and their corresponding answers in a vector library, wherein the answers include reply sentences and action labels corresponding to the reply sentences; A data enhancement module is configured to obtain training data, each sample in the training data includes a historical conversation and its corresponding answer; Constructing a soft label generation model and a generative dialogue model, inputting the answer of each sample in the training data into the soft label generation model, guiding the soft label generation model to modify the action label in the answer of each sample into an instruction sentence according to the action label and the reply sentence in the answer of each sample through prompt engineering, and outputting soft label data, wherein the soft label generation model includes the instruction sentence and the reply sentence; A model building module is configured to train the generative dialogue model using the historical dialogue and soft label data of each sample in the training data to obtain a trained generative dialogue model; The retrieval generation module is configured to obtain the historical dialogue to be replied and input it into the trained encoding model to obtain the historical dialogue vector to be replied, retrieve the top N most similar candidate historical dialogue vectors from the vector library for the historical dialogue vector to be replied, extract action labels from the answers corresponding to the top N most similar candidate historical dialogue vectors, input the historical dialogue to be replied and the action labels into the trained generative dialogue model, and generate corresponding reply sentences.
8. An electronic device comprising: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Dialogue data generation method and device, equipment and medium
CN114118101A
Training method and device of dialogue generation model and dialogue generation method and device
CN114547272A
Open domain dialogue generation method, system and device based on dialogue relationship and medium
CN115563254A
Deep learning based dialog method, apparatus, and device
US20190228070A1
Open domain dialog reply method and system based on thematic enhancement
US20240062006A1