A method, device, computer device, and storage medium for optimizing a dialogue model
By collecting and segmenting problem data in the intelligent dialogue system, setting loss function and reward function in the dialogue model, and using reinforcement learning algorithm to optimize the model, the problem that the intelligent dialogue system cannot correctly respond to user problems is solved, the dialogue quality is improved and the output distortion is avoided.
Patent Information
- Application Number
- CN202310910513.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-24
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2043-07-24
AI Technical Summary
The intelligent dialogue system cannot fully consider the various problems raised by the user, which leads to the response action that is not considered at the beginning of the design, and the response action may be a random response and cannot correctly respond to the user's questions.
The input problem data in the pre-trained dialogue model is collected through the application program interface, divided into three parts according to the preset ratio, and the loss function and reward function are set in the dialogue model. The optimized dialogue model is obtained through training, and the model is further optimized using reinforcement learning algorithm.
The dialogue quality of the dialogue model is improved, and the output of abnormal results is avoided, allowing intelligent customer service to respond to user problems more accurately.
Smart Images

Figure CN117112742B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and particularly to a method, device, computer device and storage medium for optimizing a dialogue model. Background Art
[0002] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include, for example, sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technologies, operation / interaction systems, mechatronics and other technologies. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0003] Natural Language Processing (NLP) is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between humans and computers using natural language. Natural language processing is a discipline that integrates linguistics, computer science, and mathematics. Therefore, the research in this field will involve natural language, that is, the language people use in daily life, so it has a close connection with the research of linguistics. Natural language processing technologies usually include text processing, semantic understanding, machine translation, robot question answering, knowledge graph and other technologies.
[0004] Machine learning is an interdisciplinary subject involving multiple fields such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. Machine learning specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include: artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, teaching learning and other technologies.
[0005] With the research and progress of artificial intelligence technology, artificial intelligence technology has been studied and applied in multiple fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, driverless, autonomous driving, drones, robots, smart healthcare, smart customer service, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.
[0006] Regarding intelligent customer service, due to the limitations of designers' thinking and the limitations of data, storage, and computing capabilities, the intelligent dialogue system cannot fully consider all kinds of questions raised by users. When the intelligent dialogue system encounters problems not considered at the beginning of its design, the response actions to the above problems can be regarded as random responses, often unable to correctly answer the questions raised by users, making users feel that the answers are off-topic. Summary of the Invention
[0007] Based on this, in view of the above technical problems, it is necessary to provide a dialogue model optimization method, device, computer device, and storage medium that can improve the dialogue quality of the dialogue model and avoid outputting abnormal results.
[0008] In a first aspect, a dialogue model optimization method is provided, and the method includes:
[0009] Collect question data input in the pre-trained dialogue model through an application programming interface, and divide the question data into three parts according to a preset ratio, namely first data, second data, and third data;
[0010] Set a first loss function in the pre-trained dialogue model, and train the pre-trained dialogue model based on the first data with labeled answers to obtain a trained dialogue model, so that the value of the first loss function is minimized;
[0011] Input the second data into the trained dialogue model to obtain a corresponding number of replies and label them with serial numbers;
[0012] Set a difference function in the pre-trained reward model, and train the pre-trained reward model based on the corresponding number of replies with labeled serial numbers to obtain a trained reward model, so that the value of the difference function is maximized;
[0013] Set a second loss function according to the trained dialogue model, and obtain an optimized dialogue model based on the third data through a reinforcement learning algorithm.
[0014] In one embodiment, the obtaining the corresponding number of replies and labeling them with serial numbers includes:
[0015] According to a preset rule, sort the corresponding number of replies in descending order according to the degree of correctness and label them with the corresponding serial numbers;
[0016] Wherein, the degree of correctness refers to the degree of closeness to the answer.
[0017] In one embodiment, the setting the second loss function according to the trained dialogue model and obtaining an optimized dialogue model based on the third data through a reinforcement learning algorithm includes:
[0018] Obtain the corresponding reward value function according to the trained reward model;
[0019] Set a second loss function according to the trained dialogue model, and adjust the corresponding reward value function according to the second loss function to obtain an adjusted reward value function;
[0020] Obtain an adjusted reward model according to the adjusted reward value function;
[0021] Input the third data into the trained dialogue model to output a reply result;
[0022] Input the reply result into the adjusted reward model, and output a reward value according to the adjusted reward value function;
[0023] Update the trained dialogue model according to the reward value to obtain an optimized dialogue model;
[0024] Wherein, the second loss function represents the similarity between the optimized dialogue model and the trained dialogue model.
[0025] In one embodiment, the first loss function represents the similarity between the reply obtained by inputting the first data into the pre-trained dialogue model and the answer annotated with the first data.
[0026] In one embodiment, the updating the trained dialogue model according to the reward value to obtain an optimized dialogue model includes:
[0027] Update the trained dialogue model by the gradient descent method according to the magnitude of the reward value to obtain an optimized dialogue model.
[0028] In one embodiment, the second loss function includes relative entropy divergence, and the reinforcement learning algorithm includes proximal policy optimization algorithm.
[0029] In one embodiment, the pre-trained dialogue model includes a multi-head attention layer and a feed-forward neural network layer, and the feed-forward neural network layer performs a non-linear transformation on the output of the multi-head attention layer.
[0030] In a second aspect, a dialogue model optimization device is provided, and the device includes:
[0031] An acquisition and division module, configured to acquire question data input in a pre-trained dialogue model through an application programming interface, and divide the question data into three parts according to a preset ratio, namely first data, second data, and third data;
[0032] The first setting and training module is used to set a first loss function in the pre-trained dialogue model and train the pre-trained dialogue model based on the first data with labeled answers to obtain a trained dialogue model, minimizing the value of the first loss function;
[0033] The input acquisition module is used to input the second data into the trained dialogue model to obtain several corresponding replies and label them with serial numbers;
[0034] The second setting and training module is used to set a difference function in the pre-trained reward model and train the pre-trained reward model based on several corresponding replies with labeled serial numbers to obtain a trained reward model, maximizing the value of the difference function;
[0035] The setting acquisition module is used to set a second loss function according to the trained dialogue model and obtain an optimized dialogue model through a reinforcement learning algorithm based on the third data.
[0036] In a third aspect, a computer device is provided, which includes one or more processors; and a memory associated with the one or more processors, where the memory is used to store program instructions, and when the program instructions are read and executed by the one or more processors, the steps of the dialogue model optimization method according to any one of the above first aspects are executed.
[0037] In a fourth aspect, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the dialogue model optimization method according to any one of the above first aspects are executed.
[0038] In the above dialogue model optimization method, device, computer device, and storage medium, by setting the first loss function and the difference function, and based on the first data with labeled answers and several replies with standard serial numbers, the pre-trained dialogue model and the pre-trained reward model are respectively trained to obtain a trained dialogue model and a trained reward model. According to the set second loss function and the third data, an optimized dialogue model is obtained through a reinforcement learning algorithm, achieving the improvement of the dialogue quality of the intelligent customer service dialogue model while avoiding the generation of abnormal results. Description of the Drawings
[0039] Figure 1 It is a schematic flowchart of the dialogue model optimization method in an embodiment;
[0040] Figure 2 It is a structural block diagram of the dialogue model optimization device in an embodiment;
[0041] Figure 3Internal structure diagram of a computer device in an embodiment. Detailed implementation
[0042] Reinforcement learning is different from deep learning. The problem discussed in reinforcement learning (RL) is how an agent can maximize the rewards it can obtain in a complex and uncertain environment. Reinforcement learning consists of two parts: the agent and the environment. During the reinforcement learning process, the agent and the environment are constantly interacting. After the agent obtains a certain state in the environment, it will use this state to output an action, which is also called a decision. Then this action will be executed in the environment, and the environment will output the next state and the reward brought by the current action according to the action taken by the agent. The purpose of the agent is to obtain as many rewards as possible from the environment. The following takes playing video games as an example to illustrate. Suppose there is a video game where you need to control a small ball to reach the other end from one end and avoid obstacles. You can regard this problem as a reinforcement learning task. In this task, the small ball is the intelligent agent, the game map is the environment, and the small ball can obtain rewards or punishments during the process of moving from one point to the end point. The rules and objectives of the game are as follows:
[0043] - The starting position of the small ball is fixed, and there are 4 optional directions for each move, namely up, down, left, and right.
[0044] - If the small ball hits an obstacle, it will be punished to encourage it to avoid hitting the obstacle.
[0045] - If the small ball walks from the starting point to the end point, it will obtain a reward to encourage it to reach the end point.
[0046] To train the small ball to complete this task, we can use the reinforcement learning method. In reinforcement learning, the intelligent agent takes actions by interacting with the environment and tries to maximize the cumulative reward to learn the best strategy.
[0047] In order to make the purpose, technical solution and advantages of this application clearer, the following further details this application in combination with the accompanying drawings and embodiments. It should be understood that the specific embodiments described here are only used to explain this application and are not used to limit this application.
[0048] Embodiment 1
[0049] In one embodiment, as Figure 1 shown, a method for optimizing a dialogue model is provided. The method includes the following steps:
[0050] S10. Collect the question data input into the pre-trained dialogue model through the application programming interface, and divide the question data into three parts according to a preset ratio, namely the first data, the second data, and the third data;
[0051] Further, the preset ratio is the first data: the second data: the third data = 4:2:4. The pre-trained dialogue model is trained with a large amount of data and the final optimized dialogue model is obtained through the reinforcement learning algorithm, so as to improve the dialogue quality.
[0052] Further, the pre-trained dialogue model includes a multi-head attention layer and a feed-forward neural network layer, and the feed-forward neural network layer performs a non-linear transformation on the output of the multi-head attention layer.
[0053] Furthermore, the pre-trained dialogue model includes the Source 1.0 model. In the Source 1.0 model, the basic model structure consists of an embedding layer, a multi-head attention layer, and a feed-forward neural network layer. The multi-head attention layer allows the model to interact and focus among inputs at different positions, and the feed-forward neural network layer then performs a non-linear transformation on the output of the multi-head attention layer. Between each sub-layer, residual connections and layer normalization are also added. In the multi-head attention layer, the input sequence is divided into multiple heads, and each head learns a different representation. Then, each head applies a weighted function similar to the multi-head attention layer to determine the importance of each position to other positions. This way allows the model to efficiently process long sequences. In the feed-forward neural network layer, the model inputs the output of the multi-head attention layer into a fully connected neural network to learn the non-linear relationship between feature representations. The final output is composed of multiple levels and generates the target sequence through the decoder or serves as the output for classification or regression tasks.
[0054] Moreover, in the Source 1.0 model, we used a variety of parallel algorithms, including data parallelism, tensor parallelism, and pipeline parallelism, etc., and achieved good parallel efficiency on super-large servers. At the same time, due to the increasing number of model parameters, a single server can no longer store a model of such a large size, so a good model parallel strategy is indispensable.
[0055] S11. Set a first loss function in the pre-trained dialogue model, and train the pre-trained dialogue model based on the first data with labeled answers to obtain a trained dialogue model, so that the value of the first loss function is minimized;
[0056] Specifically, the trained dialogue model can generate higher-quality answers when conducting intelligent conversations.
[0057] Further, the first loss function represents the similarity between the reply obtained by inputting the first data into the pre-trained dialogue model and the answer labeled for the first data.
[0058] Specifically, by training the pre-trained dialogue model, the responses output by the dialogue model are made closer to the answers.
[0059] S12. Input the second data into the trained dialogue model to obtain a corresponding number of responses and label them with serial numbers.
[0060] Further, the corresponding number of responses generally refers to 4-9 true responses.
[0061] S13. Set a difference function in the pre-trained reward model, and train the pre-trained reward model based on the corresponding number of responses with labeled serial numbers to obtain a trained reward model, making the difference function value the largest.
[0062] Specifically, the scoring result of the trained reward model for the output response, that is, the reward value, is closer to the artificial scoring standard.
[0063] Further, the obtaining of the corresponding number of responses and labeling them with serial numbers includes:
[0064] According to the preset rules, sort the corresponding number of responses in descending order according to the degree of correctness and label them with the corresponding serial numbers.
[0065] Among them, the degree of correctness refers to the degree of closeness to the answer.
[0066] Specifically, sort them in descending order according to the degree of closeness of the response to the answer, and label the corresponding replies with serial numbers, so that the reward value output by the trained reward model is close to the artificial scoring standard, where the serial number is a positive integer.
[0067] Further, the preset rules include whether it conforms to the facts, whether the format specification is correct, and the degree of detail of the answer. Label the best response result with serial number 1, the second best response result with serial number 2, and so on.
[0068] Even further, the difference function represents the similarity between the response result with serial number 1 and the response result with serial number 2, which can be measured by the reward value.
[0069] S14. Set a second loss function according to the trained dialogue model, and obtain an optimized dialogue model based on the third data through the reinforcement learning algorithm.
[0070] Further, after 2 epochs (a complete training), the reinforcement learning is completed, or it can also be other preset numbers of times, which can improve the dialogue quality of the optimized dialogue model while avoiding the singleness of the output response caused by overfitting.
[0071] Further, the second loss function includes relative entropy divergence, and the reinforcement learning algorithm includes the Proximal Policy Optimization (PPO) algorithm.
[0072] Further, setting the second loss function according to the trained dialogue model and obtaining an optimized dialogue model through a reinforcement learning algorithm based on the third data includes:
[0073] Obtaining a corresponding reward value function according to the trained reward model;
[0074] Setting the second loss function according to the trained dialogue model and adjusting the corresponding reward value function according to the second loss function to obtain an adjusted reward value function;
[0075] Obtaining an adjusted reward model according to the adjusted reward value function;
[0076] Inputting the third data into the trained dialogue model to output a reply result;
[0077] Inputting the reply result into the adjusted reward model and outputting a reward value according to the adjusted reward value function;
[0078] Updating the trained dialogue model according to the reward value to obtain an optimized dialogue model;
[0079] Wherein, the second loss function represents the similarity degree between the optimized dialogue model and the trained dialogue model.
[0080] Further, updating the trained dialogue model according to the reward value to obtain an optimized dialogue model includes:
[0081] Updating the trained dialogue model by the gradient descent method according to the magnitude of the reward value to obtain an optimized dialogue model.
[0082] That is, if the reward value result is low, the trained dialogue model is adjusted by offset, and the offset can adopt the gradient descent method.
[0083] Specifically, through the second loss function and the reinforcement learning algorithm, an optimized dialogue model is obtained, realizing the improvement of the dialogue quality of the dialogue model and avoiding the similarity degree between the optimized dialogue model and the trained dialogue model being too large, resulting in a deformed result.
[0084] Further, subtracting the relative entropy divergence from the corresponding reward value function to obtain an adjusted reward value function, which is specifically expressed as:
[0085]
[0086] Among them, since the responses output by the dialogue model for a problem are output one by one, refers to the reward value for obtaining the complete response result, x and y represent two adjacent output characters, x is output before y, and E (x,y) refers to the expected value of obtaining the complete response result, r θ (x, y) represents the reward value for both outputting x and outputting y, β is a preset parameter that can be adjusted according to specific circumstances, π refers to the probability of an event, RL refers to the optimized dialogue model, SFT refers to the trained dialogue model, y|x refers to the situation of outputting y based on outputting x, and π RL (y|x) is the probability of the optimized dialogue model outputting y based on outputting x, and π SFT (y|x) is the probability of the trained dialogue model outputting y based on outputting x.
[0087] It should be understood that although Figure 1 the steps in the flowchart of are shown in sequence according to the arrows, these steps do not necessarily have to be executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, Figure 1 at least a part of the steps in may include multiple sub-steps or multiple stages. These sub-steps or stages do not necessarily have to be executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages does not necessarily have to be sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.
[0088] Embodiment 2
[0089] In one embodiment, as Figure 2 shown, a dialogue model optimization device is provided, and the device includes:
[0090] A collection and division module 20, which is used to collect the input problem data in the pre-trained dialogue model through an application programming interface, and divide the problem data into three parts according to a preset ratio, namely the first data, the second data, and the third data;
[0091] Furthermore, the pre-trained dialogue model includes a multi-head attention layer and a feed-forward neural network layer, and the feed-forward neural network layer performs a non-linear transformation on the output of the multi-head attention layer.
[0092] The first setting and training module 21 is configured to set a first loss function in the pre-trained dialogue model and train the pre-trained dialogue model based on the first data with labeled answers to obtain a trained dialogue model, minimizing the value of the first loss function;
[0093] Further, the first loss function represents the similarity between the reply obtained by inputting the first data into the pre-trained dialogue model and the answer labeled for the first data.
[0094] The input acquisition module 22 is configured to input the second data into the trained dialogue model to obtain a corresponding number of replies and label them with serial numbers;
[0095] The second setting and training module 23 is configured to set a difference function in the pre-trained reward model and train the pre-trained reward model based on the corresponding number of replies with labeled serial numbers to obtain a trained reward model, maximizing the value of the difference function;
[0096] Further, the second setting and training module is further configured to:
[0097] Sort the corresponding number of replies in descending order according to the correctness degree according to a preset rule and label them with corresponding serial numbers;
[0098] Wherein, the correctness degree refers to the degree of proximity to the answer.
[0099] The setting and acquisition module 24 is configured to set a second loss function according to the trained dialogue model and obtain an optimized dialogue model based on the third data through a reinforcement learning algorithm.
[0100] Further, the second loss function includes relative entropy divergence, and the reinforcement learning algorithm includes the proximal policy optimization algorithm.
[0101] Further, the setting and acquisition module is further configured to:
[0102] Obtain a corresponding reward value function according to the trained reward model;
[0103] Set a second loss function according to the trained dialogue model and adjust the corresponding reward value function according to the second loss function to obtain an adjusted reward value function;
[0104] Obtain an adjusted reward model according to the adjusted reward value function;
[0105] Input the third data into the trained dialogue model and output a reply result;
[0106] Input the reply result into the adjusted reward model, and output a reward value according to the adjusted reward value function;
[0107] Update the trained dialogue model according to the reward value to obtain an optimized dialogue model;
[0108] Wherein, the second loss function represents the similarity between the optimized dialogue model and the trained dialogue model.
[0109] Further, the setting and obtaining module is further configured to:
[0110] Update the trained dialogue model by the gradient descent method according to the magnitude of the reward value to obtain an optimized dialogue model.
[0111] For the specific limitations of the dialogue model optimization device, reference may be made to the limitations of the dialogue model optimization method in the foregoing text, which will not be elaborated herein. Each module in the above dialogue model optimization device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or independent thereof, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.
[0112] Embodiment III
[0113] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the following steps are implemented:
[0114] Collect the question data input in the pre-trained dialogue model through the application programming interface, and divide the question data into three parts according to a preset ratio, namely first data, second data, and third data;
[0115] Set a first loss function in the pre-trained dialogue model, and train the pre-trained dialogue model based on the first data with labeled answers to obtain a trained dialogue model, so that the value of the first loss function is minimized;
[0116] Input the second data into the trained dialogue model to obtain a corresponding number of replies and label them with serial numbers;
[0117] Set a difference function in the pre-trained reward model, and train the pre-trained reward model based on the corresponding number of replies with labeled serial numbers to obtain a trained reward model, so that the value of the difference function is maximized;
[0118] Set a second loss function according to the trained dialogue model, and obtain an optimized dialogue model based on the third data through a reinforcement learning algorithm.
[0119] When the program instructions are read and executed by the one or more processors, they can also perform operations corresponding to the respective steps in the above method embodiments. Reference can be made to the descriptions in the foregoing text, and details are not repeated here. Figure 3 , which exemplarily shows the architecture of a computer device. Specifically, it may include a processor 310, a video display adapter 311, a disk drive 312, an input / output interface 313, a network interface 314, and a memory 320. The above-mentioned processor 310, video display adapter 311, disk drive 312, input / output interface 313, network interface 314, and the memory 320 can be communicatively connected through a communication bus 330.
[0120] Among them, the processor 310 can be implemented in a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided by this application.
[0121] The memory 320 can be implemented in the form of a read-only memory (ROM), a random access memory (RAM), a static storage device, a dynamic storage device, etc. The memory 320 can store an operating system 321 for controlling the operation of the computer device 300, and a basic input / output system (BIOS) 322 for controlling the low-level operations of the computer device 300. In addition, a web browser 323, a data storage management 324, an icon font processing system 325, etc. can also be stored. The above-mentioned icon font processing system 325 can be the application program that specifically implements the foregoing respective step operations in the embodiments of this application. In short, when implementing the technical solutions provided by this application through software or firmware, the relevant program codes are stored in the memory 320 and are called and executed by the processor 310.
[0122] The input / output interface 313 is used to connect to an input / output module to implement information input and output. The input / output module can be configured as a component in the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Among them, the input device can include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device can include a display, a speaker, a vibrator, an indicator light, etc.
[0123] The network interface 314 is used to connect to a communication module (not shown in the figure) to enable communication and interaction between this device and other devices. The communication module can achieve communication through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).
[0124] The bus 330 includes a path for transmitting information between various components of the device (such as the processor 310, video display adapter 311, disk drive 312, input / output interface 313, network interface 314, and memory 320).
[0125] In addition, the computer device 300 can also obtain information on specific collection conditions from the conditional information database 341 of the virtual resource object for conditional judgment, etc.
[0126] It should be noted that although the above computer device 300 only shows the processor 310, video display adapter 311, disk drive 312, input / output interface 313, network interface 314, memory 320, bus 330, etc., in the specific implementation process, the computer device may also include other components necessary for normal operation. In addition, those skilled in the art can understand that the above device may also only include the components necessary to implement the solution of this application, and does not necessarily include all the components shown in the figure.
[0127] From the description of the above embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, cloud server, or network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of this application.
[0128] Embodiment 4
[0129] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0130] Collect the question data input in the pre-trained dialogue model through the application programming interface, and divide the question data into three parts according to a preset ratio, namely the first data, the second data, and the third data;
[0131] Set a first loss function in the pre-trained dialogue model, and train the pre-trained dialogue model based on the first data with labeled answers to obtain a trained dialogue model, so that the value of the first loss function is minimized;
[0132] Input the second data into the trained dialogue model to obtain several corresponding replies and label them with serial numbers;
[0133] Set a difference function in the pre-trained reward model, and train the pre-trained reward model based on the several replies corresponding to the labeled serial numbers to obtain a trained reward model, so that the value of the difference function is maximized;
[0134] Set a second loss function according to the trained dialogue model, and obtain an optimized dialogue model based on the third data through a reinforcement learning algorithm.
[0135] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0136] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0137] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation to the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.
Claims
1. A method for optimizing a dialogue model, characterized in that, the method includes: Collecting the question data input in the pre-trained dialogue model through an application programming interface, and dividing the question data into three parts according to a preset ratio, namely the first data, the second data, and the third data; Setting a first loss function in the pre-trained dialogue model, and training the pre-trained dialogue model based on the first data with labeled answers to obtain a trained dialogue model, so that the value of the first loss function is minimized; Inputting the second data into the trained dialogue model to obtain a corresponding number of replies and numbering them; Setting a difference function in the pre-trained reward model, and training the pre-trained reward model based on the corresponding number of replies with numbered labels to obtain a trained reward model, so that the value of the difference function is maximized; Setting a second loss function according to the trained dialogue model, and obtaining an optimized dialogue model through a reinforcement learning algorithm based on the third data, which specifically includes: Obtaining a corresponding reward value function according to the trained reward model; Setting a second loss function according to the trained dialogue model, and adjusting the corresponding reward value function according to the second loss function to obtain an adjusted reward value function; Obtaining an adjusted reward model according to the adjusted reward value function; Inputting the third data into the trained dialogue model to output a reply result; Inputting the reply result into the adjusted reward model, and outputting a reward value according to the adjusted reward value function; Updating the trained dialogue model according to the reward value to obtain an optimized dialogue model; wherein, the second loss function represents the similarity between the optimized dialogue model and the trained dialogue model.
2. The method according to claim 1, characterized in that, the obtaining the corresponding number of replies and numbering them includes: Sorting the corresponding number of replies in descending order according to the correct degree according to a preset rule, and labeling the corresponding serial numbers; wherein, the correct degree refers to the degree of closeness to the answer.
3. The method according to claim 2, characterized in that, the first loss function represents the similarity between the reply obtained by inputting the first data into the pre-trained dialogue model and the answer labeled by the first data.
4. The method according to claim 1, characterized in that, the updating the trained dialogue model according to the reward value to obtain an optimized dialogue model includes: Updating the trained dialogue model by the gradient descent method according to the magnitude of the reward value to obtain an optimized dialogue model.
5. The method according to claim 1, characterized in that, the second loss function includes relative entropy divergence, and the reinforcement learning algorithm includes proximal policy optimization algorithm.
6. The method according to claim 1, characterized in that, the pre-trained dialogue model includes a multi-head attention layer and a feed-forward neural network layer, and the feed-forward neural network layer performs a non-linear transformation on the output of the multi-head attention layer.
7. A dialogue model optimization device for implementing the method described in claim 1, characterized in that, the device comprises: a collection and division module, configured to collect the input question data in the pre-trained dialogue model through an application programming interface, and divide the question data into three parts according to a preset ratio, namely first data, second data, and third data; a first setting and training module, configured to set a first loss function in the pre-trained dialogue model, and train the pre-trained dialogue model based on the first data with labeled answers to obtain a trained dialogue model, such that the value of the first loss function is minimized; an input acquisition module, configured to input the second data into the trained dialogue model to obtain a corresponding number of replies and label the serial numbers; a second setting and training module, configured to set a difference function in the pre-trained reward model, and train the pre-trained reward model based on the corresponding number of replies with labeled serial numbers to obtain a trained reward model, such that the value of the difference function is maximized; a setting and acquisition module, configured to set a second loss function according to the trained dialogue model, and obtain an optimized dialogue model based on the third data through a reinforcement learning algorithm.
8. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, when the processor executes the computer program, the steps of the dialogue model optimization method described in any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium, having a computer program stored thereon, characterized in that, when the computer program is executed by a processor, the steps of the dialogue model optimization method described in any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Model training method, device, terminal and storage medium based on data processing
CN109460463A
Conversation method and device, electronic equipment and readable storage medium
CN111753076A