A dialogue management model construction method based on a deep reinforcement learning A3C algorithm
Patent Information
- Application Number
- CN202210562268.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-23
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2042-05-23
AI Technical Summary
对话管理是任务型对话系统中的一个重要模块,其根据多轮对话的对话状态选择对应的对话回复策略,是任务型对话系统的核心模块,然而现有的对话管理模型训练效率低、准确率低
[0030]与现有技术相比,有益效果在于:1)本发明的基于深度强化学习A3C算法的对话管理模型构建方法,通过构建多个子线程网络(Actor策略网络和Critic价值网络)、全局网络,能够对Actor策略网络输出的策略进行动态地评估,通过Critic价值网络的评价能力,在训练时及时调整策略,实现了对话管理模型内部能够动态地改善对话策略的学习好坏,提升数据模型的准确率;同时以多线程地方式训练多个子线程网络,通过全局网络共享参数、分发参数,提升模型的训练效率。
Smart Images

Figure CN114818613B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of deep learning technology, and in particular to a method for constructing a dialogue management model based on the deep reinforcement learning A3C algorithm. Background Technology
[0002] Deep reinforcement learning is a technique that combines deep learning and reinforcement learning. It leverages the perceptual capabilities of deep learning and the decision-making abilities of reinforcement learning to select and optimize strategies based on data characteristics, ultimately achieving the optimal strategy. Dialogue management is a crucial module in task-oriented dialogue systems. It selects corresponding dialogue response strategies based on the dialogue states of multiple turns and is a core module of such systems. However, existing dialogue management models suffer from low training efficiency and accuracy.
[0003] Therefore, it is necessary to provide a novel method for constructing dialogue management models based on the deep reinforcement learning A3C algorithm to overcome the above-mentioned shortcomings. Summary of the Invention
[0004] The purpose of this invention is to provide a method for constructing a dialogue management model based on the deep reinforcement learning A3C algorithm, which can improve the accuracy and training efficiency of the data model.
[0005] To achieve the above objectives, this invention provides a method for constructing a dialogue management model based on the deep reinforcement learning A3C algorithm, comprising the following steps:
[0006] S1, Vector transformation;
[0007] S2. Define the TD error and set the dialogue reward function;
[0008] S3. Construct a Critic value network;
[0009] S4. Define the user simulator;
[0010] S5. The global network stores the parameters in the Actor policy network and the Critic value network, and updates the parameters in the global network composed of the Actor policy network and the Critic value network.
[0011] S6, Training Data Model.
[0012] Preferably, one-hot encoding is used to convert the dialogue state and dialogue policy into vectors; in step S1, the dialogue state and policy text information is converted into an information form that can be accepted by the neural network model, that is, converted into a feature vector X that can be accepted by the neural network model. i (i = 1, 2, 3…n).
[0013] Preferably, step S2 includes step S21: based on the existence of three states in a multi-round dialogue—success, failure, and incomplete—set a dialogue reward function, use the value of each round's dialogue strategy and the set reward function rules to construct a custom TD error, and define the value of the dialogue state and the initial selection probability of the dialogue strategy.
[0014] Preferably, the reward function is set as follows:
[0015]
[0016] Step S202: Set the current round's dialogue state as state S. t The selected dialogue strategy is taken as action A. t A t Once selected, it enters state S. t+1 Let W be the probability of each action being selected in state S. i , i = 1, 2, ..., n, E w Let V be the average value of dialogue actions selected in this state, then the value V corresponding to each state is:
[0017]
[0018] The multi-turn dialogue reward function R(A) t ,S t ) and dialogue state value V(S) t A t Combined with the construction of a custom TD error:
[0019] TD = R(A) t ,S t )+γV(S t+1 A t+1 )-V(S t A t (2);
[0020] Where γ is the discount factor. After receiving the TD error using the Actor policy network constructed with an LSTM network, the corresponding loss function L is constructed as follows:
[0021] L(θ A ) = [log W(A t ,S t ;θ A )×TD] 2 (3);
[0022] Meanwhile, the Actor policy network continuously selects the corresponding dialogue policy to maximize the accumulated reward value, calculated as follows:
[0023] Q(S,A)=maxπE[R(A t ,S t )+γR(A t+1 ,S t+1 )+γ2R(At+2,St+2)+…|At=A,St=S](4).
[0024] Preferably, step S3 includes step S31, where the input to the Critic value network is the state, and the output is the TD error. The TD error is calculated using the formula: TD = R(A t ,S t )+γV(S t+1 A t+1 )-V(S t A t (5).
[0025] Preferably, step S4 includes step S41: determining the corresponding slot fields based on dialogues in different domains, constructing a slot information collection library for different domains, and having the user send the trained message to the server.
[0026] S42. Build a user intent library; In multi-turn dialogues, users have various intents and different needs. Build a user intent library and a user simulator to support the simulation of various user intents.
[0027] S43. Define user reply templates; define different reply phrases for different system replies, so that the user simulator can dynamically fill information according to the current dialogue slots and select the corresponding reply.
[0028] Preferably, step S6 further includes step S61: using web crawling technology to crawl question-answer pair sets for different scenarios, and using manual annotation to construct multi-turn dialogue corpus from the question-answer pairs to obtain a multi-turn dialogue set;
[0029] S62, The user simulator interacts with the data model; the user simulator displays the current dialogue state S t The input is passed to the Actor policy network and the Critic value network. The Actor policy network then processes the input dialogue state S. t Generate a distribution of dialogue policy probabilities, select the dialogue policy 'a' with the highest probability, and the user simulator will then use S. t The Critic value network is input to obtain the TD error, and the parameters of the Critic value network are updated. Then the Actor policy network receives the TD error and outputs the dialogue policy a to the user simulator. The user simulator outputs the new dialogue state according to the rules until the iteration is complete.
[0030] Compared with existing technologies, the beneficial effects are as follows: 1) The dialogue management model construction method based on the deep reinforcement learning A3C algorithm of the present invention can dynamically evaluate the policy output by the Actor policy network by constructing multiple sub-thread networks (Actor policy network and Critic value network) and a global network. Through the evaluation capability of the Critic value network, the policy can be adjusted in time during training, thereby realizing the ability of the dialogue management model to dynamically improve the learning quality of the dialogue policy and improve the accuracy of the data model. At the same time, multiple sub-thread networks are trained in a multi-threaded manner, and parameters are shared and distributed through the global network, thereby improving the training efficiency of the model.
[0031] 2) With the interaction between the global network and the sub-thread networks, the global network first passes its initialized network parameters to each sub-thread network. Then, each sub-thread network begins to interact, trains independently, and updates its parameters. Finally, the Critic value network and Actor policy network parameters of each sub-thread network are asynchronously updated to the global network. This approach enables the data model to converge faster and more quickly. Attached Figure Description
[0032] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0033] Figure 1 The flowchart shows the method for constructing a dialogue management model based on the deep reinforcement learning A3C algorithm provided by this invention.
[0034] Figure 2 This is a schematic diagram of the global network structure of the dialogue management model construction method based on the deep reinforcement learning A3C algorithm provided by the present invention. Detailed Implementation
[0035] To make the objectives, technical solutions, and beneficial effects of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described in this specification are merely illustrative of the invention and are not intended to limit the invention.
[0036] Please see Figures 1 to 2 This invention provides a method for constructing a dialogue management model based on the deep reinforcement learning A3C algorithm, comprising the following steps:
[0037] S1, Vector transformation;
[0038] Specifically, one-hot encoding is used to convert the dialogue state and dialogue policy into vectors; in step S1, the dialogue state and policy text information is converted into an information form that the neural network model can accept, that is, into a feature vector X that the neural network model can accept. i (i = 1, 2, 3…n).
[0039] S2. Define the TD error and set the dialogue reward function;
[0040] Specifically, step S2 includes step S21: based on the existence of three states in a multi-turn dialogue—success, failure, and incomplete—set a dialogue reward function, use the value of each round of dialogue strategy and the set reward function rules to construct a custom TD error, and define the value of the dialogue state and the initial selection probability of the dialogue strategy.
[0041] Step S21 further includes step S201. Dialogue failure refers to a situation where, during the matching of slots, the slots do not conform to the rules, or the system fails to recognize the user's true intent or slot information, and the number of dialogue rounds is less than the maximum allowed number of rounds, causing a deviation from the correct dialogue flow. Dialogue success, on the other hand, means that the system has successfully matched all the slots in the dialogue and has completed the dialogue task.
[0042] The specific reward function is set as follows:
[0043]
[0044] S202. Set the current round's dialogue state as state S. t The selected dialogue strategy is taken as action A. t A t Once selected, it enters state S. t+1 Let W be the probability of each action being selected in state S. i , i = 1, 2, ..., n, E w Let V be the average value of dialogue actions selected in this state, then the value V corresponding to each state is:
[0045]
[0046] Simultaneously, the multi-round dialogue reward function R(A) will be applied. t ,S t ) and dialogue state value V(S) t A t Combined with the construction of a custom TD error:
[0047] TD = R(A) t ,S t )+γV(S t+1 A t+1 )-V(St A t (2);
[0048] Where γ is the discount factor, the Actor policy network constructed using an LSTM network receives the TD error and constructs its corresponding loss function L, which is as follows:
[0049] L(θ A ) = [log W(A t ,S t ;θ A )×TD] 2 (3);
[0050] Meanwhile, the Actor policy network continuously selects the corresponding dialogue policy to maximize the accumulated reward value, as shown in the following formula:
[0051] Q(S,A)=maxπE[R(A t ,S t )+γR(A t+1 ,S t+1 )+γ2R(At+2,St+2)+…|At=A,St=S](4);
[0052] The specific process of constructing the Actor policy network using an LSTM network is as follows: the input of the Actor policy network is the dialogue state of the current round, and the output is the probability distribution of the actions. The network includes an input layer, an output layer, and a hidden layer. In this paper, 64 hidden units are selected. The neural network predicts the probability value corresponding to each dialogue policy by inputting the feature vector Xi, i.e., the dialogue state vector obtained by S1, and returns the probability distribution of each action.
[0053] For example, a user asks: "I want to apply for an ID card." Based on its neural network learning, the system provides the following three probability distributions for responses: Statement 1: Do you want to apply for an ID card? Statement 2: Please repeat. Statement 3: I don't know. The probabilities of the three responses are 0.7, 0.2, and 0.1, respectively. Therefore, the Actor policy network selects the response with the highest probability, i.e., Statement 1, and outputs it to the user.
[0054] S3. Construct a Critic value network;
[0055] Specifically, step S3 includes step S31, where the input to the Critic value network is the state, and the output is the TD error. The formula for calculating the TD error is: TD = R(A t ,S t )+γV(S t+1 A t+1 )-V(S t A t(5);
[0056] S4. Define the user simulator;
[0057] Specifically, step S4 includes step S41: determining the corresponding slot fields based on dialogues in different domains, constructing a slot information collection library for different domains, and sending the trained message to the server.
[0058] Specifically, in the process of obtaining a property ownership certificate, the slot fields would include the property area and whether it is the individual's first residential property. Similarly, in the process of obtaining an ID card, the slot fields would include whether the applicant is a local resident and whether the application is for a lost or newly issued ID card. Different fields require different fields to be collected, and this part needs to be defined manually.
[0059] S42. Build a user intent library;
[0060] Specifically, users have various intentions and needs in multi-turn dialogues. A user intent library is built, and a user simulator supports simulating various user intents.
[0061] S43. Define user response templates;
[0062] Specifically, different response utterances need to be defined for different system responses. The user simulator can dynamically fill information based on the current dialogue slots and select the corresponding response. This invention artificially constructs various types of response corpora, providing more than 300 templates, including three types of expressions: positive, ask, and negative.
[0063] S5. The global network stores the parameters in the Actor policy network and the Critic value network, and updates the parameters to the global network, which consists of an Actor policy network and a Critic value network, at regular intervals.
[0064] S6, Training Data Model;
[0065] Step S6 also includes step S61, which uses web crawling technology to crawl a set of question-answer pairs covering 14 different scenarios, including housing provident fund processing, real estate processing, business license processing, social insurance, and household registration processing, containing a total of 1004 question-answer pairs. At the same time, using manual annotation, the question-answer pairs are constructed into a multi-turn dialogue corpus, resulting in a set of 206 multi-turn dialogues. The example dialogue flow of the dataset is as follows:
[0066] User: I want to transfer my household registration.
[0067] System: Where is your original household registration address?
[0068] User: No. XX, XX District, XX City, XX Province.
[0069] System: May I ask what address you moved out of?
[0070] User: XX Province, XX City, XX Road
[0071] System: Sorry, please state your full address from which you moved out.
[0072] User: No. XX, XX Road, XX District, XX City, XX Province.
[0073] System: Your situation involves relocation within the same area of the central urban area, development zone, and new urban area.
[0074] System: Please bring the following documents to the Public Security Bureau's Household Registration Window on the third floor: 1. A completed "Household Registration Application Form" by the applicant or their legal guardian. 2. The applicant's household registration booklet and resident ID card. 3. Property ownership certificate. 4. If moving to the spouse's address, a marriage certificate, spouse's household registration booklet and resident ID card are required; if children are moving with the applicant, a "Birth Certificate" is also required. If moving to the location of the parents' or children's household registration, the household registration booklet of the new address, the parents' or children's resident ID cards, a "Birth Certificate," or proof of parent-child relationship is required.
[0075] S62. User simulator interacts with data model: In a certain dialogue state S t Next, the user simulator will display the current dialogue state S. t The input is passed to the Actor policy network and the Critic value network. The Actor policy network then processes the input dialogue state S. t Generate a distribution of dialogue policy probabilities, select the dialogue policy 'a' with the highest probability, and the user simulator will then use S. t The Critic value network is input to obtain the TD error, and the parameters of the Critic value network are updated. Then the Actor policy network receives the TD error and outputs the dialogue policy a to the user simulator. The user simulator outputs a new dialogue state according to the rules. Finally, the above process is repeated until the iteration is completed.
[0076] The technical solutions provided in the embodiments of this specification enable the dialogue management model to dynamically improve the learning quality of dialogue policies and enhance the accuracy of the data model by constructing multiple sub-thread networks (Actor policy network and Critic value network) and a global network. At the same time, multiple sub-thread networks are trained in a multi-threaded manner, and parameters are shared and distributed through the global network to improve the training efficiency of the model.
[0077] In summary, this method evaluates the accuracy of the dialogue management model using two indicators: average number of dialogue rounds and dialogue success rate. For training efficiency, this method uses the training time required to reach the same number of rounds as the evaluation criterion.
[0078] This method can dynamically evaluate the policies output by the Actor policy network and adjust the policies in a timely manner during training by leveraging the evaluation capabilities of the Critic value network, thereby ensuring the accuracy of the trained model.
[0079] Through the interaction between the global network and the sub-thread networks, the global network first passes its initialized network parameters to each sub-thread network. Then, each sub-thread network begins to interact, trains independently, and updates its parameters. Finally, the Critic value network and Actor policy network parameters of each sub-thread network are asynchronously updated to the global network. This approach enables the data model to converge faster and more effectively.
[0080] The present invention is not limited to the description in the specification and embodiments, and thus other advantages and modifications can be readily realized by those skilled in the art. Therefore, the present invention is not limited to the specific details, representative devices and examples shown and described herein without departing from the spirit and scope of the general concept as defined by the claims and their equivalents.
Claims
1. A method for constructing a dialogue management model based on the deep reinforcement learning A3C algorithm, characterized in that, Includes the following steps: S1. Vector Conversion: One-hot encoding is used to convert the dialogue state and dialogue strategy into vectors. In step S1, the dialogue state and strategy text information is converted into an information form acceptable to the neural network model, i.e., converted into feature vectors acceptable to the neural network model. X i ,i =1, 2, 3…n; S2. Define the TD error and set the dialogue reward function; Step S2 includes step S21: Based on the existence of three states in a multi-turn dialogue—success, failure, and incomplete—a dialogue reward function is set. A custom TD error is constructed using the value of each round's dialogue strategy and the set reward function rules. The value of the dialogue state and the initial selection probability of the dialogue strategy are defined. The set reward function is as follows: ; Step S202: Set the current round's dialogue state as the state. S t The chosen dialogue strategy is treated as an action. A t , A t Once selected, the state is entered. S t+1 Let the state be S The probability of each action being selected is W i , i =1, 2, ..., n , E w The value corresponding to each state is the average value of the dialogue actions selected in that state. V for: (1); The multi-turn dialogue reward function R( A t , S t ) and dialogue state value V ( S t , A t Combined with the construction of a custom TD error: (2); in, γ The loss function L is constructed by using an Actor policy network built with an LSTM network to receive the TD error, and the loss function L is as follows: (3); Meanwhile, the Actor policy network continuously selects the corresponding dialogue policy to maximize the accumulated reward value, calculated as follows: Q(S,A)=maxπE[ + + … ] (4); S3. Construct the Critic value network; Step S3 includes step S31, where the input to the Critic value network is the state, and the output is the TD error. The formula for calculating the TD error is: (5); S4. Define the user simulator; S5. The global network stores the parameters in the Actor policy network and the Critic value network, and updates the parameters in the global network composed of the Actor policy network and the Critic value network. S6, Training Data Model.
2. The method for constructing a dialogue management model based on the deep reinforcement learning A3C algorithm as described in claim 1, characterized in that, Step S4 includes step S41: Determine the corresponding slot fields based on dialogues in different domains, build a slot information collection library for different domains, and send the trained message to the server. S42. Build a user intent library; In multi-turn dialogues, users have various intents and different needs. Build a user intent library and a user simulator to simulate various user intents. S43. Define user reply templates; define different reply phrases for different system replies, so that the user simulator can dynamically fill information according to the current dialogue slots and select the corresponding reply.
3. The method for constructing a dialogue management model based on the deep reinforcement learning A3C algorithm as described in claim 1, characterized in that, Step S6 also includes step S61, using web crawling technology to crawl question-answer pair sets for different scenarios, and using manual annotation to construct multi-turn dialogue corpus from the question-answer pairs to obtain multi-turn dialogue set; S62, The user simulator interacts with the data model; the user simulator displays the current dialogue state. S t The input is passed to the Actor policy network and the Critic value network. The Actor policy network then processes the input dialogue state... S t Generate a probability distribution for the generated dialogue strategies, select the dialogue strategy 'a' with the highest probability, and the user simulator will... S t The Critic value network is input to obtain the TD error, and the parameters of the Critic value network are updated. Then the Actor policy network receives the TD error and outputs the dialogue policy a to the user simulator. The user simulator outputs the new dialogue state according to the rules until the iteration is completed.
Citation Information
Patent Citations
A financial transaction method based on a deep reinforcement learning A3C algorithm
CN109816530A
A deep reinforcement learning countermeasure attack-oriented model enhancement defense method
CN112069504A