Conversation method, apparatus, computer device and medium
By generating a model using action sequences, the problems of high modeling costs and poor performance in existing technologies are solved, a unified response model is achieved, and response efficiency is improved.
Patent Information
- Application Number
- CN202010457373.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-05-26
- Publication Date
- 2025-12-19
- Estimated Expiration
- 2040-05-26
AI Technical Summary
The lack of a unified approach to establishing response models in existing technologies leads to high modeling costs, poor performance, and low response efficiency for different scenarios or domains.
An action sequence generation model is adopted. First, the action sequence of the response is generated, and then the response content is generated based on the action sequence. The hierarchical structure of the elements in the response is preserved, and a unified response model is established.
It reduces modeling costs, improves modeling performance and response efficiency, and can provide a unified response model for different scenarios or domains.
Smart Images

Figure CN113722448B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the field of artificial intelligence, and more particularly, to a query automatic answering method and device, a computer device and a medium. BACKGROUND
[0002] The current agent selling robot needs to reply automatically according to different scenarios or fields. Generally, different types of agent selling goods constitute different scenarios or fields, for example, booking train tickets, air tickets, hotels, and restaurants constitute different fields. In the process of booking or selling the same kind of goods, different stages need to be experienced, for example, first through the commodity consultation stage, then enter the bargaining stage, and finally negotiate the specific transaction method. Each of these different stages also constitutes a different scenario or field. In the prior art, different response models are constructed for each different scenario or field. Since each change in the type of goods or the change in the different sales stages, the scenario or field is different, a large number of different response models need to be trained, which requires a large amount of labeling cost and training cost. The prior art lacks a unified response model establishment method for different scenarios or fields, has high modeling cost, poor performance, and low response efficiency.
[0003] DISCLOSURE
[0004] Therefore, the present disclosure aims to provide a unified response model establishment mechanism for different scenarios or fields, reduce the modeling cost, and improve the modeling performance.
[0005] To achieve this purpose, according to one aspect of the present disclosure, a dialogue method is provided, comprising:
[0006] receiving a dialogue input of a user;
[0007] generating an action sequence according to the dialogue input, the action sequence being used to represent elements to be reflected by the generated response content;
[0008] generating response content based on the dialogue input and the action sequence.
[0009] Optionally, the elements include a field to which the dialogue input belongs, an action, and a slot.
[0010] Optionally, the generating an action sequence according to the dialogue input comprises:
[0011] generating a first element in the action sequence according to the dialogue input;
[0012] from the second element, determining a current element according to the generated element and the dialogue input.
[0013] Optionally, the generating response content based on the dialogue input and the action sequence comprises:
[0014] determining an attention score for each element in the action sequence by using a dynamic attention mechanism;
[0015] generating a response content for the dialogue input according to the dialogue input, the action sequence, and the attention score of each element in the action sequence.
[0016] Optionally, the method further comprises jointly training the task of generating the action sequence and the task of generating the response content.
[0017] Optionally, the jointly training the task of generating the action sequence and the task of generating the response content comprises:
[0018] inputting the same set of query samples to the task of generating the action sequence and the task of generating the response content;
[0019] determining a first loss rate of a result obtained by the task of generating the action sequence and a second loss rate of a result obtained by the task of generating the response content, wherein the first loss rate is a ratio of query samples in the same set of query samples in which the action sequence obtained by the task of generating the action sequence is inconsistent with an action sequence label, and the second loss rate is a ratio of query samples in the same set of query samples in which the response content obtained by the task of generating the response content is inconsistent with a response content label;
[0020] constructing a loss function by using the first loss rate, the second loss rate, a first weight of the task of generating the action sequence, and a second weight of the task of generating the response content, and stopping training if a value of the loss function is lower than a predetermined threshold.
[0021] Optionally, after constructing the loss function, the method further comprises determining values of the first weight and the second weight at which the loss function is minimized, and training the task of generating the action sequence and the task of generating the response content by using the determined values of the first weight and the second weight.
[0022] Optionally, after generating the response content, the method further comprises querying a database for a specific proper noun to replace a generic proper noun in the response content.
[0023] Optionally, the generating the response content based on the dialogue input and the action sequence comprises:
[0024] generating a first word of the response content according to the dialogue input and the action sequence;
[0025] determining a current word according to the generated word and the dialogue input, starting from a second word of the response content.
[0026] Optionally, after receiving the dialog input of the user, the method further comprises: encoding the dialog input into a dialog input code;
[0027] The generating an action sequence according to the dialog input comprises: generating an action sequence according to the dialog input code;
[0028] The generating a response content based on the dialog input and the action sequence comprises: generating a response content based on the dialog input code and the action sequence.
[0029] Optionally, the encoding the current dialog, the previous dialog, the response content to the previous dialog in the dialog input, and the attribute in the database into the dialog input code comprises: converting the current dialog, the previous dialog, the response content to the previous dialog in the dialog input, and the attribute in the database into word vectors, and concatenating the word vectors into the dialog input code.
[0030] Optionally, the encoding the current dialog, the previous dialog, the response content to the previous dialog in the dialog input, and the attribute in the database into the dialog input code comprises: converting the current dialog, the previous dialog, the response content to the previous dialog in the dialog input, and the attribute in the database into word vectors, and concatenating the word vectors into the dialog input code.
[0031] Optionally, the generating an action sequence according to the dialog input code comprises: generating an action sequence according to the dialog input code and a first masking strategy; and the generating a response content based on the dialog input and the action sequence comprises: generating a response content based on the dialog input, the action sequence, and a second masking strategy.
[0032] According to an aspect of the present disclosure, a query automatic response apparatus is provided, comprising:
[0033] An action sequence generation model is configured to generate an action sequence based on a dialog input of a user, the action sequence being used to represent an element to be reflected by a response content to be generated;
[0034] A response generation model is configured to generate a response content based on the dialog input and the action sequence.
[0035] Optionally, the element comprises: a domain to which the dialog input belongs, an action, and a slot.
[0036] Optionally, the apparatus further comprises: an attention model configured to give an attention score to each element in the action sequence; and wherein the response generation model is further configured to generate a response content to the dialog input according to the dialog input, the action sequence, and the attention score of each element in the action sequence.
[0037] According to one aspect of the present disclosure, a computer device is provided, comprising: a memory for storing computer executable code; and a processor for executing the computer executable code to implement the method as described above.
[0038] According to one aspect of the present disclosure, a computer readable medium is provided, comprising computer executable code which, when executed by a processor, implements the method as described above.
[0039] According to one aspect of the present disclosure, a responsive robot is provided, comprising the computer device as described above, or the computer readable medium as described above.
[0040] The present inventors have realized that the responsive action is hierarchical, for example, the response is first directed to the domain of a query (for example, product consultation or bargaining, the response manners in the product consultation stage and the bargaining stage are quite different), the response has different actions (whether to ask the user, or to confirm the user's requirement, or to recommend to the user, different response types determine that the response manners are quite different), and the response also uses different slots (different candidate labels, such as self-pickup, delivery location, etc.). The embodiments of the present disclosure model the elements (domain, action, slot, etc.) used in the response into an action sequence, thereby retaining the hierarchical structure of the elements in the response, and then generating the response content according to the action sequence. Since the hierarchical structure of the elements required by the response is retained in the action sequence, it is not necessary to construct different response models for each different scenario or domain as in the prior art. Through the manner of first generating the action sequence and then generating the response content according to the action sequence, the response generation model can establish a unified response model establishment method for different scenarios or domains, thereby reducing the modeling cost and improving the modeling performance. BRIEF DESCRIPTION OF DRAWINGS
[0041] The above and other objects, features and advantages of the present disclosure will become more apparent from the following description of embodiments of the present disclosure taken in conjunction with the accompanying drawings, in which:
[0042] Figure 1A FIG. F shows the interface change diagram in the scenario of the agent selling robot to which the query automatic response method according to one embodiment of the present disclosure is applied.
[0043] Figure 2 FIG. G shows a flowchart of the whole process of the agent selling robot conversation.
[0044] Figure 3 FIG. H shows a hierarchical structure diagram of the action sequence used in the response.
[0045] Figure 4 FIG. I shows a schematic diagram in which the hierarchical structure of the action sequence of the embodiments of the present disclosure replaces the single label structure of the prior art.
[0046] Figure 5 A schematic diagram showing the principle of query automatic answering according to an embodiment of the present disclosure is shown.
[0047] Figure 6 A performance comparison diagram of query automatic answering according to an embodiment of the present disclosure compared with query answering in the prior art is shown.
[0048] Figure 7 A flowchart of a query automatic answering method according to an embodiment of the present disclosure is shown.
[0049] Figure 8 A hardware diagram of a computer device performing a query automatic answering method according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0050] The present disclosure is described below based on embodiments, but the present disclosure is not limited to only these embodiments. In the following detailed description of the present disclosure, some specific details are described in detail. The present disclosure can also be fully understood without the description of these details by those skilled in the art. In order to avoid confusion of the essence of the present disclosure, the well-known methods, processes, and flows are not described in detail. In addition, the drawings are not necessarily drawn to scale.
[0051] Figure 1A -F shows an interface change diagram in a proxy selling robot scenario to which a query automatic answering method according to an embodiment of the present disclosure is applied.
[0052] A proxy selling robot is an intelligent product that can intelligently interact with a buyer when the buyer has a purchase demand, understand the user's demand, and at the same time understand various goods or services for proxy selling, so as to match the user's demand with the goods or services for proxy selling, intelligently recommend for the buyer, and complete the entire proxy selling without human participation. The difference between it and a general customer service assistant is that the customer service assistant usually only belongs to a merchant and answers the buyer for the purchase of the goods or services of the merchant, while the proxy selling robot answers the buyer for the purchase of the goods or services of all merchants on the platform and matches them. In addition, the customer service assistant cannot answer the entire purchase process, for example, it can only solve the problem of consulting goods or the problem of transaction mode, and it needs manual processing for bargaining, the proxy selling robot can intelligently process all fields including bargaining, and is more universal.
[0053] There are various implementations of the reselling robot. One is to implement in the background server of the reselling website, and the buyer logs in the webpage of the reselling website to interact with the code robot. Another is to implement as a function of the reselling application (APP). For example, the buyer installs the application on the mobile phone, and interacts with the code robot on the interface of the application to realize the purchase process. There is also a way of tangible robot provided by the mall and the like, which has a touch screen or a function of voice interaction with people, and the user interacts with the reselling robot by inputting text on the touch screen or speaking voice to realize the purchase process. Figure 1A -F mainly shows the interface change diagram of interaction with the reselling robot by the way of the buyer logging in the reselling website, and the interaction interface of other ways (such as the buyer logging in the application or the tangible robot and the like) is similar.
[0054] When the buyer logs in the reselling website, a greeting page may appear. The user enables the reselling robot by greeting, for example, inputting “Hello” on the webpage. The reselling robot may reply “Hello” on the page. Then, as shown in Figure 1A , the inquiry webpage of the reselling robot appears, which displays “What can I do for you?” and an empty query input box. The buyer inputs the query 101 in the query input box. In Figure 1A , the buyer inputs “Do you have Ferrari cars for sale?”.
[0055] The reselling robot replies according to the input query 102. As shown in Figure 1B , the reply content of the reselling robot is “What color? New or used?”, and an empty query input box is displayed. The buyer inputs further query 101 in the query input box again. In Figure 1B , the buyer inputs “Black used”.
[0056] The reselling robot continues to reply according to the input query 102. As shown in Figure 1C , the reply content of the reselling robot is “XXX production model XX picture link: XXXXX price: 400,000”, and an empty query input box is displayed. The buyer inputs further query 101 in the query input box again. In Figure 1C , the buyer starts bargaining, and inputs “Can you accept 3 million?”.
[0057] The reselling robot replies to the user's bargaining. As shown in Figure 1D , the reply content of the reselling robot is “The lowest price is 3.5 million”, and an empty query input box is displayed. The buyer inputs query 101 in the query input box again. In Figure 1D , the buyer accepts the offer of 3.5 million, and inputs “OK”.
[0058] The agent robot continues to respond 102 according to the input of the user. As shown in Figure 1E , since the bargaining has ended, the transaction mode needs to be agreed upon. The response content of the agent robot is "pick up or delivery", with a blank query input box displayed. The user inputs a query 101 in the query input box, i.e. "pick up".
[0059] The agent robot continues to respond 102 according to the input of the user. As shown in Figure 1F , since the user has chosen pick up, at this time, it is only necessary to inform the user of the pick up location, so in Figure 1F , the response content of the agent robot is "the pick up location is XXXXXXXXX, please come to the location to pick up the goods on XXXX year X month X day". The user inputs "OK" in the query input box. The entire unmanned agent selling process ends.
[0060] The above process is described by taking the text interaction between the user and the agent robot on the interface as an example, but in fact, the user and the agent robot can also perform voice interaction. In the case of voice interaction, the query of the user is converted into text through voice recognition, and then the text response content of the agent robot is also converted into corresponding voice and output from the loudspeaker. A text on one end and a voice on the other end can also be used. For example, the user inputs a query through text, and the agent robot responds through voice, or the user performs a query through voice, and the agent robot outputs text response content through the interface. In addition, the user and the agent robot can also interact through other ways (such as gestures).
[0061] As can be seen from Figure 1A -F, generally, the entire agent selling process is divided into four stages of greeting, commodity consultation, bargaining, and agreement on transaction mode, as shown in Figure 2 . First, the user 201 greets the agent robot, and the terminal 202 (in the case of interaction through the agent selling website page, the terminal 202 is the agent selling website server; in the case of interaction through the page of the application, the terminal 202 is the terminal of the user on which the application is installed) responds to the greeting. Then, the user 201 performs commodity consultation 230, and the terminal 202 responds to the commodity consultation 240, such as finalizing the color, new or old, etc. with the user 210, as in the process of Figure 1A -B. Then, the user 201 starts bargaining 250, and the terminal 202 reoffers the price or agrees to the price of the user 201 260, as in the process of Figure 1C -D. Then, the user starts to agree on the transaction mode with the agent robot 270, and the terminal 202 asks the user whether to use the pick up or delivery mode 280. The agent selling process ends.
[0062] In the above process of robot selling, the query and response have different characteristics at different stages (commodity consultation, bargaining, and transaction mode negotiation). In addition, even at the same node, such as the commodity consultation stage, different actions are used. Action is the general direction of the response. For example, sometimes the user's intention is not clear and needs to be further asked, and the inquiry is the direction of the response at this time, that is, the action is inquiry; sometimes the user expresses a clear intention and only needs to be confirmed, and the confirmation is the direction of the response at this time, that is, the action is confirmation; sometimes the user's intention is clear and needs to recommend several commodities for the user to choose, and the recommendation is the direction of the response at this time, that is, the action is recommendation. Different actions have completely different response sentence patterns. Therefore, in the prior art, different response models are constructed for each different stage or field (here, the field is equivalent to the above different stages of commodity consultation, bargaining, and transaction mode negotiation) and the combination of different actions (inquiry, confirmation, or recommendation). Each time a field (different selling stage) or an action (for example, inquiry to recommendation) is changed, a different response model is used. Building different models for each field and each action requires a large amount of labeling cost and training cost. The prior art lacks a unified response model building method for different fields and actions, has high modeling cost, poor performance, and low response efficiency.
[0063] The prior art lacks a unified response model building method for different fields and actions because the prior art does not establish a unified model for the elements (fields, actions, slots, etc.) used by the response as follows. Figure 3The hierarchical structure of the action sequence shown makes it impossible to generate responses using action sequences. Instead, it can only build models based on classification labels. The inventors of this disclosure have discovered that responses have action sequences. Action sequences include several elements that influence the way responses are made, such as domain 310, action 320, and slot 330. Domain 310 refers to the type of product being queried and the different stages in the consignment process. Different product types and different stages in the consignment process result in different response methods. For example, the response to a cosmetics inquiry is completely different from the response to a flight ticket inquiry. In the consignment process of the same product, the response method of the consignment robot should also be different when inquiring about the product or when bargaining. Therefore, domain 310 is an important element affecting the response method. Domain 310 includes product inquiry 311, bargaining 312, and product transaction method 313. In addition, action 320 is also an element affecting the response method. Action refers to the direction of the response, such as inquiry 321, confirmation 322, and recommendation 323. When the user's intent is not particularly clear, further inquiry is required. When a user expresses a clear intention, no questioning is needed, only confirmation. When a user clearly expresses their intention, neither questioning nor confirmation is needed; several products should be recommended for the user to choose from. Therefore, different actions result in drastically different response methods. Action 320 has a significant impact on the response method. Slot 330 contains candidate tags used in the response, such as color 331, new / old 332, area 333, self-pickup 334, and shipping location 335. Different candidate tags result in different response sentence structures. For example, when asking a user about their acceptable room size, the sentence structure is "What is the desired room size?" rather than "Where is the desired room size?" or "Who is the desired room size?". Therefore, slot 330 also affects the response method. This disclosure embodiment adopts... Figure 3 The queue sequence hierarchy is shown. Figure 3 In this model, each domain 310 has an arrow pointing to an action 320 that it may use, and each action 320 has an arrow pointing to a slot 330 that it may use. Therefore, the arrows between domains 310, actions 320, and slots 330 indicate their interrelationships. This embodiment of the disclosure does not directly generate the response to the query using a generative model, but rather first uses an action sequence generation model to generate a response... Figure 3 The code demonstrates the hierarchical structure of the elements (domains, actions, slots, etc.) used in the response, followed by an action sequence. Then, based on the query and this action sequence, a response generation model is used to generate the response content. Because the action sequence reflects the hierarchical structure of the elements used in the response, a unified method for building response models is provided for different domains and actions. This eliminates the need to build different models for each domain and action, reducing modeling costs, improving performance, and increasing response efficiency.
[0064] Figure 4 This illustration shows a schematic diagram of the principle of replacing the single-label structure of the prior art with a hierarchical structure of action sequences according to embodiments of the present disclosure. Embodiments of the present disclosure view the elements used in the response as a hierarchical structure of domains 310, actions 320, and slots 330, i.e., an action sequence. This hierarchical structure has a root 301, from which various domains 310 are derived. Figure 4 In this context, sector 310 includes different sectors such as hotels 314, restaurants 315, and air tickets 316. Figure 3 Area 310 in the text mainly refers to different stages in the consignment process, such as product consultation, bargaining, and agreeing on transaction methods. Figure 4 The situation is different in China. Figure 4 The "domain" refers to the type of product being queried. In other words, "domain 310" could refer to the product type itself, or it could represent different stages in the consignment process. Different product types and different stages in the consignment process will require different responses. For example, the responses to hotel reservations (314), restaurant reservations (315), and flight reservations (316) might be completely different. Furthermore, the responses will also differ during the product inquiry, negotiation, and agreement on transaction methods stages. Figure 4 In this context, action 320 includes notification 324 and request 235. When the modern sales robot understands the user's query intent, it sends notification 324 to the user, informing them of the relevant information found. When the modern sales robot needs to further clarify the user's query intent, it sends request 325 to the user to request further information. Figure 4 In the context of the response, slot 330 includes name 336, area 337, business hours 338, and telephone number 339. These are all candidate words that may be used in the response. For example, the response may need to inform the user of the business hours of the candidate restaurants for the user to choose from.
[0065] Existing technologies do not provide a unified response model for different domains and actions, but rather model according to a single label. For example, each label 302 can be a combination of a domain and a response type, and a response generation model is built for each label 302. Figure 4 For example, "hotel + notification," "hotel + request," "restaurant + notification," "restaurant + request," "flight + notification," and "flight + request" are each labeled 302, and models are built for each separately. Alternatively, labels 302 can be created according to domain + response type + slot, and a model can be built for each label. For example, "hotel + notification + phone" is a label 302, and a generative model is built for it. Therefore, in the prior art, because the hierarchical structure of the elements used in the response is not recognized and the method of first predicting the action sequence reflecting the hierarchical structure of the elements and then generating the response accordingly is adopted, a large number of models are needlessly built, resulting in high modeling costs and poor performance. This disclosure embodiment, for example... Figure 4As shown, the action sequence including start symbol 303, domain 310, action 320, and slot 330 is predicted first, and the response content is generated accordingly, which improves modeling efficiency and performance.
[0066] like Figure 7 As shown, an automatic query response method is provided. An automatic query response method refers to a method that automatically answers queries without manual intervention. In a scenario where a query response robot is installed on a website server, this method is executed by the website server. In a scenario where an application with a query response robot is installed on a user terminal, this method is executed by the user terminal. For example... Figure 7 As shown, the method includes:
[0067] Step 510: Receive user dialogue input;
[0068] Step 520: Generate an action sequence based on the dialogue input, wherein the action sequence is used to represent the elements that the response content to be generated should reflect;
[0069] Step 530: Generate response content based on the dialogue input and the action sequence.
[0070] The above steps are described in detail below.
[0071] The dialogue input in step 510 refers to the text of the query entered by the user. For example... Figure 1A -E specifies the text the user enters in the search input box. When the user uses voice input, the input speech is converted into text through speech recognition and used as the search query.
[0072] In one embodiment, after step 510, the dialogue input can be encoded into dialogue input codes, which can be used... Figure 5 The encoder 410 shown performs this operation. Dialogue input code refers to the code representing the dialogue input. Dialogue input is often in the form of text, and it generally needs to be encoded into code (such as vectors) before it can be input into the subsequent deep learning model for processing.
[0073] In one embodiment, encoding the dialogue input into dialogue input code is achieved by encoder 410 in the following manner: for each word in the current dialogue, encoder 410 queries a word vector table to obtain the word vector corresponding to each word, and concatenates the word vectors in word order to encode the dialogue input code. The current dialogue refers to the dialogue currently input by the user. For example, as... Figure 1A As shown, the current dialogue is "Do you have any Ferraris for sale?"; Figure 1BThe current dialogue is "black second-hand". The word vector table is a table that stores various words and their corresponding vectors in advance, through which the word vector corresponding to the word can be found. In the word vector table, the word vector corresponding to each word is different, so that the word can be uniquely distinguished by the word vector. There are other ways to encode the dialogue input into an input code representation, which are not described here.
[0074] In the above embodiment, the input code is obtained only from the current dialogue itself. However, in fact, the response content can not be determined only by the current dialogue. What kind of dialogue has been input before and what is the corresponding response content can also determine the response content of the current dialogue. For example, if the user inputs the current query "book tickets for the park for our whole family", if the user has input the number of people in his family before, the user does not need to be asked about the number of people in his family in the response content. If the user has not input the number of people in his family before, the user needs to be asked about the number of people in his family in the response content. Therefore, in an embodiment, the current dialogue, the previous dialogue and the response content of the previous dialogue are input into the encoder, and the encoder encodes them into a dialogue input code. The encoding method can be that, for the words of the current dialogue, the previous dialogue and the response content of the previous dialogue, the word vector table is searched one by one, and the obtained word vectors are concatenated in the order of the words of the current dialogue, the previous dialogue and the response content of the previous dialogue, to obtain the dialogue input code.
[0075] Here, the previous dialogue and the response content of the previous dialogue can be the dialogue input by the user before and the response content of all dialogues input by the user before. In this case, the information input is more comprehensive, and the obtained response content is more accurate. In another case, the previous dialogue and the response content of the previous dialogue can be the dialogue input by the user before n times and the response content of the n dialogues before, and n is a positive integer. Figure 5 The case where n = 1 is shown, where U t is the current dialogue of the user, which has a total of 4 words, generating 4 word vectors, represented by 4 ellipses; U t-1 is the dialogue input by the user before, which has a total of 4 words, generating 4 word vectors, represented by 4 ellipses; R t-1 is the response content of the dialogue input by the user before, which has a total of 4 words, generating 4 word vectors, represented by 4 ellipses. In this case, the amount of data is small, and the processing efficiency is higher.
[0076] In addition, the attributes in the database can also determine the response content. The attributes in the database are indexes of the data stored in the database. For example, for hotel data, the room type, area, floor, telephone number, etc. are stored, and when the user queries "what kind of room", the response content can contain these attributes (indexes), for example, the response content can be "what floor do you need, what type of room do you need, how big do you need the room". If the database only stores the room type and room price, the response content can be "what type of room do you need, what price do you need the room". Therefore, in an embodiment, the current dialogue, the previous dialogue, the response content to the previous dialogue, and the attributes in the database are input into the encoder, and the encoder encodes them into the dialogue input code. Specifically, each word in the current dialogue, the previous dialogue, the response content to the previous dialogue, and the attributes in the database is converted into a word vector by the encoder, and the word vectors are concatenated into the dialogue input code in the order of the current dialogue, the previous dialogue, the response content to the previous dialogue, and the attributes in the database. Figure 5 In the above embodiment, DB is the attribute in the database, and there are four attributes, generating four word vectors, represented by four ellipses. This embodiment not only considers the current dialogue and the dialogue response history when generating the response content, but also considers the attributes in the database, thereby improving the accuracy of generating the response content.
[0077] Next, in step 520, an action sequence is generated according to the dialogue input. The action sequence is used to represent the elements that the generated response content should reflect. It can be completed by the action sequence generation model 420 in Figure 5
[0078] The action sequence generation model 420 generates an action sequence based on the dialogue input code. The action sequence is a sequence composed of elements, including the domain to which the dialogue input belongs, actions, slots, etc. As shown in Figure 5 , the generated action sequence is hotel + recommend + name + area, where name and area are two slots.
[0079] In addition, the action sequence generation model 420 can generate two identical elements in the action sequence at the same time, for example, two identical slots. At this time, a deduplication operation can be performed to delete one of the duplicate elements. For example, the generated action sequence is hotel + name + recommend + name + area, and since the two "name" are the same, one of them can be deleted.
[0080] In addition, for multiple elements of the same type in the action sequence, such as multiple slots, the order of the action sequence can be arranged according to the order of the first letter of the Chinese pinyin of the first word in the alphabet, for example, "telephone" is arranged in front of "name".
[0081] Action sequence generation models can be machine learning models, especially deep neural networks. When training a deep neural network, a set of dialogue input samples consisting of a large number of dialogue input samples can be constructed. For each dialogue input sample, an expert pre-labels the action sequence associated with the response to that input sample. Then, the dialogue input code obtained by the encoder for each input sample is input into the deep neural network, which generates the determined action sequence. The determined action sequence for each input sample is compared with the action label to determine the ratio of the determined action sequences matching the action label across all dialogue input samples. Once this ratio falls below a predetermined ratio (e.g., 95%), the node coefficients in the deep neural network are adjusted until the ratio exceeds the predetermined ratio; if the predetermined ratio is exceeded, the action sequence generation model is considered successfully trained. Inputting the user's dialogue input code into the model allows it to generate action sequences.
[0082] In one embodiment, when generating an action sequence, subsequent elements are generated based on already generated elements. That is, the first element in the action sequence is generated based on the dialogue input. Starting with the second element, the current element is determined based on the already generated elements and the dialogue input. For example, when determining the second element, if the current element is the second element, the second element is determined based on the first element and the dialogue input; when determining the third element, if the current element is the third element, the third element is determined based on the first and second elements and the dialogue input. When the dialogue input is encoded, the elements in the action sequence are generated not only based on the dialogue input code but also based on the elements preceding that element in the action sequence. Figure 5 As shown, the action sequence generation model 420 generates a domain 310, namely "hotel," based on the dialogue input code and start symbol 303 output by encoder 410. Then, based on the code output by encoder 410 and the domain "hotel," it generates a dialogue 320, namely "recommendation." Next, based on the code output by encoder 410 and the two already generated elements "hotel" and "recommendation," it generates the first slot 320, namely "name." Then, based on the code output by encoder 410 and the three already generated elements "hotel," "recommendation," and "name," it generates the second slot 320, namely "area." This embodiment fully considers the dependencies between elements, taking into account previously generated elements when generating subsequent elements, thus improving the accuracy of the generated action sequence.
[0083] Next, in step 530, response content is generated based on the dialogue input and the action sequence.
[0084] Based on the dialogue input and the action sequence, a response can be generated. Figure 5The response generation model 430 can utilize a machine learning model, particularly a deep neural network. In training the deep neural network, a set of dialogue input samples can be constructed from a large number of dialogue input samples. For each dialogue input sample, a response label is attached to the dialogue input sample in advance by an expert. The dialogue input code of each dialogue input sample obtained by the encoder is input into the trained action sequence generation model to obtain an action sequence. The action sequence and the dialogue input code of the dialogue input sample are input into the deep neural network of the response generation model, and the determined response content is obtained by the deep neural network. The determined response content is compared with the response content label to determine the ratio of the determined response content consistent with the response content label in all query samples. Once the ratio is lower than a predetermined ratio (e.g., 95%), the node coefficients in the deep neural network are adjusted until the ratio exceeds the predetermined ratio, and the response generation model is considered to be successfully trained. The response generation model can generate the response content when input with the real dialogue input and the action sequence generated by the action sequence generation model for the dialogue input.
[0085] In one embodiment, when generating the response content, the following word is generated according to the word that has been generated in the response content. That is, the first word in the response content is generated according to the dialogue input and the action sequence. From the second word in the response content, the current word is determined according to the generated word and the dialogue input. The word in the response content is generated by the response generation model based on the dialogue input code and the word before the word in the response content. As shown in the following table, the first word in the response content is generated according to the code output by the encoder 410, the action sequence, and the start symbol. The first word is a generic proper noun, i.e., “<hotel name>”, which is to be replaced with a specific proper noun in the following process. The generic proper noun refers to a generic name for a person or thing, such as a person's name, a hotel name, a restaurant name, etc. The specific proper noun refers to a specific name for a person or thing, such as Li Si, Friendship Hotel, and Fu Gui Restaurant. Then, the second word “how” is generated according to the code output by the encoder 410, the action sequence, and the first word “<hotel name>” that has been generated. Then, the third word “area” is generated according to the code output by the encoder 410, the action sequence, and the generated words “<hotel name>” and “how”. Then, the fourth word “is” is generated according to the code output by the encoder 410, the action sequence, and the generated words “<hotel name>”, “how”, and “area”. This embodiment fully considers the contextual dependency relationship between the words in the response content, takes into account the previously generated words when generating the following words in the response content, and improves the accuracy of generating the response content. Figure 5 As shown in the table, the response generation model 430 generates the first word in the response content according to the code output by the encoder 410, the action sequence, and the start symbol. The first word is a generic proper noun, i.e., “<hotel name>”, which is to be replaced with a specific proper noun in the following process. A generic proper noun refers to a generic name for a person or thing, such as a person's name, a hotel name, a restaurant name, etc. A specific proper noun refers to a specific name for a person or thing, such as Li Si, Friendship Hotel, and Fu Gui Restaurant. Then, the second word “how” is generated according to the code output by the encoder 410, the action sequence, and the first word “<hotel name>” that has been generated. Then, the third word “area” is generated according to the code output by the encoder 410, the action sequence, and the generated words “<hotel name>” and “how”. Then, the fourth word “is” is generated according to the code output by the encoder 410, the action sequence, and the generated words “<hotel name>”, “how”, and “area”. This embodiment fully considers the contextual dependency relationship between the words in the response content, takes into account the previously generated words when generating the following words in the response content, and improves the accuracy of generating the response content.
[0086] In the above embodiments, each element in the action sequence is input to the response generation model 430 equally, but in fact, their effects on the response are not equally influential. Some elements should be paid special attention when responding, and some elements should not. Therefore, in an embodiment, step 530 comprises: giving an attention score to each element in the action sequence by using a dynamic attention mechanism (e.g. the attention model 440 in Figure 3 Based on the dialogue input, the action sequence, and the attention score of each element, generating the response content for the dialogue input.
[0087] Since the attention model 440 can obtain the attention score of each element (domain, action, slot) in the action sequence, which reflects the degree to which each element needs to be paid attention to when generating the response content, the element to which attention should be paid can be paid attention to in time when generating the response content, filtering out the interference brought by irrelevant elements, and improving the response quality. The attention model 440 pays attention to different actions in the action sequence at different times according to the current state, which can enhance the response generation model 430, make up for the deficiency of the model that cannot remember all the information when the input is too long, and improve the response accuracy.
[0088] The attention score is an inherent concept of the attention model 440, and is not described in detail. In the attention model 440 in Figure 5 The attention score of “name” is the highest, followed by “hotel”, “recommend”, “is”, and the start symbol in turn. In the case of the attention model 440, as shown in Figure 5 As shown in Figure 5 The response generation model 430 generates the first word in the response content, i.e. “<hotel name>”, according to the code output by the encoder 410, the action sequence output by the attention model 440, and the attention score of each element in the action sequence, and the start symbol. Then, the second word “how” is generated according to the encoding output by the encoder 410, the action sequence output by the attention model 440, and the attention score of each element in the action sequence, and the first word “<hotel name>” that has been generated. Then, the third word “area” is generated according to the code output by the encoder 410, the action sequence output by the attention model 440, and the attention score of each element in the action sequence, and the words “<hotel name>”, “how” that have been generated. Then, the fourth word “is” is generated according to the code output by the encoder 410, the action sequence output by the attention model 440, and the attention score of each element in the action sequence, and the words “<hotel name>”, “how”, “area” that have been generated.
[0089] The attention model can be a machine learning model, especially a deep neural network. In training the deep neural network, a set of action sequence samples can be constructed, each of which is composed of a large number of action sequence samples. For each action sequence sample, the attention scores of the elements therein are determined in advance by an expert, and are labeled with attention score labels. The action sequence is input into the attention model, and the attention scores of the elements in the action sequence are output by the attention model, which are compared with the attention score labels. The ratio of the determined attention scores consistent with the attention score labels in all action sequence samples is determined. Once the ratio is lower than a predetermined ratio (e.g., 95%), the node coefficients in the deep neural network are adjusted until the ratio exceeds the predetermined ratio, and the attention model is considered to be successfully trained. When a real action sequence is input into the model, the attention scores of the elements therein can be generated by the model.
[0090] Before the response generation model generates the response content to the query based on the dialogue input code, the action sequence, and the attention scores of the elements in the action sequence, the response generation model can be pre-trained as follows: for each dialogue input sample in the set of dialogue input samples, the response label for the dialogue input sample is labeled in advance by an expert. The dialogue input code obtained by the encoder for each dialogue input sample is input into the trained action sequence generation model to obtain an action sequence. The action sequence is input into the trained attention model to obtain the attention scores of the elements in the action sequence. The attention scores of the elements are input into the deep neural network of the response generation model with the dialogue input code, and the determined response content is obtained by the deep neural network. The response content is compared with the response content label, and the ratio of the determined response content consistent with the response content label in all dialogue input samples is determined. Once the ratio is lower than a predetermined ratio (e.g., 95%), the node coefficients in the deep neural network are adjusted until the ratio exceeds the predetermined ratio, and the response generation model is considered to be successfully trained. When the dialogue input code and the attention scores of the elements in the action sequence are input into the model, the response content can be generated by the model.
[0091] In the above process, the action sequence generation model and the response generation model are trained separately. In an embodiment, the action sequence generation model and the response generation model can be jointly trained, so that information (e.g., the dialogue input code of the input query output by the encoder) can be shared, improving the performance and efficiency of training.
[0092] When jointly trained, the action sequence generation model and the response generation model can share the same set of dialogue input samples. For each dialogue input sample in the set of dialogue input samples, the dialogue input sequence label and the response label are labeled in advance by an expert.
[0093] For the set of dialogue input samples, the dialogue input code of each dialogue input sample in the set of dialogue input samples is encoded by the encoder, and an action sequence generation model is used to generate an action sequence, which is compared with the action sequence label of the dialogue input sample, so as to determine the first loss rate of the set of dialogue input samples generated by the action sequence generation model. The first loss rate is equal to the ratio of the number of generated action sequences that are inconsistent with the action sequence label of the dialogue input sample to the number of dialogue input samples in the set.
[0094] Then, the dialogue input code of the dialogue input sample in the set of dialogue input samples and the action sequence generated by the action sequence generation model are input into the response generation model to generate the response content of the dialogue input sample, which is compared with the response content label of the dialogue input sample, so as to determine the second loss rate of the set of dialogue input samples generated by the response generation model. The second loss rate is equal to the ratio of the number of generated response contents that are inconsistent with the response content label of the dialogue input sample to the number of dialogue input samples in the set.
[0095] Then, the first loss rate, the second loss rate, the first weight of the response element sequence generation model, and the second weight of the response generation model are used to construct a loss function. For example, an uncertainty loss function can be constructed. This function uses homoscedastic uncertainty to measure task-related uncertainty. The method of constructing the loss function is known and will not be described here. After the loss function is constructed, the values of the first weight and the second weight that minimize the loss function are determined. Then, the action sequence generation model and the response generation model are trained using the determined values of the first weight and the second weight. That is, the weighted sum of the first loss rate and the second loss rate is calculated using the determined values of the first weight and the second weight. When measuring whether the action sequence generation model and the response generation model have been well trained, it is determined whether the weighted sum is less than a predetermined loss rate weighted sum threshold (e.g., 5%). If not, it means that the two models have not been well trained and the node coefficients in the two models need to be adjusted so that the finally calculated weighted sum is less than the predetermined threshold (e.g., 5%), and the training is stopped after the predetermined threshold is less than the predetermined threshold.
[0096] This way of joint training fully considers the dependency between generating an action sequence and generating a response content, improves the quality of the trained model, and can share information (e.g., the dialogue input code output by the encoder), improving the efficiency of training.
[0097] As mentioned above, the action sequence generation model and the response generation model can share the dialogue input code output by the encoder, but the two models can have different focuses and thus set different masking strategies. When the action sequence generation model generates an action sequence, the action sequence can be generated based on the dialogue input code and a first masking strategy. When the response generation model generates response content, the response content can be generated based on the dialogue input code, the action sequence, and a second masking strategy. For example, the action sequence generation model can not be interested in attributes in the database, and thus the first masking strategy can be to mask the part of the code that encodes the attributes in the database. The response generation model can not be interested in user queries and response history, and thus the second masking strategy can be to mask the part of the code that encodes the previous dialogue and the response content of the previous dialogue. Through different masking strategies, information is shared, and information that is useless to each model is prevented from interfering with the model, improving model performance.
[0098] In the above embodiment, the action sequence and the response content are sequentially generated, i.e., one generated action sequence is generated by the action sequence generation model, and one response content is generated by the response generation model based on the complete action sequence. In another embodiment, the generation of the action sequence and the generation of the response content are cross-performed, i.e., one element in the action sequence is generated, and then one or more words in the response content are generated based on the element, and then one element in the action sequence is generated, and then one or more words in the response content are generated based on the element, and so on, until the action sequence and the response content are completely generated.
[0099] In this embodiment, the action sequence is generated element by element, and the response content is generated word by word. When generating the response content to the dialogue input, the action sequence generation model generates one element each time, and inputs the dialogue input code and the action sequence generated by the action sequence generation model to the response generation model, and generates one or more corresponding words in the response content based on the response generation model. Figure 5 As shown, the action sequence generation model 420 generates the element "hotel", and then the response generation model 430 generates "<hotel name>" based on "hotel" and the code obtained by encoding the dialogue input by the encoder 410. Then, the action sequence generation model 420 generates the element "recommend", and then the response generation model 430 generates "how about" based on the two elements "hotel" and "recommend" and the code obtained by encoding the dialogue input by the encoder 410. Next, the action sequence generation model 420 generates the element "area", and then the response generation model 430 generates the two words "area" and "is" based on the three elements "hotel", "recommend", and "area" and the code obtained by encoding the dialogue input by the encoder 410, and so on. This way of cross-generating the action sequence and the response content improves the efficiency of generating the response content.
[0100] As shown in Figure 5 , the answer content output by the answer generation model 430 can contain a generic proper noun such as "<hotel name>" and in the actual answer, it needs to be replaced with a specific proper noun such as "Happy Hotel". In the database 450, for a generic proper noun, all candidate specific proper nouns thereunder are stored. For example, for "<hotel name>", many specific hotel names such as "Happy Hotel", "Friendship Hotel", etc. are stored. Which specific proper noun to recommend to the user needs to be determined according to the current dialogue, previous dialogue and keywords extracted from the answer content of the previous dialogue, and the keywords are used for retrieval. For example, the user mentioned in the previous dialogue that "close to the Youth Palace" and "room type area within 30 square meters" and other requirements, then extract keywords such as "Youth Palace" and "within 30 square meters" from them, and use these keywords to search in the database 450, and replace the generic proper noun in the answer content output by the answer generation model 430 with the hit specific proper noun. As shown in Figure 6 , the final answer content displayed to the user is "How about Happy Hotel? The room type area is 25 square meters".
[0101] The inventors of the present disclosure realize that the elements are hierarchical, for example, the answer is first directed to the field of a query, the answer has different actions, and the answer also uses different slots (different candidate labels). The embodiments of the present disclosure model the elements used in these answers as action sequences, thereby preserving the hierarchical structure of the elements in the answers and considering the relationship between these elements, and first generating the action sequence, and generating the answer content according to the action sequence. Since the hierarchical structure of the elements is preserved in this action sequence, it is not necessary to construct different answer models for each different scenario or field as in the prior art. Therefore, a unified answer model construction method can be established for different scenarios or fields, reducing the modeling cost and improving the modeling performance.
[0102] In addition, in the embodiments of the present disclosure, the action sequence and the answer content are generated simultaneously, and the attention mechanism is used, so that the importance of each element can be dynamically focused on when generating the answer content.
[0103] In addition, in the embodiments of the present disclosure, an uncertain loss function is used to model the loss function of the joint training of the action sequence generation model and the answer model, so that the model training process is more stable and the training result is better.
[0104] Figure 6A performance comparison chart of query automatic answering according to the embodiment of the present disclosure compared with query answering in the prior art is shown. The performance is expressed by a performance score, which includes at least two parts of scores, one part is a completion condition score, and the other part is a completion fluency score. When testing, a certain number (for example, 1000) of users are adopted, which respectively make the same query to the vending robot of the prior art and the vending robot of the embodiment of the present disclosure, and the number of users for which the vending robot successfully completes the vending is counted. The completion ratio is obtained by dividing the number of successfully vending users by the total number of users, and the completion condition score is obtained based on the completion ratio, for example, the completion condition score can be proportional to the completion ratio. When determining the completion fluency score, the form of investigating the successfully vending users can be adopted, and if after the user investigation, the number of users who think that the overall answer is fluent is divided by the total number of successfully vending users, the fluency ratio is obtained, and the completion fluency score is obtained based on the fluency ratio, for example, the completion fluency score can be proportional to the fluency ratio. The completion condition score and the completion fluency score are added to obtain the performance score. Figure 6 The performance scores of the vending robots of the prior art and the performance scores of the vending robots of the embodiment of the present disclosure are determined respectively for users who make reservations for hotels, trains, restaurants, scenic spots, and taxis. When targeting a certain field, the users participating in the test can all be users who make reservations in the field. For example, when testing the performance score of hotel reservation, 1000 users participating in the test can all be users who make hotel reservations. When testing the performance score of train ticket reservation, 1000 users participating in the test can all be users who make train ticket reservations, and so on. From Figure 6 It can be seen from Table 1 that in any field, the performance of the vending robot of the embodiment of the present disclosure is much higher than that of the vending robot of the prior art.
[0105] In addition, Figure 6 The average performance scores of the vending robots of the prior art and the average performance scores of the vending robots of the embodiment of the present disclosure in single fields of hotels, trains, restaurants, scenic spots, and taxis are also determined in Table 2. The performance of the vending robot of the embodiment of the present disclosure is better than that of the vending robot of the prior art. In addition, Figure 6 The performance scores of the vending robots of the prior art and the performance scores of the vending robots of the embodiment of the present disclosure when the tested population is multi-field reservation users are also shown in Table 3. When testing, 1000 users participating in the test can include users who make reservations in various fields, and finally the performance score of the vending robot of the prior art and the performance score of the vending robot of the embodiment of the present disclosure are obtained for the 1000 users. As Figure 5As shown, the embodiments of the present disclosure far exceed the prior art on multi-domain dialogue data sets, with overall scores increasing from 99.50 to 105.47. The performance of the embodiments of the present disclosure on all domains far exceeds the prior art, and the superiority is more obvious in the multi-domain scenario.
[0106] According to one embodiment of the present disclosure, as Figure 8 As shown, a query automatic answering apparatus 400 is provided, comprising:
[0107] an action sequence generation model 420, configured to generate an action sequence based on a dialogue input of a user, the action sequence being used to represent elements to be reflected by generated answer content;
[0108] an answer generation model 430, configured to generate answer content based on the dialogue input and the action sequence.
[0109] Optionally, the elements include a domain to which the dialogue input belongs, an action, and a slot.
[0110] Optionally, the action sequence generation model 420 is further configured to:
[0111] generate a first element in the action sequence according to the dialogue input;
[0112] starting from a second element, determine a current element according to the dialogue input and the generated element.
[0113] Optionally, the apparatus 400 further comprises an attention model 440, configured to give an attention score to each element in the action sequence; and the answer generation model 430 is further configured to generate answer content for the query based on the input query, the action sequence, and the attention score of each element in the action sequence.
[0114] Optionally, the action sequence generation model 420 and the answer generation model 430 are jointly trained in advance.
[0115] Optionally, the joint training comprises:
[0116] inputting a same query sample set into the action sequence generation model 420 and the answer generation model 430;
[0117] determine a first loss rate of results obtained by the action sequence generation model 420 and a second loss rate of results obtained by the response generation model 430, wherein the first loss rate is a ratio of query samples in the same set of query samples in which the action sequence generated by the action sequence generation model 420 is inconsistent with the action sequence label, and the second loss rate is a ratio of query samples in the same set of query samples in which the response content generated by the response generation model 430 is inconsistent with the response content label;
[0118] construct a loss function using the first loss rate, the second loss rate, a first weight of the action sequence generation model 420, and a second weight of the response generation model 430, and stop training if the value of the loss function is lower than a predetermined threshold.
[0119] Optionally, the joint training further comprises: determining the values of the first weight and the second weight when the loss function is at a minimum, and training the action sequence generation model 420 and the response generation model 430 using the determined values of the first weight and the second weight.
[0120] Optionally, the apparatus 400 further comprises a database 450, wherein after generating the response content, the database 450 is queried to replace the generic proper noun in the response content with the specific proper noun.
[0121] Optionally, the response generation model 430 is further configured to:
[0122] generate a first word in the response content according to the dialogue input and the action sequence;
[0123] determine a current word according to the generated word and the dialogue input, starting from a second word in the response content.
[0124] Optionally, the apparatus 400 further comprises an encoder 410 configured to encode the dialogue input into a dialogue input code after receiving the dialogue input of the user. The action sequence generation model 420 generates an action sequence according to the dialogue input code. The response generation model 430 generates response content based on the dialogue input code and the action sequence.
[0125] Optionally, the encoder 410 is further configured to encode the current dialogue, the previous dialogue, the response content to the previous dialogue in the dialogue input, and the attributes in the database into the dialogue input code.
[0126] Optionally, the encoder 410 is further configured to convert the current dialogue, the previous dialogue, the response content to the previous dialogue in the dialogue input, and the attributes in the database into word vectors, and concatenate the word vectors into the dialogue input code.
[0127] Optionally, the action sequence generation model 420 generates an action sequence according to the dialogue input code and a first mask strategy; and the response generation model 430 generates response content based on the dialogue input, the action sequence, and a second mask strategy.
[0128] The implementation details of the query automatic response apparatus 400 have been described in detail in the method embodiments above, and thus are not described herein.
[0129] The query automatic response according to one embodiment of the present disclosure can be implemented by a computer device 800. Figure 8 In the scenario where a response robot is installed on a website server, the computer device 800 is the website server. In the scenario where an application with a response robot is installed on a user terminal, the computer device 800 is the user terminal.
[0130] The computer device 800 according to an embodiment of the present disclosure is described below with reference to Figure 8 The computer device 800 shown is merely one example and should not be taken as limiting the scope of the present embodiments. Figure 8 The computer device 800 shown is merely one example and should not be taken as limiting the scope of the present embodiments.
[0131] As shown in Figure 7 The computer device 800 is in the form of a general-purpose computing device. The components of the computer device 800 can include, but are not limited to, the at least one processing unit 810, the at least one storage unit 820, and a bus 830 connecting different system components, including the storage unit 820 and the processing unit 810.
[0132] The storage unit stores program codes that can be executed by the processing unit 810, so that the processing unit 810 performs the steps of various exemplary embodiments of the present disclosure described in the description part of the exemplary method above. For example, the processing unit 810 can perform various steps as shown in
[0133] The storage unit 820 can include a readable medium in the form of a volatile storage unit, such as a random access memory (RAM) 8201 and / or a cache memory 8202, and can further include a read-only memory (ROM) 8203.
[0134] The storage unit 820 can further include program / utility 8204 having a set of program modules 8205, including but not limited to, an operating system, one or more application programs, other program modules, and program data, each of which can include implementation of a network environment, alone or in combination.
[0135] Bus 830 can be one of several types of bus structures including a storage bus or a memory bus, a peripheral bus, a graphics bus, a processor bus, or a local bus using any of a variety of bus architectures.
[0136] Computer device 800 can also communicate with one or more external devices 700 such as a keyboard or pointing devices, a display, a printer, etc. via user input and / or output (I / O) interface(s) 850. Note that that the computer device 800 can also communicate with one or more devices 700 selected via a communication interface 860. In this regard, communication interface 860 can enable computer device 800 to communicate with other computer devices and / or systems, for example, via an electronic communications network, over the Internet, etc. using any one of a number of commercially available protocols (e.g., TCP / IP, Ethernet, Bluetooth®, etc.). Note that these technologies are known in the art and need not be discussed at length here.
[0137] It is to be understood that the above-referenced elements of the present disclosure are merely preferred embodiments of the present disclosure and that numerous variations of the embodiments described can be made and still fall within the scope of the present disclosure. Any modifications, equivalents, improvements, combinations, or the like not described above are intended to be within the scope of the present disclosure.
[0138] It should be understood that each of the embodiments described in the specification ensure progressive presentation of concepts and that identical or similar parts found in the different embodiments can be mutually substituted for each other in order to describe the different embodiments.
[0139] It should be understood that the detailed description of the specific embodiments of the present disclosure is not intended to limit the present disclosure. Other embodiments within the scope of the claims will be apparent to those skilled in the art from consideration of the specification and may be practiced without departing from the spirit or essential character of the claims. It is furthermore understood that the steps recited in the claims need not be performed in the order recited in the claims. Furthermore, the processes depicted in the accompanying figures need not be performed in the order depicted in the figures. Indeed, all of the depicted processes can be performed in parallel, or in any order. In some embodiments, multiple tasks can be performed at the same time, or in any order.
[0140] It should be understood that the use of a singular form to describe an element or to show only one element in the accompanying drawings does not imply that the number of such element is limited to one. Furthermore, modules or elements described or shown as separate herein may be combined into a single module or element, and modules or elements described or shown as single herein may be broken down into multiple modules or elements.
[0141] It should also be understood that the terminology and expressions used herein are for descriptive purposes only, and one or more embodiments described herein should not be limited to these terms and expressions. The use of these terms and expressions does not exclude any illustrative and descriptive equivalent features (or parts thereof), and it should be recognized that various modifications that may exist should also be included within the scope of the claims. Other modifications, variations, and substitutions may also exist. Accordingly, the claims should be considered to cover all such equivalents.
Claims
1. A dialogue method, comprising: receiving a dialogue input of a user; generating an action sequence according to the dialogue input, the action sequence being used to represent an element to be reflected by a generated response content; generating a response content based on the dialogue input and the action sequence; the method further comprises: constructing a loss function by using a first loss rate, a second loss rate, a first weight of a task of generating the action sequence, and a second weight of a task of generating the response content, wherein the loss function is used to jointly train the task of generating the action sequence and the task of generating the response content, the first loss rate is based on a same set of query samples, and is obtained by performing the task of generating the action sequence, and the second loss rate is based on the same set of query samples, and is obtained by performing the task of generating the response content.
2. The method of claim 1, wherein, the element comprises a domain to which the dialogue input belongs, an action, and a slot.
3. The method of claim 2, wherein, the generating the action sequence according to the dialogue input comprises: generating a first element in the action sequence according to the dialogue input; starting from a second element, determining a current element according to the generated element and the dialogue input.
4. The method of claim 2, wherein, the generating the response content based on the dialogue input and the action sequence comprises: using a dynamic attention mechanism to give an attention score to each element in the action sequence; generating a response content for the dialogue input according to the dialogue input, the action sequence, and the attention score of each element in the action sequence.
5. The method of claim 1, wherein, the jointly training the task of generating the action sequence and the task of generating the response content comprises: inputting a same set of query samples to the task of generating the action sequence and the task of generating the response content; determining the first loss rate of a result obtained by the task of generating the action sequence and the second loss rate of a result obtained by the task of generating the response content, wherein the first loss rate is a ratio of query samples in which an action sequence obtained by the task of generating the action sequence is inconsistent with an action sequence label in the same set of query samples, and the second loss rate is a ratio of query samples in which a response content obtained by the task of generating the response content is inconsistent with a response content label in the same set of query samples; stopping training if a value of the loss function is lower than a predetermined threshold.
6. The method of claim 5, wherein, after constructing the loss function, the method further comprises: determining values of the first weight and the second weight when the loss function is at a minimum, and training the task of generating the action sequence and the task of generating the response content by using the determined values of the first weight and the second weight.
7. The method of claim 1, wherein, after generating the response content, the method further comprises: querying a database for a specific proper noun to replace a generic proper noun in the response content.
8. The method of claim 1, wherein, the generating the response content based on the dialogue input and the action sequence comprises: generating a first word of the response content according to the dialogue input and the action sequence; starting from a second word of the response content, determining a current word according to the generated word and the dialogue input.
9. The method of claim 1, wherein, after receiving the dialogue input of the user, the method further comprises: encoding the dialogue input into a dialogue input code; The generating the action sequence according to the dialogue input comprises: generating the action sequence according to the dialogue input code. The generating the response content based on the dialogue input and the action sequence comprises: generating the response content based on the dialogue input code and the action sequence.
10. The method of claim 9, wherein, The encoding the dialogue input into the dialogue input code comprises: The encoding the dialogue input into the dialogue input code comprises:
11. The method of claim 10, wherein, The encoding the dialogue input into the dialogue input code comprises: converting the current dialogue, the previous dialogue, the response content to the previous dialogue, and the attributes in the database in the dialogue input into word vectors, and concatenating the word vectors into the dialogue input code.
12. The method of claim 9, wherein, The generating the action sequence according to the dialogue input code comprises: generating the action sequence according to the dialogue input code and a first mask strategy. The generating the response content based on the dialogue input and the action sequence comprises: generating the response content based on the dialogue input, the action sequence, and a second mask strategy.
13. An automatic response device for query, comprising: an action sequence generation model configured to generate an action sequence based on dialogue input of a user, the action sequence being used to represent an element to be reflected by response content to be generated; a response generation model configured to generate response content based on the dialogue input and the action sequence; The device is further configured to construct a loss function by using a first loss rate, a second loss rate, a first weight of a task of generating the action sequence, and a second weight of a task of generating the response content, wherein the loss function is used to jointly train the task of generating the action sequence and the task of generating the response content, the first loss rate is obtained based on a same query sample set by performing the task of generating the action sequence, and the second loss rate is obtained based on the same query sample set by performing the task of generating the response content.
14. The apparatus of claim 13, wherein, The element comprises: a domain to which the dialogue input belongs, an action, and a slot.
15. The apparatus of claim 14, further comprising: An attention model gives an attention score to each element in the action sequence. The response generation model is further configured to generate the response content to the dialogue input based on the dialogue input, the action sequence, and the attention score of each element in the action sequence.
16. A computer device, comprising: a memory configured to store computer executable code; a processor configured to execute the computer executable code to implement the method of any one of claims 1-12.
17. A computer readable medium characterized by Computer executable code, when executed by a processor, implements the method of any one of claims 1-12.
18. A responsive robot comprising: The computer device of claim 16, or the computer readable medium of claim 17.
Citation Information
Patent Citations
Topic-based intelligent dialogue method and system
CN106095834A
Generating dialogue responses in end-to-end dialogue systems utilizing a context-dependent additive recurrent neural network
US20200090651A1