A dialogue method, dialogue device and intelligent device
By integrating the news recognition model and user portrait model in the smart device, the smart device can actively extract key news information and user interests, generate target opening statements, solving the problem of passiveness of the existing dialogue system, and realizing meaningful and active dialogue between the smart device and the user.
Patent Information
- Application Number
- CN202210283666.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-22
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2042-03-22
AI Technical Summary
Most existing dialogue systems are passive, and users need to actively wake up the device to conduct dialogue, which leads to a decrease in user enthusiasm for dialogue and makes it difficult for the device to fully demonstrate its use value.
The trained news recognition model extracts key information of the day's news, generates candidate opening statements, and determines the target opening statements through the user portrait model to realize active dialogue on the smart device.
Enable smart devices to actively have meaningful conversations with users. The topics are based on the news and user interests of the day, improve users' enthusiasm for dialogue and make full use of the functions of the device.
Smart Images

Figure CN114756646B_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of artificial intelligence technology, and in particular relates to a dialogue method, a dialogue device, an intelligent device, and a computer-readable storage medium. Background Art
[0002] With the maturity and widespread adoption of artificial intelligence (AI) technology, conversational systems are gradually being applied across various industries, revolutionizing numerous fields, including automotive, smart home, smart wearables, and smart speakers. However, most existing conversational systems are passive, requiring the user to activate and actively ask questions before the device can respond. Because users are unaware of the range of conversations a device can handle, conversations often remain limited to common topics like weather and traffic inquiries. Over time, this can dampen user engagement, making it difficult for the device's conversational system to truly deliver value. Summary of the Invention
[0003] The present application provides a conversation method, a conversation apparatus, an intelligent device, and a computer-readable storage medium, which enable the intelligent device to proactively engage in meaningful conversations with users.
[0004] In a first aspect, the present application provides a dialogue method, comprising:
[0005] Through the trained news recognition model, the key information of all news that has occurred on that day is extracted;
[0006] Generate candidate opening sentences based on the above key information;
[0007] Extract the user's user profile through the trained user profile model;
[0008] Determine a target opening statement from the candidate opening statements based on the user profile.
[0009] Based on the above target opening statement, initiate a dialogue with the above user.
[0010] In a second aspect, the present application provides a conversation device, comprising:
[0011] The first extraction module is used to extract key information of all news that has occurred on the day through the trained news recognition model;
[0012] A generation module, configured to generate candidate opening statements based on the above key information;
[0013] The second extraction module is used to extract the user profile of the user through the trained user profile model;
[0014] a determination module, configured to determine a target opening statement from the candidate opening statements according to the user profile;
[0015] The dialogue initiation module is used to initiate a dialogue with the user based on the target opening statement.
[0016] In a third aspect, the present application provides an intelligent device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the method according to the first aspect are implemented.
[0017] In a fourth aspect, the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method of the first aspect are implemented.
[0018] In a fifth aspect, the present application provides a computer program product, which includes a computer program. When the computer program is executed by one or more processors, it implements the steps of the method of the first aspect.
[0019] Compared to the prior art, the present application has the following advantages: the smart device first uses a trained news recognition model to extract key information from all news that has occurred that day. It then generates candidate opening statements based on this key information. It then uses a trained user profile model to extract the user's profile. From the candidate opening statements, it determines a target opening statement based on the user profile. Finally, it initiates a conversation with the user based on the target opening statement. This process enables the smart device to proactively engage in a conversation with the user, and the topics of proactive conversations have the following characteristics: First, they are generated based on the day's news, allowing the user to obtain the latest information; second, they are filtered based on the user profile, allowing the output topics to capture the user's interests. In this way, the smart device can proactively engage in meaningful conversations with the user, fully stimulating the user's enthusiasm for engaging with the smart device. It is understood that the beneficial effects of the second to fifth aspects described above can be found in the relevant description of the first aspect and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments or descriptions of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0021] Figure 1This is a schematic diagram of the implementation flow of the dialogue method provided in the embodiment of the present application;
[0022] Figure 2 This is an example output diagram of the news recognition model in the dialogue method provided in the embodiment of the present application;
[0023] Figure 3 This is an example output diagram of the user portrait model in the dialogue method provided in the embodiment of the present application;
[0024] Figure 4 This is a structural block diagram of a conversation device provided in an embodiment of the present application;
[0025] Figure 5 It is a schematic diagram of the structure of the smart device provided in the embodiment of the present application. DETAILED DESCRIPTION
[0026] In the following description, specific details such as specific system structures and techniques are provided for purposes of illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present application with unnecessary detail.
[0027] In order to illustrate the technical solution proposed in this application, a specific embodiment is provided below.
[0028] The following uses a robot as an example to illustrate the conversation method proposed in the embodiment of the present application. Of course, the smart device can also be a smart phone, tablet computer, smart home or smart speaker, etc. The type of smart device is not limited here. Figure 1 , the implementation process of the dialogue method is detailed as follows:
[0029] Step 101: extract key information of all news that has occurred on that day through the trained news recognition model.
[0030] In an embodiment of the present application, a robot can perform offline processing of all news items occurring on a given day at a specified time point or time period, obtaining key information for each piece of news that occurred up to that specified time point or time period. This offline processing primarily relies on a trained news recognition model. As an example, the key news information extracted by the news recognition model includes the following: entity, event, agent, and recipient. The agent and recipient are actually entities, and can be considered special entities related to the event.
[0031] Among them, entities can express specific things such as time, place or people, for example: the Chinese team.
[0032] Among them, events can express things that are triggered by an entity or things that happen to an entity, for example: the Chinese team won the championship.
[0033] In some embodiments, for any piece of news, the robot can input the title, rather than the full text, into the trained news recognition model to extract the news's key information. Because the number of words in a news title is far less than the number of words in the full text, and the title summarizes the core information of the news, this can significantly reduce the processing pressure on the news recognition model and effectively increase the speed of extracting key information from the news.
[0034] Step 102: Generate candidate opening statements based on the key information.
[0035] In this embodiment of the present application, the robot can generate candidate opening statements based on the key information of each news item of the day extracted in step 101. Generally speaking, at least one candidate opening statement can be generated for each news item. It is understood that the candidate opening statements proposed in this embodiment of the present application are actually topics related to the news that the robot proactively proposes. They can be declarative statements or interrogative statements, and the form of the candidate opening statements is not limited here.
[0036] Step 103: extract the user profile of the user through the trained user profile model.
[0037] In an embodiment of the present application, the robot may first obtain the historical chat records between it and the user, and then input the historical chat records into the trained user portrait model, and the user portrait model may thereby output the user portrait of the user.
[0038] Of course, after obtaining the user portrait of the user output by the user portrait model, the robot can also adjust the user portrait based on the user information of the user to obtain a more accurate user portrait, which is not limited here.
[0039] Step 104 : Determine a target opening statement from the candidate opening statements based on the user profile.
[0040] In this embodiment of the present application, the robot can use the user profile to determine which news the user is most interested in. To motivate the user in subsequent conversations, the robot can consider the news that the user may be interested in from the candidate opening statements based on the user profile and select the target opening statement from the candidate opening statements generated from these news.
[0041] Step 105: Initiate a dialogue with the user based on the target opening statement.
[0042] In this embodiment of the present application, if the target opening statement is a complete sentence, the robot can directly output the target opening statement as the opening statement of the conversation. It is understood that, depending on the interaction mode currently adopted by the robot, the robot can initiate the conversation through voice or text, and the robot can also initiate the conversation through text, which is not limited here.
[0043] In some embodiments, to obtain a trained news recognition model, the robot may perform the following operations:
[0044] First, data is collected, for example, historical news is collected, and the titles, descriptions, and / or texts of the historical news are stored in a database in a structured form by category. In addition, the entities mentioned in the historical news (such as names of people, places, and organizations, etc.) and events can also be stored in the database in a structured form by category. Specifically, the titles, descriptions, and / or texts of historical news can be stored in one table in the database, and the entities and events of historical news can be stored in another table in the database, and these two tables can be linked by a primary key-foreign key. Of course, the data related to historical news can also be stored in the form of files instead of in the database, which is not limited here.
[0045] Then, the model to be used is selected. For example, the robot can use existing entity recognition models as the news recognition model, including but not limited to the Hidden Markov Model (HMM), Conditional Random Field (CRF), Bi-directional Long Short-Term Memory (BiLSTM), and BiLSTM+CRF. Alternatively, the robot can use pre-trained models that offer better results but higher computational overhead as the news recognition model, including but not limited to BERT (Bidirectional Encoder Representation from Transformers), BERT+CRF, and ALBERT (A Lite BERT).
[0046] Next, specify the two types of output of the model and the corresponding two loss functions.
[0047] Among them, the function of one type of output is to identify entities and events, and the function of the other type of output is to identify the agent and the patient. Both types of outputs are annotated in the BIO format.
[0048] Among them, when identifying entities and events, the specific annotation method is as follows: O represents the characters outside the entities and events; B_e represents the first character of an entity; I_e represents the non-first characters of an entity; B_a represents the first character of an event; I_a represents the non-first characters of an event. <x <x
[0049] When identifying the agent and the patient, the specific annotation method is as follows: O represents the characters outside the agent and the patient; B_s represents the first character of the agent; I_s represents the non-first characters of the agent; B_o is the first character of the patient; I_o is the non-first character of the patient. <x <x
[0050] Please refer to <x Figure 2 , <x Figure 2 which gives examples of two different types of outputs obtained based on the same input text. From <x Figure 2 it can be seen that for the input "Shocking! Italy won the European Cup", the first type of output of the news recognition model will identify that: "Shocking", "!", and "了" are characters outside the entities and events; "意" in "意大利" is the first character of an entity, and "大" and "利" are the non-first characters of the entity; "欧" in "欧洲杯" is the first character of an entity, and "洲" and "杯" are the non-first characters of the entity; "夺" in "夺冠" is the first character of an event, and "冠" is the non-first character of the event. The second type of output of the news recognition model will identify that: "Shocking", "!", "夺", "冠", and "了" are characters outside the agent and the patient; "意" in "意大利" is the first character of the agent, and "大" and "利" are the non-first characters of the agent; "欧" in "欧洲杯" is the first character of the patient, and "洲" and "杯" are the non-first characters of the patient. <x <x
[0051] For the first type of output (i.e., output used to identify entities and events), the first loss function is used, and the calculated loss value is recorded as Loss1; for the second type of output (i.e., output used to identify the agent and the recipient), the second loss function is used, and the calculated loss value is recorded as Loss2. Both the first and second loss functions can adopt some of the current mainstream sequence labeling loss functions, such as cross entropy. It should be noted that Loss1 and Loss2 are completely independent, that is, Loss1 is only related to the first type of output, and Loss2 is only related to the second type of output. Generally speaking, for the same input text, the robot can consider the average of Loss1 and Loss2 as the overall loss value; or, it can set the weights of Loss1 and Loss2 according to their importance, so as to obtain the weighted average of Loss1 and Loss2 as the overall loss value, such as the overall loss value Loss = (α*Loss1 +β*Loss2) / 2, where α is the weight of Loss1, β is the weight of Loss2, and α+β=1.
[0052] It can be understood that after the model is selected and the output and loss function are specified, a large number of historical news titles need to be annotated using the annotation method specified by the output of the model to create labels.
[0053] Finally, the model is trained. When the overall loss reaches convergence or the number of training iterations meets a preset threshold, a trained news recognition model is obtained. This trained news recognition model can then be put into use. The training and tuning methods for this model are similar to those used in current mainstream deep learning training and will not be detailed here.
[0054] In some embodiments, the robot may randomly select one of the models that the news recognition model shown above may adopt (such as HMM, CRF, BiLSTM, BiLSTM+CRF, BERT, BERT+CRF, and ALBERT, etc.) as the user profile model. However, different from the news recognition model, this user profile model has only one type of output, and the function of this type of output is: to identify person names, entities, and relationships. Among them, person names belong to a special type of entity, such as you, me, or him, etc. Specifically, person names are used to distinguish and determine whether a certain relationship or entity is related to the user; relationships are used to express the types of user profiles, such as: likes, interests, hobbies, idols, and hometowns that express positive relationships, etc.; entities are used to express the specific attributes of the user profile, such as: military news. This type of output is also annotated in the BIO format, and its annotation method is specifically: B_p represents the first character of a person name; I_p represents the non-first character of a person name; B_r represents the first character of a relationship; I_r represents the non-first character of a relationship respectively; B_e represents the first character of an (non-person name) entity; I_e represents the non-first character of an (non-person name) entity; O represents the characters other than entities and relationships.
[0055] Please refer to Figure 3 , Figure 3 which gives an example of the output of the user profile model. As Figure 3 can be seen, for the input "I really like military news", the output of the user profile model will identify that: "really" and "呢" are the characters other than entities and relationships; "I" is the first character of a person name; "喜" in "like" is the first character of a relationship, and "欢" is the non-first character of a relationship; "军" in "military news" is the first character of an entity, and "事", "新", and "闻" are the non-first characters of an entity.
[0056] It can be understood that the training and tuning methods of this user profile model are also similar to the current mainstream deep learning training methods, which will not be elaborated here. After obtaining the trained user profile model, this trained user profile model can be put into application. During the application process, the input of this trained user profile model is each sentence in the user's historical chat records. This trained user profile model can extract entities, relationships, and person names from each sentence in this historical chat record, and a user profile can be obtained based on this information.
[0057] Specifically, the robot can determine whether the extracted relationships and entities are relevant to the user based on the person. If so, the robot can then add the entities to the user profile table based on the extracted relationships to form a user profile for that user. For example, based on a user's historical chat record "I really like military news," the robot extracts the person "I," the relationship "like," and the entity "military news." Based on the person "I," the robot determines that the relationship "like" and the entity "military news" are relevant to the user. Then, based on the relationship "like" and the entity "military news," the entity "military news" is added to the "Preferences" column of the user's user profile table.
[0058] In some embodiments, based on the news recognition model described above, the key information extracted in step 101 includes the following four categories: entity, event, agent, and recipient. For each piece of news that has occurred on that day, step 102 specifically includes:
[0059] A1. Combine the entities, events, agents, and recipients of the news to obtain all possible combination results.
[0060] A2. Generate candidate opening statements related to the news based on the combination result and a preset template corresponding to the combination result.
[0061] To facilitate understanding, the following is an explanation using specific examples:
[0062] 1. Generate candidate opening statements based on a single entity and a first preset template. The first preset template may be "[entity] is on the headlines today." For example: Zhou is on the headlines today.
[0063] 2. Generate candidate opening statements based on the individual event and a second preset template. The second preset template may be "There is the latest [event] information." For example: There is the latest championship winning information.
[0064] 3. Generate candidate opening statements based on the individual agent and a third preset template. The third preset template may be "[Agent] did something big today."
[0065] 4. Generate candidate opening statements based on the individual recipient and a fourth preset template. The fourth preset template may be "[The recipient] is in trouble today."
[0066] 5. Generate candidate opening statements based on the multiple entities and the fifth preset template. The fifth preset template may be "[Entity 1] and [Entity 2] have melons."
[0067] Of course, there are more combinations and corresponding preset templates, which are not detailed here due to space limitations. As can be seen, by extracting key information from a piece of news, the robot can usually generate multiple candidate opening statements.
[0068] In some embodiments, in order to ensure the rationality of the target opening statement selected by the robot when dealing with various types of users, step 104 may include:
[0069] B1. Determine whether the user profile is empty.
[0070] The input to the user profile model is the user's chat history. If the user is a new user who has not interacted with the robot before, or if the user is a guest user logging in temporarily, then the user's chat history will inevitably be empty. Obviously, when the user's chat history is empty, since the input to the user profile model is empty, its output will also be empty. In other words, in application scenarios where the user is a new user or a guest user, the robot will not be able to obtain a valid user profile for that user.
[0071] B2. If the user profile is empty, determine the target opening statement from the candidate opening statements based on hot news.
[0072] If the user profile is empty, it indicates that the user is a new user or a visitor. Since the user profile is empty, it's impossible to determine the user's news preferences, and therefore, personalized, targeted conversations with the user are no longer considered. In this case, hot news can be filtered from all the news that has already occurred that day, and the target opening statement can be determined based on this hot news. Specifically, any statement from the candidate opening statements generated based on the hot news can be determined as the target opening statement.
[0073] Hot news can be: the N news items with the highest cumulative click-through rate among all news items that have occurred on the day, where N is a preset positive integer; or, the news items with a cumulative click-through rate exceeding a first preset number of clicks among all news items that have occurred on the day; or, the N news items with the highest click-through rate within the most recent unit of time (e.g., the last hour) among all news items that have occurred on the day; or, the news items with a click-through rate exceeding a second preset number of clicks within the most recent unit of time among all news items that have occurred on the day. The method for determining hot news is not limited herein.
[0074] B3. If the user profile is not empty, determine a target opening statement from the candidate opening statements based on the user's preference indicated by the user profile.
[0075] If the user profile is not empty, it can be determined that the user is an experienced user. To motivate the user to interact with the robot, the robot can select news that the user may be interested in based on the user's preferences indicated by the user profile and determine any sentence from the candidate opening sentences generated based on the news as the target opening sentence.
[0076] It is understood that the target opening statement may include only one candidate opening statement. For example, the target opening statement may be: Italy won the championship today. Alternatively, the target opening statement may include two or more candidate opening statements. When the target opening statement includes two or more candidate opening statements, the two or more candidate opening statements may be derived from the same news (i.e., generated based on the same news), or the two or more candidate opening statements may be derived from different news (i.e., generated based on different news). For example, the target opening statement may be: Italy won the championship today, Biden visited China today.
[0077] In some embodiments, step B3 may specifically include:
[0078] B31. Detect whether there is a candidate opening statement that matches the user's preference.
[0079] In some embodiments, based on the user profile model described above, it can be seen that a user's preferences typically include one, two, or more entities. Based on this, the robot can match the entities included in the user's preferences with the various news items that have occurred that day. To improve matching efficiency, the robot can first extract the summaries of each news item, and then match the entities included in the user's preferences with the summaries of the various news items that have occurred that day. If a match is successful, the corresponding news item is the news that matches the user's preferences; correspondingly, the candidate opening statement generated based on the news item is the candidate opening statement that matches the user's preferences.
[0080] B32. If there are candidate opening statement sentences that match the user's preference, determine a target opening statement sentence from the candidate opening statement sentences that match the user's preference.
[0081] B33. If there is no candidate opening statement that matches the user's preference, determine a target opening statement from the candidate opening statements that are associated with the user's preference.
[0082] The robot can find topics that are related to the user's preferences through data mining or knowledge graphs. The news that matches these topics is the news that is related to the user's preferences; correspondingly, the candidate opening statements generated based on the news are the candidate opening statements associated with the user's preferences.
[0083] The following is an explanation of the data mining method:
[0084] The robot can use association rules in data mining to perform entity recognition on a large number of news items, treating each news item as a transaction and assigning a number to each item to obtain a transaction sequence number. For each news item, all entities contained in the content are combined into a data item. This forms a data mining table, as shown in Table 1 below:
[0085]
[0086] Table 1
[0087] Based on the constructed data mining table, the confidence of a particular entity B relative to another entity A can be calculated. Specifically, this represents the frequency with which entity B appears when entity A appears, denoted as {A→B}. In other words, the confidence of entity B relative to entity A is the ratio of the number of transactions containing both A and B to the number of transactions containing A. This formula can be simply expressed as: confidence of {A→B} = P(A|B) / P(A). For illustrative purposes only, this confidence can be calculated using algorithms such as Apriori and is not detailed here.
[0088] Based on the concept of confidence presented above, the robot can perform association rule analysis based on data mining tables to calculate the confidence of a particular entity relative to another entity. If the confidence is sufficiently high, the entity can be added to the relationship of the other entity to the user profile. For example, if "Yao" and "Rockets" appear in a large number of news stories, and the confidence of "Rockets" relative to "Yao" is high, and the user is known to like "Yao," then there is a high probability that the user will also like "Rockets." This allows the robot to find news related to the user's preferences. For example, if the user's profile indicates that the user prefers "Yao," but there is no news related to "Yao" on that day, data mining will reveal that the confidence of "Rockets" relative to "Yao" is high, and there is news related to "Rockets" on that day, then the news related to "Rockets" is the news related to the user's preferences. The target opening sentence can then be identified from the candidate opening sentences generated based on the news related to "Rockets."
[0089] The following is an explanation of the knowledge graph method:
[0090] The robot has multiple pre-generated knowledge graphs. In the knowledge graph, each entity can be linked to other entities through given relationships. When the robot extracts persons, relationships, and entities from historical chat logs, if the relationships match those shown in the knowledge graph, it can transfer information through the knowledge graph to find other entities associated with the user's preferred entity. The news corresponding to these other entities is the news associated with the user's preference. The robot can then determine the target opening statement from the candidate opening statements generated by these news.
[0091] In some embodiments, to prevent the robot from having difficulty understanding the user's response, when designing the wording used in the candidate opening statements generated by the robot, the wording should be designed to guide the user to respond with interest or disinterest, and to prevent the user's response from being too open. For example, the candidate opening statements can be designed to guide the user to obtain a news introduction or summary. For the conversation actively initiated by the robot based on the target opening statement, the user's response can be divided into the following types: positive, meaning interested and wanting to continue chatting to obtain the details of the news; neutral, meaning not interested and wanting to switch to another news; negative, meaning not interested and not wanting to chat about the news. Based on this, the robot can pre-train a classifier, which is set with three categories: positive, neutral and negative. The robot can sort out a large amount of corpus that may contain these three categories and use a commonly used text classification algorithm to train the classifier. Based on the classifier, after step 105, the conversation method also includes:
[0092] C1. After receiving the user's response to the target opening statement, classify the response.
[0093] It is understood that after the robot initiates a conversation based on the target opening statement, the user can respond to the target opening statement. After receiving the response, the robot can classify the response using the classifier proposed above to determine the user's attitude towards the target opening statement.
[0094] C2. If the answer is positive, output the summary of the target news.
[0095] If the user's response is classified as positive, it indicates that the user is interested in further understanding the target news item, which refers to the news item corresponding to the target opening statement. The robot extracts a summary of the target news item and, when the user gives a positive response, outputs this summary to the user as feedback. This way, when a user is interested in a proactive news topic, the robot can provide a concise and clear summary, avoiding the feeling of lengthy discussion and achieving a natural conversation with the user.
[0096] The summary of the target news can be extracted in the following ways:
[0097] A text summarization model is pre-configured in the robot. Its input is a news paragraph, and its output is a summary of that paragraph. This text summarization model uses a seq2seq architecture and can be any of the following network models: recurrent neural networks (RNNs), LSTMs, and transformers, though these are not limited here. For example, when using a transformer network model, the encoder extracts the features of the input paragraph, while the decoder generates a summary related to the features of the paragraph.
[0098] C3. If the answer is a neutral answer, the target opening statement is switched and the process returns to step 105 and subsequent steps based on the switched target opening statement.
[0099] If the user's response is classified as neutral, it indicates that the user wishes to switch to a different news topic. The robot can then search for other news that matches the user's preferences. Alternatively, it can use the knowledge graph or data mining methods described above to find other news related to the user's preferences. The robot can then determine a new target opening statement from the candidate opening statements generated from these other news items, switching to the target opening statement (i.e., the news topic). The robot can then initiate a new conversation with the user based on the switched target opening statement.
[0100] C4. If the reply is negative, the chat system is triggered to continue the conversation with the user.
[0101] If the user's reply is classified as a negative reply, it indicates that the user currently does not want to have a conversation with the robot about the news. At this time, the robot can trigger the operation of the chat system and chat with the user based on the chat system.
[0102] C5. If the response cannot be classified, analyze the response and continue the conversation with the user based on the analysis results.
[0103] If the user fails to respond according to the pre-set script, the response likely cannot be categorized by the classifier as positive, neutral, or negative. In cases where this response cannot be classified, the robot may consider it an open-ended response. In this case, the robot needs to further analyze the response and, based on the analysis, determine which conversation strategy to adopt to continue the conversation with the user. Based on this, step C5 specifically includes:
[0104] C51. Analyze the correlation between the target news and the responses.
[0105] As just one example, the robot can determine the relevance based on the frequency of occurrence of the keyword in the reply in the target news. Obviously, the higher the frequency, the higher the relevance. In other words, the frequency is directly proportional to the relevance.
[0106] Alternatively, the robot may determine the relevance based on the similarity between the reply and the summary of the target news. Obviously, the higher the similarity, the higher the relevance. In other words, the similarity is directly proportional to the relevance.
[0107] C52. If the relevance is higher than the preset relevance threshold, the multi-round dialogue system is triggered to continue the dialogue with the user;
[0108] The robot pre-sets a relevance threshold. When the relevance calculated in step C51 exceeds this threshold, the response is considered to be a response from the user surrounding the target opening statement, indicating that the user is still interested in the target news. Given that the response is open-ended, to achieve a natural conversational experience, the robot can trigger a multi-turn dialogue system to continue the conversation with the user.
[0109] C53. If the relevance is equal to or lower than the relevance threshold, the chat system is triggered to continue the conversation with the user.
[0110] If the relevance calculated in step C51 is not higher than the relevance threshold, the reply is deemed to be largely unrelated to the target news. In other words, the user is not interested in the specific content of the target news. In this case, the robot can trigger the chat system to provide feedback on the user's reply. Specifically, the chat system includes a Q&A pair matching library and a generative model, which will not be detailed here.
[0111] In some embodiments, the multi-turn dialogue system works as follows:
[0112] The core of the multi-turn dialogue system is a trained language model. The robot can combine a summary of the target news, the robot's proactive statements, and the user's responses into a dialogue sequence, which is then fed into the language model for prediction. The language model can employ algorithms such as RNN or GPT (Generative Pre-Training), though this is not a limitation here. It should be understood that in this embodiment of the present application, the structure of the language model itself remains unchanged; the actual improvement lies in the input of the language model, specifically, the inclusion of a summary of the target news in the input dialogue sequence. This enables the robot to provide feedback focused on the target news.
[0113] It can be understood that the statement initiated by the robot in the first round is actually the target opening statement; therefore, after receiving the user's reply to the target opening statement, if the reply is an open-ended reply and the reply is highly relevant to the target news, the robot enters a multi-round dialogue state. The multi-round dialogue system can input a dialogue sequence into the language model based on the first round of dialogue. The dialogue sequence is specifically:
[0114] [CLS] Summary of target news [SEP] Opening statement of target [SEP] User response [SEP]
[0115] It can be understood that in the dialogue sequence, the target opening statement is the statement output by the robot in the first round, and the user's response is the user's response in the first round. Based on this dialogue sequence, the language model can predict the statement to be output by the robot in the second round.
[0116] Specifically, the language model predicts the first word after the conversation sequence, appends it to the conversation sequence, and then predicts the next word. This process repeats until the next word is predicted as [SEP] or a termination criterion is met (for example, the number of words in the predicted sentence reaches a preset threshold). At this point, the robot uses the predicted sentence as the output for this round (i.e., the second round).
[0117] Similarly, when a multi-turn dialogue system predicts the output statement for the Nth round of dialogue, the input dialogue sequence is:
[0118] [CLS] Summary of target news [SEP] Target opening statement [SEP] User's first-round response [SEP] Robot's second-round output statement [SEP] User's second-round response [SEP] Robot's third-round output statement [SEP] User's third-round response [SEP] ... Robot's N-1th-round output statement [SEP] User's N-1th-round response [SEP]
[0119] In some embodiments, before the language model is put into use, the robot needs to train the language model. The training process is briefly described as follows:
[0120] First, pre-build a large amount of sequence-type corpus, such as:
[0121] [CLS] News summary [SEP] Conversation initiated by the robot [SEP] User's response [SEP] The robot continues to speak [SEP] The user continues to speak [SEP]... [SEP]...
[0122] Then, based on the existing sequence, the next possible word can be predicted. Its operating logic is the same as that of the language model during application. Because the robot has pre-built a large amount of sequence-type corpus, through continuous training, the language model can learn the patterns of this type of conversation.
[0123] In some embodiments, there are no restrictions on when the user profile model is run. For example, real-time processing can be performed for all users, regardless of response speed. This real-time processing means that during the conversation between the robot and the user, the latest chat history is input into the user profile model in real time, thereby achieving real-time updates of the user profile. Alternatively, the real-time processing described above can be performed only for new and guest users, while offline processing is performed for existing users. This offline processing means that after the current conversation between the user and the robot ends, the current chat history (i.e., the latest historical chat history) is input into the user profile model, thereby achieving offline updates of the user profile. Of course, offline processing can also be performed for all user profiles, which is not limited here.
[0124] As can be seen from the above, through the embodiments of the present application, the smart device first extracts the key information of all the news that has occurred on the day through the trained news recognition model, then generates candidate opening statements based on the key information, and then uses the trained user portrait model to extract the user's user portrait. Among the candidate opening statements, the target opening statement is determined based on the user portrait. Finally, a conversation can be initiated with the user based on the target opening statement. Through the above process, the smart device can actively engage in a conversation with the user, and the topics of active conversation have the following characteristics: on the one hand, they are generated based on the news that has occurred on the day, allowing users to obtain the latest information; on the other hand, they are filtered based on the user portrait, so that the output topics can capture the user's interests. In this way, the smart device can actively engage in meaningful conversations with the user, fully mobilizing the user's enthusiasm for conversation with the smart device.
[0125] Corresponding to the dialogue method provided above, the embodiment of the present application also provides a dialogue device. Figure 4 As shown, the dialogue device 400 includes:
[0126] The first extraction module 401 is used to extract key information of all news that has occurred on the day using the trained news recognition model;
[0127] A generating module 402 is used to generate candidate opening statements based on the key information;
[0128] The second extraction module 403 is used to extract the user profile of the user through the trained user profile model;
[0129] A determination module 404 is configured to determine a target opening statement from the candidate opening statements based on the user profile.
[0130] The dialogue initiation module 405 is configured to initiate a dialogue with the user based on the target opening statement.
[0131] Optionally, the key information includes: entity, event, agent, and recipient; for each piece of news that has occurred on that day, the generation module 402 includes:
[0132] A key information combination unit is used to combine the entities, events, actors and recipients of the above news to obtain all possible combination results;
[0133] The candidate opening statement generating unit is used to generate candidate opening statement related to the news according to the combination result and the preset template corresponding to the combination result.
[0134] Optionally, the determining module 404 includes:
[0135] A judgment unit, used to judge whether the user profile is empty;
[0136] A first determining unit is configured to determine the target opening statement from the candidate opening statements based on hot news if the user profile is empty;
[0137] The second determining unit is configured to determine the target opening statement from the candidate opening statements based on the user's preference indicated by the user portrait if the user portrait is not empty.
[0138] Optionally, the second determining unit includes:
[0139] a detection subunit, configured to detect whether there is the candidate opening statement that matches the user's preference;
[0140] a first determining subunit configured to determine the target opening statement from among the candidate opening statements that match the user's preference, if there are any such candidate opening statements that match the user's preference;
[0141] The second determining subunit is configured to determine the target opening statement from the candidate opening statements associated with the user's preference if there is no candidate opening statement matching the user's preference.
[0142] Optionally, the dialogue device 400 further includes:
[0143] a classification module, configured to classify, after receiving a response from the user to the target opening statement, the response;
[0144] The processing module is used to output a summary of the target news if the above reply is a positive reply, wherein the above target news is: the news corresponding to the above target opening statement; if the above reply is a neutral reply, switch the above target opening statement and trigger the operation of the dialogue initiation module 405 again; if the above reply is a negative reply, trigger the chat system to continue the dialogue with the above user; if the above reply cannot be classified, analyze the above reply and continue the dialogue with the above user based on the analysis result.
[0145] The above-mentioned processing module is specifically used to analyze the correlation between the above-mentioned target news and the above-mentioned reply when the above-mentioned reply cannot be classified; if the above-mentioned correlation is higher than the preset correlation threshold, the multi-round dialogue system is triggered to continue the dialogue with the above-mentioned user; if the above-mentioned correlation is equal to or lower than the above-mentioned correlation threshold, the chat system is triggered to continue the dialogue with the above-mentioned user.
[0146] Optionally, the above analysis of the correlation between the target news and the reply includes: determining the correlation based on the frequency of occurrence of keywords in the reply in the target news; or determining the correlation based on the similarity between the reply and the summary of the target news.
[0147] As can be seen from the above, through the embodiments of the present application, the smart device first extracts the key information of all the news that has occurred on the day through the trained news recognition model, then generates candidate opening statements based on the key information, and then uses the trained user portrait model to extract the user's user portrait. Among the candidate opening statements, the target opening statement is determined based on the user portrait. Finally, a conversation can be initiated with the user based on the target opening statement. Through the above process, the smart device can actively engage in a conversation with the user, and the topics of active conversation have the following characteristics: on the one hand, they are generated based on the news that has occurred on the day, allowing users to obtain the latest information; on the other hand, they are filtered based on the user portrait, so that the output topics can capture the user's interests. In this way, the smart device can actively engage in meaningful conversations with the user, fully mobilizing the user's enthusiasm for conversation with the smart device.
[0148] Corresponding to the dialogue method provided above, the embodiment of the present application also provides a smart device. Figure 5 The smart device 5 in the embodiment of the present application includes: a memory 501, one or more processors 502 ( Figure 5Only one is shown) and a computer program stored in memory 501 and executable on the processor. Memory 501 is used to store software programs and units. Processor 502 executes the software programs and units stored in memory 501 to perform various functional applications and diagnose, thereby obtaining resources corresponding to the aforementioned preset events. Specifically, when processor 502 executes the computer program stored in memory 501, it implements the following steps:
[0149] Through the trained news recognition model, the key information of all news that has occurred on that day is extracted;
[0150] Generate candidate opening sentences based on the above key information;
[0151] Extract the user's user profile through the trained user profile model;
[0152] Determine a target opening statement from the candidate opening statements based on the user profile.
[0153] Based on the above target opening statement, initiate a dialogue with the above user.
[0154] Assuming that the above is the first possible implementation, in a second possible implementation provided based on the first possible implementation, the key information includes: entity, event, agent, and recipient; for each news item that has occurred on that day, the candidate opening statement generated based on the key information includes:
[0155] Combine the entities, events, agents and recipients of the above news to obtain all possible combination results;
[0156] A candidate opening statement related to the news is generated based on the combination result and a preset template corresponding to the combination result.
[0157] In a third possible implementation provided based on the first possible implementation, determining a target opening statement from the candidate opening statements according to the user profile includes:
[0158] Determine whether the above user portrait is empty;
[0159] If the user profile is empty, then determine the target opening statement from the candidate opening statements based on hot news;
[0160] If the user portrait is not empty, the target opening statement is determined from the candidate opening statements based on the user's preference indicated by the user portrait.
[0161] In a fourth possible implementation provided based on the third possible implementation, determining the target opening statement from the candidate opening statements based on the user's preference indicated by the user profile includes:
[0162] Detecting whether there is the candidate opening statement that matches the user's preference;
[0163] If there are candidate opening statements that match the user's preference, determining the target opening statement from the candidate opening statements that match the user's preference;
[0164] If there is no candidate opening statement that matches the user's preference, the target opening statement is determined from the candidate opening statements that are associated with the user's preference.
[0165] In a fifth possible implementation provided as a basis for the first possible implementation, after initiating a dialogue with the user based on the target opening statement, the processor 502 further implements the following steps when executing the computer program stored in the memory 501:
[0166] After receiving the user's response to the target opening statement, classifying the response;
[0167] If the answer is positive, then output a summary of the target news, wherein the target news is: the news corresponding to the target opening statement;
[0168] If the response is a neutral response, switching to the target opening statement and returning to the step of initiating a conversation with the user based on the target opening statement and subsequent steps based on the switched target opening statement;
[0169] If the above reply is a negative reply, the chat system is triggered to continue the conversation with the above user;
[0170] If the reply cannot be classified, the reply is analyzed, and the conversation with the user is continued based on the analysis result.
[0171] In a sixth possible implementation provided based on the fifth possible implementation, analyzing the reply and continuing the dialogue with the user based on the analysis result includes:
[0172] Analyze the relevance between the above target news and the above responses;
[0173] If the correlation is higher than the preset correlation threshold, the multi-round dialogue system is triggered to continue the dialogue with the user.
[0174] If the correlation degree is equal to or lower than the correlation degree threshold, the chat system is triggered to continue the conversation with the user.
[0175] In a seventh possible implementation provided on the basis of the sixth possible implementation, the analyzing the relevance between the target news and the reply includes:
[0176] Determine the relevance based on the frequency of occurrence of the keywords in the responses in the target news;
[0177] or,
[0178] The relevance is determined based on the similarity between the reply and the summary of the target news.
[0179] It should be understood that in the embodiments of the present application, the processor 502 may be a central processing unit (CPU), and may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.
[0180] The memory 501 may include a read-only memory and a random access memory, and provides instructions and data to the processor 502. A portion or all of the memory 501 may also include a non-volatile random access memory. For example, the memory 501 may also store device category information.
[0181] As can be seen from the above, through the embodiments of the present application, the smart device first extracts the key information of all the news that has occurred on the day through the trained news recognition model, then generates candidate opening statements based on the key information, and then uses the trained user portrait model to extract the user's user portrait. Among the candidate opening statements, the target opening statement is determined based on the user portrait. Finally, a conversation can be initiated with the user based on the target opening statement. Through the above process, the smart device can actively engage in a conversation with the user, and the topics of active conversation have the following characteristics: on the one hand, they are generated based on the news that has occurred on the day, allowing users to obtain the latest information; on the other hand, they are filtered based on the user portrait, so that the output topics can capture the user's interests. In this way, the smart device can actively engage in meaningful conversations with the user, fully mobilizing the user's enthusiasm for conversation with the smart device.
[0182] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the above-mentioned device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, which will not be repeated here.
[0183] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.
[0184] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of external device software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0185] In the embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the system embodiments described above are merely schematic. For example, the division of the above modules or units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0186] The units described above as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0187] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application can implement all or part of the processes in the above-mentioned embodiment method by instructing the associated hardware through a computer program. The above-mentioned computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, it can implement the steps of each of the above-mentioned method embodiments. Among them, the above-mentioned computer program includes computer program code, and the above-mentioned computer program code can be in source code form, object code form, executable file or some intermediate form. The above-mentioned computer-readable storage medium may include: any entity or device capable of carrying the above-mentioned computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer-readable memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal and software distribution medium, etc. It should be noted that the content contained in the above-mentioned computer-readable storage medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable storage media does not include electrical carrier signals and telecommunication signals.
[0188] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.
Claims
1. A conversational approach, It is characterized in that include: Through the trained news recognition model, the key information of all news that has occurred on that day is extracted; Generate candidate opening statements based on the key information; Extract the user's user profile through the trained user profile model; Determining a target opening statement among the candidate opening statements according to the user portrait; Initiating a dialogue with the user based on the target opening statement; Determining a target opening statement among the candidate opening statements according to the user portrait includes: Determine whether the user portrait is empty; If the user portrait is empty, determining the target opening statement from the candidate opening statements based on hot news; If the user portrait is not empty, detecting whether there is a candidate opening statement that matches the user's preference; If there are candidate opening statements that match the user's preference, determining the target opening statement from the candidate opening statements that match the user's preference; If there is no candidate opening statement that matches the user's preference, the target opening statement is determined by searching the candidate opening statements that are associated with the user's preference through data mining or knowledge graph.
2. The dialogue method according to claim 1, It is characterized in that The key information includes: entity, event, agent and recipient; for each news item that has occurred on the day, the candidate opening statement is generated according to the key information, including: Combining the entities, events, agents and recipients of the news to obtain all possible combination results; A candidate opening statement related to the news is generated according to the combination result and a preset template corresponding to the combination result.
3. The conversation method according to claim 1, It is characterized in that After initiating a dialogue with the user based on the target opening statement, the dialogue method further includes: After receiving a response from the user to the target opening statement, classifying the response; If the reply is a positive reply, a summary of the target news is output, wherein the target news is: the news corresponding to the target opening statement; If the reply is a neutral reply, switching the target opening statement and returning to execute the step of initiating a dialogue with the user based on the target opening statement and subsequent steps based on the switched target opening statement; If the reply is a negative reply, the chat system is triggered to continue the conversation with the user; If the reply cannot be classified, the reply is analyzed and the conversation with the user is continued based on the analysis result.
4. The dialogue method according to claim 3, It is characterized in that Analyzing the reply and continuing the dialogue with the user based on the analysis result, comprises: Analyzing the relevance between the target news and the reply; If the correlation is higher than a preset correlation threshold, the multi-round dialogue system is triggered to continue the dialogue with the user; If the association degree is equal to or lower than the association degree threshold, the small talk system is triggered to continue the conversation with the user.
5. The dialogue method according to claim 4, It is characterized in that The analyzing the relevance between the target news and the reply includes: Determining the relevance based on the frequency of occurrence of the keywords in the reply in the target news; or, The relevance is determined based on a similarity between the reply and a summary of the target news.
6. A conversation device, It is characterized in that include: The first extraction module is used to extract key information of all news that has occurred on the day through the trained news recognition model; A generating module, used for generating candidate opening statements according to the key information; The second extraction module is used to extract the user portrait of the user through the trained user portrait model; A determination module, configured to determine a target opening statement from among the candidate opening statements according to the user portrait; A dialogue initiation module, used to initiate a dialogue with the user based on the target opening statement; The determining module comprises: A judging unit, used to judge whether the user portrait is empty; A first determining unit, configured to determine the target opening statement from the candidate opening statements based on hot news if the user portrait is empty; The second determining unit is used to detect whether there is a candidate opening statement matching the user's preference if the user portrait is not empty; if there is a candidate opening statement matching the user's preference, determine the target opening statement from the candidate opening statements matching the user's preference; if there is no candidate opening statement matching the user's preference, determine the target opening statement from the candidate opening statements that are associated with the user's preference found through data mining or knowledge graph.
7. An intelligent device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, It is characterized in that When the processor executes the computer program, the method according to any one of claims 1 to 5 is implemented.
8. A computer-readable storage medium storing a computer program. It is characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Conversational analysis method, device, storage medium and mobile terminal
CN108134876A
Human-machine conversation method, apparatus, electronic device and compute readable medium
CN109284357A
Conversation system integrating FAQ, tasks and active guidance
CN109977208A
Man-machine conversation method and device
CN110704703A