Conversation method and device, electronic equipment and storage medium

By generating and evaluating the content of candidate dialogues, combining the current dialogue and participant information, the timing and content of the agent's speech in multi-party dialogues is determined, and the problem of the agent participating in the dialogue at an inappropriate time is solved, and a more natural and coherent dialogue is achieved.

CN120218248APending Publication Date: 2025-06-27BAIDU COM TIMES TECH (BEIJING) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510344280.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The prior art is difficult to effectively utilize the strength of the agent's speech motivation in multi-party dialogue, which leads to the agent participating in the dialogue at an inappropriate time or content, affecting the naturalness and coherence of the dialogue.

Method used

By generating candidate conversation content and evaluating the intensity of their motivation to speak, combining the current conversation content and participant information of multi-party conversations, we determine whether and how the agent speaks in multi-party conversations, ensuring that the agent participates in the dialogue at the appropriate time and content.

Benefits of technology

It realizes the natural and coherent participation of the agent in multi-party dialogue, improves the fun and effectiveness of the dialogue, and makes full use of the agent's speech motivation intensity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120218248A_ABST
    Figure CN120218248A_ABST
Patent Text Reader

Abstract

The invention discloses a dialogue method and device, electronic equipment and a storage medium, and relates to the field of computers, in particular to the field of artificial intelligence such as large models and deep learning. According to the specific implementation scheme, candidate dialogue contents are generated according to current dialogue contents in a multi-party dialogue; obtaining an evaluation score of the candidate dialogue content; wherein the evaluation score is used for indicating the speaking motivation intensity of the candidate dialogue content; determining a next speaker of the multi-party dialogue according to the current dialogue content; and according to the evaluation score and the next speaker, determining whether the agent speaks in the multi-party dialogue and the speaking content in the multi-party dialogue.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, particularly to artificial intelligence fields such as large models and deep learning, and specifically to a dialogue method, device, electronic device, and storage medium. Background Art

[0002] With the continuous development of artificial intelligence technology, significant progress has been made in the development of large models in the direction of intelligent agent dialogue. For example, intelligent agents can participate in multi-party conversations, provide reference opinions, and enhance the fun of chatting, etc. Summary of the Invention

[0003] This application provides a dialogue method, device, electronic device, and storage medium. The specific solutions are as follows:

[0004] According to one aspect of this application, a dialogue method is provided, including:

[0005] Generating candidate dialogue content according to the current dialogue content in a multi-party conversation;

[0006] Obtaining an evaluation score of the candidate dialogue content; wherein, the evaluation score is used to indicate the intensity of the speaking motivation of the candidate dialogue content;

[0007] Determining the next speaker of the multi-party conversation according to the current dialogue content;

[0008] Determining whether the intelligent agent speaks in the multi-party conversation and the speaking content in the multi-party conversation according to the evaluation score and the next speaker.

[0009] According to another aspect of this application, a dialogue device is provided, including:

[0010] A first generation module, configured to generate candidate dialogue content according to the current dialogue content in a multi-party conversation;

[0011] A first obtaining module, configured to obtain an evaluation score of the candidate dialogue content; wherein, the evaluation score is used to indicate the intensity of the speaking motivation of the candidate dialogue content;

[0012] A first determination module, configured to determine the next speaker of the multi-party conversation according to the current dialogue content;

[0013] A second determination module, configured to determine whether the intelligent agent speaks in the multi-party conversation and the speaking content in the multi-party conversation according to the evaluation score and the next speaker.

[0014] According to another aspect of this application, an electronic device is provided, including:

[0015] At least one processor; and

[0016] A memory communicatively connected to the at least one processor; wherein,

[0017] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described in the above embodiments.

[0018] According to another aspect of the present application, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the method described in the above embodiments.

[0019] According to another aspect of the present application, there is provided a computer program product including a computer program, and the computer program implements the steps of the method described in the above embodiments when executed by a processor.

[0020] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present application, nor is it used to limit the scope of the present application. Other features of the present application will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The drawings are used to better understand the solution and do not constitute a limitation to the present application. Among them:

[0022] Figure 1 is a schematic flowchart of a dialogue method provided by an embodiment of the present application;

[0023] Figure 2 is a schematic flowchart of a dialogue method provided by another embodiment of the present application;

[0024] Figure 3 is a schematic flowchart of a dialogue method provided by another embodiment of the present application;

[0025] Figure 4 is a schematic structural diagram of a dialogue device provided by an embodiment of the present application

[0026] Figure 5 is a block diagram of an electronic device for implementing the dialogue method of the embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0027] The following describes exemplary embodiments of the present application with reference to the accompanying drawings. Various details of the embodiments of the present application are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present application. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0028] It should be noted that in the technical solution of this application, the acquisition, storage, use, processing, etc. of data all comply with the relevant provisions of national laws and regulations and do not violate public order and good customs.

[0029] Next, the dialogue method, device, electronic device, and storage medium of the embodiments of this application will be described with reference to the accompanying drawings.

[0030] Figure 1 It is a schematic flowchart of the dialogue method provided by an embodiment of this application.

[0031] The dialogue method of the embodiments of this application can be executed by the dialogue device of the embodiments of this application, and this device can be configured in an electronic device.

[0032] Among them, the electronic device can be any device with computing capabilities, such as a personal computer, a mobile terminal, a server, etc. The mobile terminal can be, for example, a vehicle-mounted device, a mobile phone, a tablet computer, a personal digital assistant, a wearable device, etc., which are hardware devices with various operating systems, touch screens, and / or display screens.

[0033] As Figure 1 shown, the dialogue method includes:

[0034] Step 101, generate candidate dialogue content according to the current dialogue content in the multi-party dialogue.

[0035] In this application, the participants in the multi-party dialogue include an agent. For example, in a chat group g, there are user A, user B, and agent C.

[0036] Among them, the current dialogue content in the multi-party dialogue can be the content of the most recent or multiple speeches in the multi-party dialogue, and there is no limitation on this. For example, the dialogue in chat group g is as follows:

[0037] User A: "I went camping in the mountains on the weekend. The air is so good!"

[0038] User B: "Sounds good! I've also been thinking about hiking recently."

[0039] Among them, the current dialogue content of chat group g can be the speech content of user B "Sounds good! I've also been thinking about hiking recently.", or it can be the speech content of user A "I went camping in the mountains on the weekend. The air is so good!" and the speech content of user B "Sounds good! I've also been thinking about hiking recently.".

[0040] In this application, candidate dialogue content can be generated according to the current dialogue content by using a large model. Among them, the candidate dialogue content can be one or more, and there is no limitation on this.

[0041] It should be noted that in this application, the participants in the multi-party conversation can be two parties or more than two parties, and there is no limitation on this.

[0042] Step 102: Obtain the evaluation score of the candidate conversation content; wherein, the evaluation score is used to indicate the intensity of the speaking motivation of the candidate conversation content.

[0043] In this application, the intensity of the speaking motivation expressed by the candidate conversation content can be evaluated to obtain an evaluation score. Among them, the higher the evaluation score, the higher the intensity of the speaking motivation.

[0044] Exemplarily, the evaluation score can be obtained according to the relevance between the candidate conversation content and the current conversation content.

[0045] Step 103: Determine the next speaker in the multi-party conversation according to the current conversation content.

[0046] In this application, the current conversation content can be parsed to determine the next speaker in the multi-party conversation.

[0047] For example, if the speech content of user B in the above chat group g does not specify the next speaker, it can be determined that the next speaker can be any participant in the multi-party conversation.

[0048] Another example, in a certain multi-party conversation, the current conversation content is the speech content of user b "I think it's great to go hiking this weekend. What do you think, agent c?" It can be seen that user b asks for the opinion of agent c, and it can be determined that the next speaker in this multi-party conversation is agent c.

[0049] Step 104: Determine whether the agent speaks in the multi-party conversation and the speech content in the multi-party conversation according to the evaluation score and the next speaker.

[0050] As a possible implementation manner, if the next speaker is the agent, then it is determined that the agent speaks in the multi-party conversation, and according to the evaluation score of the candidate conversation content, the speech content of the agent is determined from the candidate conversation content, and then the agent speaks according to the speech content.

[0051] As another possible implementation manner, if the next speaker is any participant in the multi-party conversation, it can be determined whether the agent speaks in the multi-party conversation according to the evaluation score. If it is determined that the agent speaks in the multi-party conversation, then according to the evaluation score, the speech content of the agent is determined from the candidate conversation content, and then the agent speaks according to the speech content.

[0052] In the embodiments of the present application, candidate conversation content is generated according to the current conversation content in a multi-party conversation, and an evaluation score indicating the intensity of the speaking motivation of the candidate conversation content is obtained. The next speaker of the multi-party conversation is determined according to the current conversation content, and then, based on the evaluation score and the next speaker, it is determined whether the intelligent agent will speak and what the intelligent agent will say in the multi-party conversation. Thus, in a multi-party conversation involving an intelligent agent, candidate conversation content is generated based on the current conversation content, and in combination with the intensity of the speaking motivation of the candidate conversation content, it is determined whether the intelligent agent will speak and what to say, thereby fully considering the intensity of the speaking motivation of the intelligent agent, enabling the intelligent agent to actively participate in the conversation at an appropriate time, and achieving a more natural and coherent multi-party conversation.

[0053] Figure 2 It is a schematic flowchart of a conversation method provided in another embodiment of the present application.

[0054] As Figure 2 shown, the conversation method includes:

[0055] Step 201, generate candidate conversation content according to the current conversation content in the multi-party conversation.

[0056] In the present application, candidate conversation content can be generated according to the current conversation content when the multi-party conversation meets the conversation content generation condition.

[0057] Among them, the conversation content generation condition can be any one of the following: there is new information in the multi-party conversation; the pause duration of the multi-party conversation exceeds a preset duration; the intelligent agent is designated to speak; the next speaker is any participant in the multi-party conversation. Among them, the pause duration of the multi-party conversation exceeding the preset duration can be understood as the time interval between the most recent speaking time of the participants in the multi-party conversation and the current time exceeding the preset duration.

[0058] For example, if there is new information in the multi-party conversation or the pause duration of the multi-party conversation exceeds the preset duration, candidate conversation content can be generated according to the current conversation content.

[0059] Exemplarily, second content generation requirement information can be obtained, where the second content generation requirement information can be used to require the generation of content with a complexity less than a first threshold, and according to the current conversation content and the second content generation requirement information, a second prompt message is generated, and the second prompt message is input into a large model, and the large model is guided by the second prompt message to generate content, so as to obtain candidate conversation content with a complexity less than the first threshold.

[0060] Among them, the second content generation requirement information can include, but is not limited to, that the number of words of the generated content is less than a set quantity, the relevance of the generated content to the current conversation content is greater than a relevance threshold, etc.

[0061] For example, the second prompt message is "Generate 2 short responses that are directly related to

I went hiking on the weekend

[0062] Therefore, according to the current conversation content and the requirement information for generating content with a complexity lower than the first threshold, the large model is used to generate candidate conversation contents, so that while generating candidate conversation contents with a lower complexity, the accuracy of the generated content can be improved.

[0063] Step 202: Use evaluation indicators from multiple dimensions to evaluate the speaking motivation intensity of the candidate conversation contents, so as to obtain the score values corresponding to the evaluation indicators of each dimension and the scoring probabilities of the score values.

[0064] In this application, the evaluation indicators from multiple dimensions can include but are not limited to relevance, information difference, urgency, coherence, etc.

[0065] Among them, relevance can refer to the degree of fit between the candidate conversation content and the current conversation content; information difference can refer to the degree to which the candidate conversation content supplements the missing information in the multi-party conversation; urgency can refer to the time sensitivity of the candidate conversation content; coherence can refer to the contribution of the candidate conversation content to the logical fluency of the conversation.

[0066] In this application, for each candidate conversation content, the evaluation indicators from multiple dimensions can be independently scored through an evaluation model, so the evaluation model can generate the score values corresponding to the evaluation indicators from multiple dimensions and the scoring probabilities of the score values.

[0067] Among them, the scoring probability of the score value can refer to the probability that the evaluation model predicts that the score of the evaluation indicator of the candidate conversation content is the score value.

[0068] For example, if the probability that the score of the relevance of a candidate conversation content is 4.5 is 0.8, then the scoring probability is 0.8.

[0069] Step 203: Determine the evaluation score according to the score values and scoring probabilities corresponding to the evaluation indicators from multiple dimensions.

[0070] As a possible implementation, the score values corresponding to the evaluation indicators from multiple dimensions can be weighted according to the scoring probabilities corresponding to the evaluation indicators from multiple dimensions to obtain the evaluation score.

[0071] As the time since the agent's last speech becomes longer, the motivation intensity of the agent's speech may become stronger. In this regard, as another possible implementation, the time interval between the last speech time of the agent in the multi-party conversation and the current time can be determined, and based on the time interval, the current cumulative enhancement factor of the motivation intensity of the agent can be determined. According to the score value, score probability, and cumulative enhancement factor of the motivation intensity corresponding to the evaluation index of each dimension, the index score of the evaluation index of each dimension can be determined, and then based on the index scores of the evaluation indexes of multiple dimensions, the evaluation score can be determined.

[0072] Exemplarily, for the evaluation index of each dimension, the score value, score probability, and cumulative enhancement factor of the motivation intensity can be multiplied to obtain the index score.

[0073] Exemplarily, the index scores of the evaluation indexes of multiple dimensions can be added to obtain the evaluation score.

[0074] For example, the following formula (1) can be used to calculate the evaluation score:

[0075]

[0076] where IM is the evaluation score; p i is the score probability of the i-th evaluation index, and can also be understood as the score probability of the evaluation index of dimension i; s i is the score value of the i-th evaluation index; n is the number of evaluation indexes; d p is the cumulative enhancement factor of the motivation intensity, d p = λ t-τ d p represents the cumulative effect of the time interval from the last speech time of the agent to the current time on the speech motivation, λ is the cumulative rate, λ is greater than 1, t is the current time, and τ is the last speech time of the agent.

[0077] For example, the time interval from the last speech of the agent to the current time is 5, λ = 1.02, and d p is calculated to be 1.104.

[0078] Exemplarily, the average value of the index scores of the evaluation indexes of multiple dimensions can also be calculated, and this average value can be used as the evaluation score.

[0079] Thus, in the process of evaluating the speech motivation intensity, the influence of the agent's long-term silence on the speech motivation intensity is considered, thereby improving the accuracy of the evaluation of the speech motivation intensity.

[0080] Step 204, determine the next speaker in the multi-party conversation according to the current conversation content.

[0081] In this application, step 204 can adopt any implementation manner in the embodiments of this application, so it will not be elaborated here.

[0082] Step 205: Determine whether the agent speaks in the multi-party conversation and the speech content in the multi-party conversation according to the evaluation score and the next speaker.

[0083] Exemplarily, if the next speaker is the agent, it can be determined that the agent speaks in the multi-party conversation, and according to the evaluation score, the first target conversation content is determined from the candidate conversation contents, and then the first target conversation content is used as the speech content of the agent. For example, the candidate conversation content with the highest evaluation score can be used as the first target conversation content.

[0084] Thus, when the next speaker is the agent, the speech content of the agent can be selected from the candidate conversation contents according to the evaluation score, so as to ensure that the agent's speech is appropriate and improve the naturalness of the multi-party conversation.

[0085] Exemplarily, if the next speaker is any participant in the multi-party conversation, it is determined whether there is a candidate conversation content in the candidate conversation contents whose evaluation score is greater than the second threshold. If so, it can be determined that the agent speaks in the multi-party conversation, and the second target conversation content is determined from the candidate conversation contents whose evaluation score is greater than the second threshold, and the second target conversation content is determined as the speech content of the agent.

[0086] For example, any candidate conversation content can be selected from the candidate conversation contents whose evaluation score is greater than the second threshold as the second target conversation content. Or, the candidate conversation content with the highest evaluation score can be selected from the candidate conversation contents whose evaluation score is greater than the second threshold as the second target conversation content.

[0087] Thus, if the next speaker is not specified explicitly, when there is a candidate conversation content whose evaluation score is greater than the second threshold, it can be determined that the agent speaks, and the speech content of the agent is selected according to the evaluation score, so that the speaking opportunity and speech content of the agent are appropriate.

[0088] Exemplarily, if the next speaker is other participants in the multi-party conversation except the agent, it is determined whether there is a candidate conversation content in the candidate conversation contents whose evaluation score is greater than the third threshold. If so, it is determined that the agent speaks during the speech of other participants, and the third target conversation content is determined from the candidate conversation contents whose evaluation score is greater than the third threshold, and the third target conversation content is determined as the speech content of the agent.

[0089] Therefore, if another participant is the next speaker and there is candidate conversation content with an evaluation score greater than the third threshold, then during the speech of the other participant, the agent can actively interrupt the speech of the other participant, thereby improving the initiative of the agent's speech and further improving the naturalness of the conversation.

[0090] It can be seen that the conversation method of the embodiment of the present application can support flexible participation modes in multiple scenarios, including explicit initiative, interruption mechanism, etc.

[0091] Exemplarily, the speech content of the agent can be displayed in a multi-party conversation interface to be shown to each participant in the multi-party conversation.

[0092] Optionally, the speech content of the agent can also be displayed in the multi-party conversation interfaces of the terminals held by some participants in the multi-party conversation to show the speech content of the agent to some participants. Thus, the flexibility of the display of the speech content can be improved.

[0093] For example, in chat group g, user A, user B, and agent C speak in turn. Among them, the speech content of agent C is only displayed in the chat group interface of the terminal held by user B and not in the chat group interface of the terminal held by user A.

[0094] In the embodiment of the present application, by using evaluation indicators in multiple dimensions to evaluate the speech motivation intensity of the candidate conversation content and determining the evaluation score according to the scoring values and scoring probabilities corresponding to the evaluation indicators in multiple dimensions, the accuracy of the evaluation of the speech motivation intensity is improved, and further the accuracy of the speech opportunity and speech content of the agent is improved.

[0095] Figure 3 It is a schematic flowchart of the conversation method provided by another embodiment of the present application.

[0096] As Figure 3 shown, the conversation method includes:

[0097] Step 301, perform a search according to the current conversation content to determine the search content related to the current conversation content.

[0098] As an implementation manner, a search can be performed in the knowledge base according to the current conversation content to determine the search content related to the current conversation content.

[0099] As another possible implementation manner, a search can be performed in the memory information of the agent according to the current conversation content to determine the search content related to the current conversation content.

[0100] Among them, the memory information of the agent can be the data and information acquired, stored, and invoked by the agent during the interaction with the environment. For example, the memory information of the agent can include, but is not limited to, the agent's background knowledge, experience, interests, current conversation information, etc.

[0101] Exemplarily, the target similarity between the current conversation content and the memory information of the agent can be determined, and based on the target similarity and the weight of the memory type to which the memory information belongs, the relevance between the current conversation content and the memory information can be determined. Then, based on the relevance, the retrieval content can be determined from the memory information. Thus, by determining the relevance based on the similarity between the current conversation content and the memory information of the agent and combining the weight of the memory type, and determining the retrieval content based on the relevance, the accuracy of the retrieval result can be improved.

[0102] Among them, the memory type can include long-term memory, short-term memory, etc. For example, the memory information such as the agent's background knowledge, experience, interests, etc. belongs to long-term memory, and the memory information such as the context information of the current conversation content belongs to short-term memory.

[0103] In some examples, the current conversation content can include the original conversation content and the explanatory content of the original conversation content. The first similarity between the original conversation content and the memory information can be determined, and the second similarity between the explanatory content and the memory information can be determined, and based on the first similarity and the second similarity, the target similarity can be determined.

[0104] Among them, the explanatory content can be a higher-level semantic interpretation generated based on the original conversation content. Exemplarily, the original conversation content can be semantically analyzed and summarized by a large model to obtain the explanatory content.

[0105] For example, the original conversation content is the speech content of user A: "I went camping in the mountains on the weekend. The air is so good!" The explanatory content generated by the large model is: "User A shared the camping experience on the weekend and expressed the love for fresh air."

[0106] Thus, the implicit information can be extracted from the original conversation content by the large model, enhancing the understanding ability of the original conversation content. Retrieving based on the explanatory content can improve the accuracy of the retrieval result.

[0107] For determining the target similarity based on the first similarity and the second similarity, for example, the maximum value of the first similarity and the second similarity can be used as the target similarity, or the average value of the first similarity and the second similarity can be used as the target similarity.

[0108] Thus, by determining the target similarity based on the similarities between the original conversation content and the explanatory content of the original conversation content and the memory information respectively, the accuracy of the target similarity can be improved.

[0109] In some examples, the product of the target similarity and the weight of the memory type to which the memory information belongs can be used as the relevance between the current conversation content and the memory information.

[0110] For humans, human memory may decay over time. Therefore, in some examples, a memory decay factor can be introduced. The memory decay factor can represent the degree of decay of the agent's memory information over time. The product of the target similarity, the weight of the memory type to which the memory information belongs, and the memory decay factor can be used as the relevance between the current conversation content and the memory information. Thus, by introducing the memory decay factor to determine the relevance, the accuracy of the relevance can be improved.

[0111] For example, the following formula (2) can be used to calculate the relevance between the current conversation content and the memory information:

[0112] S(x,u)=max(sim(x,u original ),sim(x,u interp ))*w x *d x (2)

[0113] where S(x,u) is the relevance between the current conversation content u and the memory information x; u original is the original conversation content; u interp is the original conversation content u original 's explanatory content; sim(x,u original ) is the cosine similarity between the memory information x and the original conversation content u original ; sim(x,u interp ) is the cosine similarity between the memory information x and the explanatory content u interp ; w x is the weight of the memory type to which the memory information x belongs; d x is the memory decay factor.

[0114] In some examples, the memory information with the highest relevance can be used as the retrieved content.

[0115] In some examples, the memory information can include long-term memory information and short-term memory information. The attribute information of the agent can be obtained, and based on the attribute information of the agent, the long-term memory information can be obtained, and based on the context information of the multi-party conversation, the short-term memory information can be obtained.

[0116] Among them, the attribute information of the agent can include the agent's background knowledge, experience, interests, etc. For example, the attribute information of the agent can be used as the long-term memory information of the agent, and the context information of the multi-party conversation can be used as the short-term memory information of the agent.

[0117] Thus, by obtaining the long-term memory information of the agent according to the attribute information of the agent, and obtaining the short-term memory information of the agent according to the context information of the multi-party conversation, the long-term and short-term memories are combined to retrieve the memory information related to the conversation, and based on the retrieved memory information, candidate conversation content is generated, which can enhance the context relevance and improve the accuracy of the candidate conversation content.

[0118] Step 302: Generate candidate conversation content according to the retrieved content.

[0119] In this application, the context information of the multi-party conversation can be obtained, and candidate conversation content can be generated according to the retrieved content and the context information. Thus, generating candidate conversation content according to the retrieved content and combining the context information can improve the accuracy of the candidate conversation content.

[0120] Exemplarily, first content generation requirement information can be obtained. The first content generation requirement information is used to require the generation of content with a complexity greater than a first threshold, and according to the retrieved content, the context information, and the first content generation requirement information, first prompt information is generated, and then the first prompt information is input into the large model, and the large model is guided by the first prompt information to generate content, so as to obtain candidate conversation content with a complexity greater than the first threshold.

[0121] Among them, the first content generation requirement information can include, but is not limited to, that the number of words of the generated content is less than a set quantity, the relevance of the generated content to the current conversation content is greater than a relevance threshold, etc.

[0122] For example, the first prompt information can be "Generate 3 conversation contents according to the following context (including the current conversation content and the retrieved content). Please ensure that the contents are diverse, conform to the context, and are less than 15 words."

[0123] Thus, according to the current conversation content, the retrieved content, and the requirement information for requiring the generation of content with a complexity greater than the first threshold, the large model is used to generate candidate conversation content through in-depth thinking, so that the depth and naturalness of the generated content can be improved.

[0124] Exemplarily, the response type can be determined according to the current conversation content, the key information can be determined according to the retrieved content and the context information, and then candidate conversation content can be generated according to the key information and the response type.

[0125] Among them, the response type can include confirmation, negation, interest expression, etc. For example, if the current conversation content includes the speech content of user b "I have been considering a self-driving tour recently.", the response type can be determined to be interest expression.

[0126] Among them, the key information can refer to the key information involved in the multi-party conversation.

[0127] For example, the current conversation content includes the speech content of user B: "I have been considering a self-driving tour recently." The retrieved content is "I like outdoor activities, such as self-driving tours". Combining the context information "considering a self-driving tour", the key information can be determined as a self-driving tour. Based on the key information and the response type, the candidate conversation content "A self-driving tour is a good choice" can be generated.

[0128] Thus, by determining the response type according to the current conversation content, generating the key information based on the retrieved content and the above information, and generating a directly relevant short response based on the key information and the response type, the naturalness of the candidate conversation content is improved.

[0129] It should be noted that when generating candidate conversation content, it is possible to generate only candidate conversation content with a higher complexity, or only generate short candidate conversation content with a lower complexity, or generate both candidate conversation content with a higher complexity and short candidate conversation content, and there is no limitation on this.

[0130] Step 303, obtain the evaluation score of the candidate conversation content; wherein, the evaluation score is used to indicate the intensity of the speech motivation of the candidate conversation content.

[0131] Step 304, determine the next speaker of the multi-party conversation according to the current conversation content.

[0132] Step 305, determine whether the agent speaks in the multi-party conversation and the speech content in the multi-party conversation according to the evaluation score and the next speaker.

[0133] In this application, steps 303 - 305 can adopt any implementation manner in the embodiments of this application, so details are not described herein again.

[0134] In the embodiments of this application, by performing a retrieval according to the current conversation content in the multi-party conversation, retrieving content related to the current conversation content, and generating candidate conversation content based on the retrieved content, not only can the relevance between the candidate conversation content and the current conversation content be improved, but also the candidate conversation content can be enriched, thereby improving the coherence, naturalness, etc. of the multi-party conversation.

[0135] To facilitate the understanding of the conversation method in the embodiments of this application, the following is described with examples.

[0136] Suppose the scenario is as follows: The chat group g includes user A, user B, and agent C, and the conversation is as follows:

[0137] User A: "I went camping in the mountains on the weekend. The air is so good!"

[0138] User B: "Sounds good! I have also been considering hiking recently."

[0139] At this time, the task of Agent C is to decide whether to participate in the conversation and select an appropriate timing and content.

[0140] When the speech of User B, "That sounds great! I've also been thinking about hiking recently.", is detected, the agent triggers the candidate conversation content generation process.

[0141] The agent retrieves memory information related to the current conversation content from the following two types of memory information:

[0142] 1. Long-term memory information (including the background knowledge of the AI):

[0143] (1) "I like outdoor activities, especially hiking and camping."

[0144] (2) "I saw a very beautiful maple forest in the mountains recently."

[0145] (3) "I know some hiking routes suitable for beginners."

[0146] 2. Short-term memory information (including the current conversation context):

[0147] (1) User A mentioned "going camping on the weekend".

[0148] (2) User B mentioned "considering hiking".

[0149] The agent calculates the relevance between the memory information and the current conversation content, and selects the memory information related to the current conversation content according to the relevance. For example, the following memory information is selected:

[0150] 1. "I like outdoor activities, especially hiking and camping." (Relevance: 0.8)

[0151] 2. "I saw a very beautiful maple forest in the mountains recently." (Relevance: 0.7)

[0152] After that, the agent generates candidate conversation content based on the retrieved memory information and the conversation context, and generates it in two ways: quick response and in-depth thinking.

[0153] Quick response:

[0154] (1) "Camping is really great!"

[0155] (2) "Hiking is a good idea!"

[0156] In-depth thinking:

[0157] (1) "Last time I went to the mountains and saw a very beautiful maple forest. You might like it."

[0158] (2) "If you need hiking routes, I can recommend several suitable for beginners."

[0159] Among them, the quick response can be generated by the large model according to the current conversation content as described above, or can be generated according to the current conversation content, retrieval content and context information as described above, and this is not limited. Additionally, the in-depth thinking can be generated by the large model according to the first content generation requirement information, retrieval content and context information. Thus, by integrating the quick response and in-depth thinking, the depth and naturalness of the conversation can be improved.

[0160] The agent can use the above evaluation method to evaluate the speech motivation intensity of the 4 generated candidate conversation contents, and the specific evaluation results are shown in Table 1 below:

[0161] Table 1

[0162]

[0163] The candidate conversation content with the highest score in Table 1 is "If you need hiking routes, I can recommend several suitable for beginners.", and the total score is 4.50.

[0164] For example, the next speaker can be any one of user A, user B, and agent C, that is, currently it is an open round and there is no clear round allocation.

[0165] Since it is detected that the total score of 4.50 of the candidate conversation content with the highest score is higher than the participation threshold (set to 4.0), it is determined that agent C will speak, and then this candidate conversation content is converted into a conversation speech:

[0166] Agent C: "If you need hiking routes, I can recommend several suitable for beginners!"

[0167] Therefore, the final conversation of this chat group g is as follows:

[0168] User A: "I went camping in the mountains on the weekend. The air is so good!"

[0169] User B: "Sounds good! I've also been considering hiking recently."

[0170] Agent C: "If you need hiking routes, I can recommend several suitable for beginners!"

[0171] To implement the above embodiments, an embodiment of the present application also proposes a conversation device. Figure 4 It is a schematic structural diagram of the conversation device provided by an embodiment of the present application.

[0172] As Figure 4 shown, this conversation device 400 includes:

[0173] The first generation module 410 is configured to generate candidate conversation content according to the current conversation content in the multi-party conversation;

[0174] The first acquisition module 420 is configured to acquire an evaluation score of the candidate conversation content; wherein, the evaluation score is used to indicate the strength of the speaking motivation of the candidate conversation content;

[0175] The first determination module 430 is configured to determine the next speaker of the multi-party conversation according to the current conversation content;

[0176] The second determination module 440 is configured to determine whether the agent speaks in the multi-party conversation and the content of the speech in the multi-party conversation according to the evaluation score and the next speaker.

[0177] Optionally, the first acquisition module 420 is configured to:

[0178] Evaluate the strength of the speaking motivation of the candidate conversation content by using evaluation indicators in multiple dimensions, so as to obtain the scoring value corresponding to the evaluation indicator of each dimension and the scoring probability of the scoring value;

[0179] Determine the evaluation score according to the scoring value and scoring probability corresponding to the evaluation indicators in the multiple dimensions.

[0180] Optionally, the first acquisition module 420 is configured to:

[0181] Determine the time interval between the most recent speaking time of the agent in the multi-party conversation and the current time;

[0182] Determine the current cumulative enhancement factor of the motivation strength of the agent according to the time interval;

[0183] Determine the index score of the evaluation indicator of each dimension according to the scoring value, scoring probability corresponding to the evaluation indicator of each dimension and the cumulative enhancement factor of the motivation strength;

[0184] Determine the evaluation score according to the index scores of the evaluation indicators in the multiple dimensions.

[0185] Optionally, the first generation module 410 is configured to:

[0186] Retrieve according to the current conversation content to determine the retrieved content related to the current conversation content;

[0187] Generate the candidate conversation content according to the retrieved content.

[0188] Optionally, the first generation module 410 is configured to:

[0189] Determine the target similarity between the current conversation content and the memory information of the agent;

[0190] Determine the relevance between the current conversation content and the memory information according to the target similarity and the weight of the memory type to which the memory information belongs;

[0191] Determine the retrieved content from the memory information according to the relevance.

[0192] Optionally, the first generation module 410 is used for:

[0193] Determine the first similarity between the original conversation content and the memory information;

[0194] Determine the second similarity between the explanatory content and the memory information;

[0195] Determine the target similarity according to the first similarity and the second similarity.

[0196] Optionally, the device may further include:

[0197] A second generation module, configured to perform semantic analysis and summary on the original conversation content through a large model to obtain the explanatory content.

[0198] Optionally, the memory information includes long-term memory information and short-term memory information, and the device further includes:

[0199] A second acquisition module, configured to acquire the attribute information of the agent, and acquire the long-term memory information according to the attribute information;

[0200] A third acquisition module, configured to acquire the short-term memory information according to the context information of the multi-party conversation.

[0201] Optionally, the first generation module 410 is used for:

[0202] Acquire the context information of the multi-party conversation;

[0203] Generate the candidate conversation content according to the retrieved content and the context information.

[0204] Optionally, the first generation module 410 is used for:

[0205] Acquire first content generation requirement information; wherein, the first content generation requirement information is used to require the generation of content with a complexity greater than a first threshold;

[0206] Generate a first prompt message according to the retrieved content, the context information and the first content generation requirement information;

[0207] Input the first prompt information into the large model and use the large model for content generation to obtain the candidate conversation content.

[0208] Optionally, the first generation module 410 is configured to:

[0209] Determine the response type according to the current conversation content;

[0210] Determine the key information according to the retrieved content and the context information;

[0211] Generate the candidate conversation content according to the key information and the response type.

[0212] Optionally, the first generation module 410 is configured to:

[0213] Obtain the second content generation requirement information; wherein, the second content generation requirement information is used to require the generation of content with a complexity less than the first threshold;

[0214] Generate the second prompt information according to the current conversation content and the second content generation requirement information;

[0215] Input the second prompt information into the large model and use the large model for content generation to obtain the candidate conversation content.

[0216] Optionally, the second determination module 440 is configured to:

[0217] In response to the next speaker being the intelligent agent, determine that the intelligent agent speaks in the multi-party conversation;

[0218] Determine the first target conversation content from the candidate conversation content according to the evaluation score;

[0219] Use the first target conversation content as the speech content.

[0220] Optionally, the second determination module 440 is configured to:

[0221] In response to the next speaker being any participant in the multi-party conversation and there being candidate conversation content with an evaluation score greater than the second threshold in the candidate conversation content, determine that the intelligent agent speaks in the multi-party conversation;

[0222] Determine the second target conversation content from the candidate conversation content with an evaluation score greater than the second threshold;

[0223] Determine the second target conversation content as the speech content.

[0224] Optionally, the second determination module 440 is configured to:

[0225] In response to the next speaker being other participants in the multi-party conversation except the intelligent agent, and there being candidate conversation contents in the candidate conversation contents whose evaluation scores are greater than a third threshold, it is determined that the intelligent agent speaks during the speech of the other participants;

[0226] Determine a third target conversation content from the candidate conversation contents whose evaluation scores are greater than the third threshold;

[0227] Determine the third target conversation content as the speech content.

[0228] It should be noted that the above explanation of the embodiment of the conversation method also applies to the conversation device of this embodiment, so it will not be elaborated here.

[0229] In the embodiment of the present application, by generating candidate conversation contents according to the current conversation content in the multi-party conversation, obtaining an evaluation score for indicating the speech motivation intensity of the candidate conversation contents, determining the next speaker of the multi-party conversation according to the current conversation content, and then determining whether the intelligent agent speaks and the speech content of the intelligent agent in the multi-party conversation according to the evaluation score and the next speaker. Thus, in a multi-party conversation involving an intelligent agent, candidate conversation contents are generated based on the current conversation content, and in combination with the speech motivation intensity of the candidate conversation contents, it is determined whether the intelligent agent speaks and the speech content, thereby fully considering the speech motivation intensity of the intelligent agent, enabling the intelligent agent to actively participate in the conversation at an appropriate time, and realizing a more natural and coherent multi-party conversation.

[0230] According to an embodiment of the present application, the present application also provides an electronic device, a readable storage medium, and a computer program product.

[0231] Figure 5 FIG. shows a schematic block diagram of an exemplary electronic device 500 that can be used to implement the embodiments of the present application. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processing, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present application described herein and / or claimed.

[0232] As Figure 5As shown, device 500 includes a computing unit 501, which can perform various appropriate actions and processes according to computer programs stored in a ROM (Read-Only Memory) 502 or computer programs loaded from a storage unit 508 into a RAM (Random Access Memory) 503. In the RAM 503, various programs and data required for the operation of device 500 can also be stored. The computing unit 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An I / O (Input / Output) interface 505 is also connected to the bus 504.

[0233] Multiple components in device 500 are connected to the I / O interface 505, including: an input unit 506, such as a keyboard, a mouse, etc.; an output unit 507, such as various types of displays, speakers, etc.; a storage unit 508, such as a magnetic disk, an optical disc, etc.; and a communication unit 509, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 509 allows device 500 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0234] The computing unit 501 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 501 include but are not limited to a CPU (Central Processing Unit), a GPU (Graphic Processing Units), various dedicated AI (Artificial Intelligence) computing chips, various computing units running machine learning model algorithms, a DSP (Digital Signal Processor), and any appropriate processor, controller, microcontroller, etc. The computing unit 501 executes the various methods and processes described above, such as the dialogue method. For example, in some embodiments, the dialogue method can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 500 via the ROM 502 and / or the communication unit 509. When the computer program is loaded into the RAM 503 and executed by the computing unit 501, one or more steps of the dialogue method described above can be executed. Alternatively, in other embodiments, the computing unit 501 can be configured to execute the dialogue method in any other appropriate manner (e.g., by means of firmware).

[0235] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, FPGAs (Field Programmable Gate Arrays), ASICs (Application-Specific Integrated Circuits), ASSPs (Application Specific Standard Products), SOCs (System On Chip), CPLDs (Complex Programmable Logic Devices), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which may be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.

[0236] The program code for implementing the methods of this application can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine, or entirely on the remote machine or server.

[0237] In the context of this application, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a RAM, a ROM, an EPROM (Electrically Programmable Read-Only Memory), or a flash memory, an optical fiber, a CD-ROM (Compact Disc Read-Only Memory), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0238] To provide for interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (Cathode-Ray Tube) or LCD (Liquid Crystal Display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).

[0239] The systems and techniques described herein can be implemented in a computing system that includes backend components (such as, for example, a data server), or a computing system that includes middleware components (such as, for example, an application server), or a computing system that includes frontend components (such as, for example, a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (such as, for example, a communication network). Examples of a communication network include: a LAN (Local Area Network), a WAN (Wide Area Network), the Internet, and a blockchain network.

[0240] A computer system may include a client and a server. The client and the server are generally far from each other and usually interact via a communication network. The relationship between the client and the server is generated by computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services (Virtual Private Server). The server may also be a server of a distributed system or a server combined with a blockchain.

[0241] According to an embodiment of the present application, the present application also provides a computer program product, which, when executed by an instruction processor in the computer program product, executes the dialogue method proposed in the above embodiments of the present application.

[0242] It should be understood that various forms of the processes shown above can be used, reordering, adding or deleting steps. For example, the steps recited in the present application can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in the present application can be achieved, and no limitations are imposed herein.

[0243] The above specific embodiments do not constitute a limitation on the protection scope of the present application. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the present application shall be included within the protection scope of the present application.

Claims

1. A conversation method, comprising: Generate candidate conversation content based on current conversation content in multi-party conversations; Obtaining an evaluation score for the candidate dialogue content; wherein the evaluation score is used to indicate the intensity of the speech motivation of the candidate dialogue content; Determining the next speaker of the multi-party conversation according to the current conversation content; According to the evaluation score and the next speaker, it is determined whether the agent speaks in the multi-party dialogue and the content of the speech in the multi-party dialogue.

2. The method of claim 1, wherein: The obtaining the evaluation score of the candidate dialogue content includes: Using evaluation indicators of multiple dimensions, the speech motivation intensity of the candidate dialogue content is evaluated to obtain a score value corresponding to the evaluation indicator of each dimension and a score probability of the score value; The evaluation score is determined according to the scoring values ​​and scoring probabilities corresponding to the evaluation indicators of the multiple dimensions.

3. The method of claim 2, wherein: Determining the evaluation score according to the scoring values ​​and scoring probabilities corresponding to the evaluation indicators of the multiple dimensions includes: Determine the time interval between the last speech time of the agent in the multi-party conversation and the current time; Determining the current motivation intensity cumulative enhancement factor of the agent according to the time interval; Determine the indicator score of the evaluation indicator of each dimension according to the score value, score probability and the motivation intensity cumulative enhancement factor corresponding to the evaluation indicator of each dimension; The evaluation score is determined according to the indicator scores of the evaluation indicators of the multiple dimensions.

4. The method of claim 1, wherein: The step of generating candidate conversation content according to the current conversation content in the multi-party conversation includes: Performing a search based on the current conversation content to determine search content related to the current conversation content; The candidate conversation content is generated according to the search content.

5. The method of claim 4, wherein: The searching according to the current conversation content to determine the search content related to the current conversation content includes: Determining a target similarity between the current conversation content and the memory information of the agent; Determining the relevance between the current conversation content and the memory information according to the target similarity and the weight of the memory type to which the memory information belongs; The search content is determined from the memory information according to the relevance.

6. The method of claim 5, wherein: The current dialogue content includes original dialogue content and interpretation content of the original dialogue content, and the determining of the target similarity between the current dialogue content and the memory information of the agent includes: Determining a first similarity between the original conversation content and the memory information; determining a second similarity between the interpretation content and the memory information; The target similarity is determined according to the first similarity and the second similarity.

7. The method of claim 6, further comprising: The original conversation content is semantically analyzed and summarized by a large model to obtain the explained content.

8. The method of claim 5, wherein: The memory information includes long-term memory information and short-term memory information, and the method further includes: Acquire attribute information of the agent, and acquire the long-term memory information according to the attribute information; The short-term memory information is acquired according to the context information of the multi-party conversation.

9. The method of claim 4, wherein: The step of generating the candidate dialogue content according to the search content includes: Obtaining context information of the multi-party conversation; The candidate conversation content is generated according to the search content and the context information.

10. The method of claim 9, wherein: The generating the candidate dialogue content according to the search content and the context information includes: Acquire first content generation requirement information; wherein the first content generation requirement information is used to require generation of content with a complexity greater than a first threshold; Generate first prompt information according to the search content, the context information and the first content generation request information; The first prompt information is input into a large model, and the large model is used to generate content to obtain the candidate dialogue content.

11. The method of claim 9, wherein: The generating the candidate dialogue content according to the search content and the context information includes: Determine a response type according to the current conversation content; Determining key information according to the search content and the context information; The candidate dialogue content is generated according to the key information and the response type.

12. The method of claim 1, wherein: The step of generating candidate conversation content according to the current conversation content in the multi-party conversation includes: Acquire second content generation requirement information; wherein the second content generation requirement information is used to require generation of content with a complexity less than a first threshold; Generate second prompt information according to the current conversation content and the second content generation request information; The second prompt information is input into the big model, and the big model is used to generate content to obtain the candidate dialogue content.

13. The method according to any one of claims 1 to 12, wherein: The step of determining whether the agent speaks in the multi-party dialogue and the content of the speech in the multi-party dialogue according to the evaluation score and the next speaker includes: In response to the next speaker being the agent, determining that the agent is to speak in the multi-party conversation; Determining a first target dialogue content from the candidate dialogue contents according to the evaluation score; The first target dialogue content is used as the speech content.

14. The method according to any one of claims 1 to 12, wherein: The step of determining whether the agent speaks in the multi-party dialogue and the content of the speech in the multi-party dialogue according to the evaluation score and the next speaker includes: In response to the next speaker being any participant of the multi-party dialogue, and the candidate dialogue contents having an evaluation score greater than a second threshold, determining that the agent is to speak in the multi-party dialogue; Determining a second target dialogue content from the candidate dialogue contents having an evaluation score greater than a second threshold; The second target dialogue content is determined as the speech content.

15. The method according to any one of claims 1 to 12, wherein: The step of determining whether the agent speaks in the multi-party dialogue and the content of the speech in the multi-party dialogue according to the evaluation score and the next speaker includes: In response to the next speaker being another participant among all the participants of the multi-party dialogue except the agent, and the candidate dialogue content having the evaluation score greater than a third threshold, determining that the agent is to speak during the process in which the other participants are speaking; Determining a third target dialogue content from the candidate dialogue content having an evaluation score greater than a third threshold; The third target dialogue content is determined as the speech content.

16. A conversation device, comprising: A first generation module, used to generate candidate conversation content according to the current conversation content in the multi-party conversation; A first acquisition module is used to acquire an evaluation score of the candidate dialogue content; wherein the evaluation score is used to indicate the intensity of the speech motivation of the candidate dialogue content; A first determination module, configured to determine the next speaker of the multi-party dialogue according to the current dialogue content; The second determination module is used to determine whether the intelligent agent speaks in the multi-party dialogue and the content of the speech in the multi-party dialogue according to the evaluation score and the next speaker.

17. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 15.

18. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-15.

19. A computer program product, comprising a computer program, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 15.