Dialogue method and apparatus, electronic device, storage medium, and product
By calling robots in user conversations, using machine learning models to analyze conversation content and generate auxiliary information, the problem of existing intelligent conversation robots lacking richness and fun in personalized online social scenarios is solved, and the richness and efficiency of conversations are improved.
Patent Information
- Application Number
- PCT/CN2024/084795
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-03-29
- Publication Date
- 2025-10-02
AI Technical Summary
Existing intelligent conversational robots cannot be effectively applied in each user's personalized online social scenarios, and lack the richness and fun of conversations.
By calling a robot in a conversation between a user and a second user, auxiliary conversation information is generated and displayed, and a machine learning model is used to analyze the conversation content to select a suitable robot, providing direct participation or reference information to enrich the conversation content.
The richness and interest of the conversation between the user and the second user are improved, and the smoothness and efficiency of the conversation are enhanced.
Smart Images

Figure CN2024084795_02102025_PF_FP_ABST
Abstract
Description
Dialogue method, device, electronic device, storage medium and product Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular to a dialogue method, device, electronic device, storage medium, and product. Background Art
[0002] With the development of artificial intelligence and machine learning technologies, machine learning models can be used to implement intelligent conversational robots. For example, intelligent conversational robots can serve as intelligent customer service representatives or virtual friends, receiving inquiries from users and providing feedback on the answers.
[0003] Summary of the Invention
[0004] This summary is provided to briefly introduce concepts that will be described in detail in the detailed description below. This summary is not intended to identify key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.
[0005] According to some embodiments of the present disclosure, a conversation method is provided, including: displaying a first conversation between a first user and a second user on a client of the first user; calling one or more robots for the first user based on information of the first conversation; generating auxiliary conversation information of the robots based on the first conversation; and displaying the auxiliary conversation information.
[0006] According to some embodiments of the present disclosure, a conversation device is provided, including: a first display module, configured to display a first conversation between a first user and a second user on a client of the first user; a calling module, configured to call one or more robots for the first user based on information of the first conversation; a generating module, configured to generate auxiliary conversation information of the robot based on the first conversation; and a second display module, configured to display the auxiliary conversation information.
[0007] According to some embodiments of the present disclosure, an electronic device is provided, including: a memory; and a processor coupled to the memory, wherein the processor is configured to execute the conversation method of any embodiment described in the present disclosure based on instructions stored in the memory.
[0008] According to some embodiments of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the program is executed by a processor, the conversation method of any embodiment described in the present disclosure is performed.
[0009] According to some embodiments of the present disclosure, a computer program product is provided. When the computer program product is run on a computer, the computer is enabled to implement the dialogue method of any embodiment described in the present disclosure.
[0010] According to some embodiments of the present disclosure, a computer program is provided, comprising: instructions, which, when executed by a processor, cause the processor to perform the dialogue method of any embodiment described in the present disclosure.
[0011] Other features, aspects and advantages of the present disclosure will become apparent from the following detailed description of exemplary embodiments of the present disclosure with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The preferred embodiments of the present disclosure are described below with reference to the accompanying drawings. The drawings described herein are used to provide a further understanding of the present disclosure. Each of the drawings, together with the following detailed description, is included in this specification and forms a part of the specification to explain the present disclosure. It should be understood that the drawings described below only relate to some embodiments of the present disclosure and do not constitute a limitation of the present disclosure. In the drawings:
[0013] FIG1 shows a flow chart of a conversation method according to some embodiments of the present disclosure.
[0014] FIG2 shows a method for calling a robot according to some embodiments of the present disclosure.
[0015] FIG3 shows a method for calling a robot according to other embodiments of the present disclosure.
[0016] FIG4 illustrates a method for calling a robot according to yet other embodiments of the present disclosure.
[0017] [Corrected 20.05.2024 according to Rule 91] Figures 5A to 5C show schematic diagrams of dialogue interfaces according to some embodiments of the present disclosure.
[0018] FIG6 shows a schematic diagram of a conversation interface according to some embodiments of the present disclosure.
[0019] FIG7 shows a schematic structural diagram of a conversation device according to some embodiments of the present disclosure.
[0020] FIG8 shows a schematic structural diagram of an electronic device according to some embodiments of the present disclosure.
[0021] FIG9 shows a schematic structural diagram of a computer system according to some embodiments of the present disclosure.
[0022] It should be understood that, for ease of description, the dimensions of the various parts shown in the drawings are not necessarily drawn to scale. The same or similar reference numerals are used throughout the drawings to indicate the same or similar parts. Therefore, once an item is defined in one drawing, it may not be discussed further in subsequent drawings. DETAILED DESCRIPTION
[0023] The following will be combined with the accompanying drawings in the embodiments of the present disclosure to clearly and completely describe the technical solutions in the embodiments of the present disclosure. However, it is obvious that the embodiments described are only some embodiments of the present disclosure, rather than all embodiments. The following description of the embodiments is actually only illustrative and is in no way intended to limit the present disclosure and its application or use. It should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments described herein.
[0024] It should be understood that the various steps described in the method embodiments of the present disclosure can be performed in different orders and / or performed in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect. Unless otherwise specifically stated, the relative arrangement, numerical expressions and numerical values of the parts and steps set forth in these embodiments should be interpreted as being merely exemplary and do not limit the scope of the present disclosure.
[0025] As used in this disclosure, the term "include" and its variations are intended to be open-ended terms that include at least the following elements / features but do not exclude other elements / features, i.e., "including but not limited to." Furthermore, the term "comprise" and its variations are intended to be open-ended terms that include at least the following elements / features but do not exclude other elements / features, i.e., "including but not limited to." Therefore, "include" and "include" are synonymous. The term "based on" means "based, at least in part, on."
[0026] Reference throughout this specification to "one embodiment," "some embodiments," or "an embodiment" means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the present invention. For example, the term "one embodiment" means "at least one embodiment," the term "another embodiment" means "at least one additional embodiment," and the term "some embodiments" means "at least some embodiments." Furthermore, the appearances of the phrases "in one embodiment," "in some embodiments," or "in an embodiment" in various places throughout this specification are not necessarily all referring to the same embodiment, but may.
[0027] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules, or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules, or units. Unless otherwise specified, concepts such as "first" and "second" are not intended to imply that the objects described in such a manner must be in a given order in time, space, ranking, or any other manner.
[0028] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".
[0029] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.
[0030] The following detailed description of the embodiments of the present disclosure is provided in conjunction with the accompanying drawings, but the present disclosure is not limited to these specific embodiments. The following specific embodiments may be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments. In addition, in one or more embodiments, specific features, structures, or characteristics may be combined in any suitable manner that will be apparent to those skilled in the art from this disclosure.
[0031] It should be understood that the present disclosure does not limit how the image to be applied / processed is obtained. In one embodiment of the present disclosure, the image can be obtained from a storage device, such as an internal memory or an external storage device. In another embodiment of the present disclosure, the image can be captured by mobilizing a camera assembly. It should be noted that the image obtained can be a captured image or a frame of a captured video, and is not particularly limited to this.
[0032] In the context of this disclosure, an image may refer to any of a variety of images, such as a color image, a grayscale image, and the like. It should be noted that, in the context of this specification, the type of image is not specifically limited. Furthermore, the image may be any suitable image, such as a raw image obtained by a camera, or an image that has undergone specific processing, such as preliminary filtering, anti-aliasing, color adjustment, contrast adjustment, normalization, and the like. It should be noted that preprocessing operations may also include other types of preprocessing operations known in the art, which will not be described in detail here.
[0033] In the related art use of intelligent conversational bots, the bots are primarily provided by application developers and serve as a conversational partner with users, responding to messages sent by users. The bots can respond to user messages based on user input and their own knowledge base. In other words, this application scenario involves a one-on-one conversation between a user and the bot. Bots cannot be applied to personalized online social interactions for each user.
[0034] An embodiment of the present disclosure provides a dialogue, which can call a robot for the first user to assist the first user in having a dialogue with the second user through information of the dialogue between the first user and the second user.
[0035] Figure 1 shows a flow chart of a conversation method according to some embodiments of the present disclosure. As shown in Figure 1 , the conversation method of this embodiment includes steps S102 to S108.
[0036] In step S102, a first conversation between the first user and the second user is displayed on the client of the first user.
[0037] The client of the first user refers to a client that is logged into the account of the first user, or a client used by the first user. The client can provide the first user with a function of online chatting with other users, for example, it can be an application client based on a social network or instant messaging, or it can also be another type of client that includes an online chatting module.
[0038] The second user is the user who is chatting with the first user. The second user can be a real user, such as a contact of the first user, or another user in the application. Of course, the second user can also be a virtual user, such as an agent created by the first user or another user, or an agent provided by the application.
[0039] A first conversation between a first user and a second user may include messages sent by both the first user and the second user. It may also include only messages sent by the second user due to the first user not checking and replying to the messages in a timely manner, or it may include only messages sent by the first user. Furthermore, it may also include messages sent by other users. That is, the embodiments of the present disclosure are applicable to scenarios where the first user chats with a single user, as well as scenarios where the first user chats with multiple users in a group.
[0040] In step S104, based on the information of the first conversation, one or more robots are called for the first user.
[0041] The information of the first conversation includes at least one of the content information and attribute information of the first conversation. Content information can be understood as the body of the conversation, namely, the text, voice, image, video, emoticon, link, and other content sent by the user. Attribute information includes, for example, attribute information of the conversation content, such as the sending time and text length, and may also include attribute information of the sender of the conversation, such as the sender's information and the relationship between the senders.
[0042] A robot can generate corresponding content based on conversations sent by other entities in a conversation scenario (e.g., the first user, the second user, and other users or robots participating in the conversation). It can be implemented using software, hardware, or a combination of software and hardware. A robot can also be referred to as a digital human or a virtual agent of a machine learning model. A robot can be implemented based on a machine learning model, such as a Large Language Model (LLM) or a Foundation Model. A machine learning model can be a generative model, which is used to output target content based on input information. The input information of a generative model includes the processing basis of the generative model during the generation process, such as the reference information used to execute the generation process, the requirements for the target output content, and so on. Generative models include, for example, models that generate content based on text or images, and the output of a generative model can include text, images, or a combination of the two. Of course, the input or output of a generative model can also be data in other modalities, such as audio, video, or a combination of multiple types of data. A generative model can be a single-modal model, such as a model that generates text from text (referred to as a "text-to-text model") or a model that generates images from images (referred to as a "image-to-image model"). Alternatively, a generative model can be a cross-modal model, that is, a model whose input and output belong to different modalities, such as a model that generates images from text (referred to as a "text-to-image model"). Alternatively, the input of a generative model can include multiple modalities, and the output can also include multiple modalities.
[0043] When invoking a bot, one or more bots corresponding to the information in the first conversation can be selected from one or more candidate bots. The invoked bots can be created by the first user, another user, or an application. The candidate bots can include bots of various types and uses, and each bot can generate a response based on its pre-configured settings, response strategy, and other information.
[0044] The bot can directly participate in a conversation between a first user and a second user. It can also generate reference information for the first user to use in the conversation. For example, the bot can display generated content only to the first user, who can then copy or edit the content and send it to the second user.
[0045] Robots can be divided into various categories based on their type, applicable scenarios, and usage methods. For example, robots can include virtual avatar robots, which serve as the virtual or digital avatar of the first user and can imitate some of the first user's habits and characteristics, thereby participating in conversations on behalf of the first user in specific scenarios. For another example, robots can be divided into tool robots, knowledge robots, and emotion robots based on their applicable scenarios. Tool robots can help users use other functions provided in the application, such as booking flights, booking hotels, and online shopping. Knowledge robots have a knowledge base and can provide users with factual information. Emotion robots can analyze the emotions in the user's conversation, provide emotional assistance to the user through the conversation, or liven up the chat atmosphere based on the relationship between the first user and other users.
[0046] After determining the robot to be called, the calling process can be carried out after obtaining the user's consent. Alternatively, the user can provide unified confirmation in advance. After the user confirms, the robot can be automatically called according to its configured logic and strategy.
[0047] In step S106 , the robot's auxiliary dialogue information is generated based on the first dialogue.
[0048] Assisted conversation information refers to information used to facilitate a conversation between a first user and a second user, making the conversation between the first and second users smoother. This information can be the message sent by the robot itself during the conversation, or it can be reference information provided to the first user. In either form, it can include at least one of reference information associated with the first conversation, a summary of the first conversation, and a reply to the second user.
[0049] The robot can use its underlying machine learning model to generate auxiliary dialogue information. For example, part or all of the content from the first dialogue can be input into the robot's machine learning model. This input information can also include instructions, such as requirements for the content, format, and style of the generated information. After obtaining the model's output, the output information, or processed output information, can be used directly as auxiliary dialogue information.
[0050] In step S108, auxiliary dialogue information is displayed.
[0051] The auxiliary conversation information is displayed in the first user's client. For example, it can be displayed in the conversation interface between the first user and the second user. The auxiliary conversation information can be visible to both the first user and the second user, or only to the first user. In some embodiments, when the auxiliary conversation information is displayed, the robot's identifier (e.g., avatar, name, etc.) can also be displayed in conjunction to indicate the source of the auxiliary conversation information.
[0052] When a robot participates in a conversation, it can be displayed separately from real users to let the second user know that the robot is currently in the conversation. For example, it can add labels such as "AI" or "robot" or annotate its messages in a different style.
[0053] In the above embodiment, during the conversation between the first user and the second user, the robot is automatically called to assist the first user in the conversation using the information of the conversation, thereby enriching the content of the conversation by using the robot, and improving the richness and interest of the conversation between the first user and the second user.
[0054] The present disclosure provides the following two exemplary methods for a robot to assist a first user in a conversation. Of course, those skilled in the art can use other methods as needed, and the present disclosure will not describe them one by one.
[0055] In one exemplary embodiment, a robot participates in a conversation between a first user and a second user. Specifically, the robot can act as a subject in the conversation, sending messages to the first and second users. This creates a group chat between the first and second users, which can optionally include more subjects. In this embodiment, auxiliary conversation information is information sent by the robot while participating in the conversation between the first and second users. Specifically, after generating the auxiliary conversation information, the robot can send it as its own message, making the auxiliary conversation information also part of the conversation between the first and second users and the robot. This directly enriches the content of the conversation.
[0056] In one exemplary embodiment, the robot does not directly participate in the conversation between the first and second users, but instead provides reference information to the first user. This reference information can be visible to both the first and second users, or only to the first user. In this embodiment, the auxiliary conversation information is provided to the first user for reference. Thus, while assisting the first user, it can reduce the impact on the nature of the conversation between the first and second users, and the robot can serve as a behind-the-scenes guide for the first user, without appearing in the conversation.
[0057] The above two methods can be selected manually by the user or automatically based on the conversation scenario. In some embodiments, based on information from the first conversation, the conversation scenario between the first user and the second user is determined; in response to the conversation scenario belonging to the first type, one or more robots are activated for the first user to participate in the conversation between the first user and the second user; in response to the conversation scenario belonging to the second type, one or more robots are activated for the first user to provide reference information.
[0058] The conversation scenario is used to help determine whether the robot needs to directly participate in the conversation between the first user and the second user. In the first type of scenario, the robot needs to directly participate in the conversation, for example, the relationship between the first user and the second user is close, the relationship between the first user and the second user is not close and the importance is low (for example, chatting with strangers), the topic of the first user and the second user is aimless, etc. In the second type of scenario, the robot may not need to directly participate in the conversation, for example, the relationship between the first user and the second user is not close and the importance is high, the chat content between the first user and the second user involves professional fields, etc. When determining the conversation scenario, the multiple dimensions included in the information of the first conversation can be input into the machine learning model, or matched with the pre-set conditions for each scenario to obtain the classification results of the conversation scenario.
[0059] After describing the application scenarios of the robots provided by the embodiments of the present disclosure, the following further describes the calling methods of the robots of some embodiments of the present disclosure.
[0060] The embodiment shown in Figure 2 describes a method for invoking a robot from a semantic perspective. Figure 2 illustrates a method for invoking a robot according to some embodiments of the present disclosure. In this embodiment, the information in the first conversation includes semantic information. As shown in Figure 2, the invoking method of this embodiment includes steps S202 to S204.
[0061] In step S202, based on the semantic information of the first dialogue, one or more robots corresponding to the first dialogue are determined.
[0062] The semantic information reflects the main content of the first conversation, which may include one or more of a summary, key content, or theme of the conversation. In some embodiments, a semantic analysis model may be used to process all or part of the conversation in the first conversation to obtain the semantic information.
[0063] Several exemplary methods for determining the target dialogue in the first dialogue are provided below to determine the semantic information of the first dialogue based on the target dialogue. Of course, those skilled in the art may also use other methods to determine the semantic information as needed.
[0064] In some embodiments, the last conversation in a first conversation with a specified number of messages is used as the target conversation to determine the semantic information of the first conversation based on the target conversation. The specified number of conversations can be conversations of a specified length (i.e., a specified number of words), conversations containing a specified number of messages, or conversations containing a specified number of rounds (one or more messages sent consecutively by the same user belong to the same round). Thus, by performing semantic analysis on the most recently generated messages with a certain amount of information, the semantic information of the current first conversation is obtained.
[0065] In some embodiments, the last conversation in the first conversation, which occurred within a specified time period, is used as the target conversation to determine the semantic information of the first conversation based on the target conversation. This approach also performs semantic analysis on the most recently generated message, but it more specifically identifies reference content from a temporal perspective and uses this content to determine the semantic information of the current first conversation.
[0066] In some embodiments, conversations within the first conversation that share the same topic are targeted as target conversations, allowing semantic information from the first conversation to be determined based on the target conversations. This approach categorizes messages by topic, allowing the focus of the conversation to be used as a basis for extracting semantic information. If the first conversation covers multiple topics, the conversation with the largest amount of information can be identified as the target conversation. Because chats between users are typically casual, with users occasionally switching topics and quickly returning to the original topic, this approach can extract the most important information from the first conversation.
[0067] The semantic information may reflect a summary of the first conversation, or may reflect key information in the first conversation. Alternatively, it may reflect both types of information. The specific type of information to be used may be determined based on the topic of the information in the first conversation.
[0068] If the topic of the first conversation is relatively simple and clear, summary information can be generated to reflect the semantics of the first conversation. In some embodiments, the topics covered in one or more conversation rounds in the first conversation are determined; in response to the number of topics covered in one or more conversation rounds not exceeding a specified value, the summary information of the first conversation is determined as the semantic information of the first conversation. Because the semantics expressed in a conversation round are generally relatively continuous, this embodiment uses conversation rounds as the minimum unit to count the topics covered in one or more conversation rounds. This conversation round or rounds may be the most recent conversation in the first conversation. Of course, other units of statistics can be used as needed, such as word count, number of messages, etc., which will not be discussed here. If the topics covered in one or more conversation rounds are not greater than a specified value, it indicates that the topics covered in these conversations are relatively clear. Therefore, generating summary information can more simply, accurately, and comprehensively reflect the semantics of the conversation.
[0069] In some embodiments, key information is extracted from the first conversation as semantic information of the first conversation. Key information is, for example, keywords and key sentences. Key information can be determined by matching a preset key information table, or by analyzing core words and core sentences in the conversation. Key information can be extracted from the first conversation as semantic information when the number of topics involved in one or more rounds of conversations is greater than a specified value. That is, when the topics involved in the conversation are relatively scattered and unclear, key information can be extracted to reflect the key content under each scattered topic. Of course, when the number of topics involved in one or more rounds of conversations is not greater than a specified value, key information can also be extracted as semantic information.
[0070] Based on the semantic information, a robot that matches the first conversation can be determined. In some embodiments, the semantic information of the first conversation is matched with the attributes of one or more candidate robots, and one or more robots corresponding to the first conversation are determined based on the matching results. The attributes of a robot can be type, setting information, or description information. When matching the semantic information with the attributes of the robot, the degree of matching can be determined by calculating the similarity between the semantic information and the attributes. For example, a specified number of robots with the highest matching degree, or robots with a matching degree above a specified threshold, are determined as matching robots. As a result, the called robot can more accurately respond to the content currently being discussed by the first and second users.
[0071] In some embodiments, whether to select a robot corresponding to the semantic information of the first conversation can be determined based on the number of topics involved in the first conversation. For example, in response to the number of topics involved in the semantic information being no greater than a specified value, the semantic information of the first conversation is matched with the attributes of one or more candidate robots, and one or more robots corresponding to the first conversation are determined based on the matching results; in response to the number of topics involved in the semantic information being greater than a specified value, or the semantic information involved in the first conversation being ambiguous, the conversation guiding robot is determined to be the robot corresponding to the first conversation. That is, when the number of topics involved in the first conversation is small (e.g., no greater than a specified value), it indicates that the topic of the first conversation is clear, and a robot semantically matching it can be called based on the semantic information of the first conversation. However, when the number of topics involved in the first conversation is large (e.g., greater than a specified value), or when the semantics are ambiguous, such as when a topic cannot be extracted, it indicates that the current conversation between the first user and the second user does not have a clear topic, and a conversation guiding robot can be used to introduce a topic to liven up the chat atmosphere between the two.
[0072] A conversation guidance robot is used to stimulate further communication between a first user and a second user using content that includes new topics. For example, when the conversation guidance robot is invoked, semantic information association information can be generated based on the semantic information of the first and second conversations, as well as information about the first and second users. Based on this association information, a message is generated for the conversation guidance robot. The semantic association information can be obtained by expanding the semantic information, or by determining topics associated with the semantic information and then generating association information based on the associated topics. When generating association information, information about the first and second users, such as their attributes or the relationship between them, can be referenced. For example, the first and second users are engaged in aimless small talk, and the first conversation consists of the following: "Have you eaten?" "Yes," "Oh," "Hmm," and other nonsensical expressions. In this case, the conversation guidance robot can be invoked to participate in the conversation. For example, if the topic of eating was previously mentioned, the conversation guidance robot can send a message such as "Who are you having dinner with? A female colleague?" to liven up the conversation. Alternatively, the conversation guidance robot can provide the first user with some conversation assistance information for reference, which is only visible to the first user, such as "Ask her if she has eaten" or "We can go eat XXX together next time", etc., so that the first user can continue the conversation with the second user based on the inspiration.
[0073] In step S204, one or more robots corresponding to the first conversation are called for the first user.
[0074] The above embodiment calls the robot based on the semantic information of the first conversation, so that the called conversation robot can match the content discussed by the first user and the second user, thereby improving the effect of the robot in assisting the first user in the conversation, thereby further improving the richness and interest of the conversation between the first user and the second user.
[0075] In some embodiments, it is further possible to determine whether the first conversation corresponds to a service intent based on semantic information. For example, the application provides some service functions, such as booking air tickets, booking hotels, online shopping, etc. If the subject or key information of the first conversation can match these service functions, it is considered that the first conversation has a service intent. In response to the first conversation corresponding to the service intent, auxiliary conversation information including at least one of a product card and a service subscription card is generated. In this way, the first user, or the first user and the second user, can quickly access the service function during the conversation, which can assist users in improving the efficiency of using the service function.
[0076] The embodiment shown in Figure 3 describes a method for invoking a robot from the perspective of the second user's category. Figure 3 illustrates a method for invoking a robot according to other embodiments of the present disclosure. In this embodiment, the information in the first conversation includes the second user's category. As shown in Figure 3, the invoking method of this embodiment includes steps S302 to S304.
[0077] In step S302, the category of the second user is determined. The category of the second user can be determined based on whether the second user is a contact of the first user, or based on a category label or grouping set for the second user by the first user. Those skilled in the art may also use other methods to determine the category, which will not be described in detail here.
[0078] In step S304, in response to the second user's category being the designated category, a conversation guiding robot is called for the first user to participate in the conversation between the first user and the second user.
[0079] The designated category can be set by the user. For users of the designated category, the first user hopes to invite the dialogue guiding robot to participate in the chat to liven up the atmosphere when having a conversation with them. For example, when the first user is chatting with a close friend, the first user hopes that the dialogue guiding robot will participate to liven up the atmosphere and keep the conversation going. In this case, the appearance condition of the dialogue guiding robot can be pre-configured to be a conversation between the first user and a second user of the designated category, so that when the first user and the second user of the designated category have a conversation, the robot automatically appears and participates in the conversation. The auxiliary dialogue information of the dialogue guiding robot is the message sent by the dialogue guiding robot during the process of participating in the conversation. The method of generating the auxiliary dialogue information can refer to the aforementioned embodiment and will not be repeated here.
[0080] The above embodiment calls a robot according to the type of the second user, so that the called robot can be automatically determined for the designated chat partner of the first user, thereby improving the fun of the chat between the first user and the second user.
[0081] The embodiment shown in Figure 4 describes a robot invocation method from the perspective of the conversation state. Figure 4 illustrates a robot invocation method according to yet other embodiments of the present disclosure. In this embodiment, the information of the first conversation includes the conversation time. As shown in Figure 4, the invocation method of this embodiment includes steps S402 to S404.
[0082] In step S402, the sender and sending time of the last conversation in the first conversation are determined.
[0083] In step S404, in response to the last conversation in the first conversation being sent by the second user and the sending time of the last conversation exceeding the specified threshold from the current time, a robot serving as the first user's virtual avatar is called for the first user to reply to the second user.
[0084] If the sender of the last conversation is the first user, it means that the first user has already replied to the message sent by the second user. If the sender of the last conversation is the second user, it means that the first user has not yet replied to the message sent by the second user. If the last message was sent some time ago, it means that the first user may not have time to check the message or does not want to reply to the second user. In this case, a robot that can replace the first user can be called to reply to the second user, so that the second user can receive a response in a timely manner, improving the smoothness of communication between the first and second users.
[0085] The first user's virtual avatar can be a robot created by the first user to communicate with other users on his or her behalf. The attributes of the virtual avatar can be configured by the first user, and some attributes can also be system defaults when it is created.
[0086] In some embodiments, the avatar of the virtual avatar robot can be generated based on a picture specified by the first user. The user can upload an image from a local or cloud gallery, or select from an image gallery provided by the application. After the user specifies the picture, it can be processed to obtain a picture with added effects as the robot's avatar. For example, a real-life photo uploaded by a user can be processed into a picture with a comic effect. In the process of processing the image, a graph-to-graph model can be used for processing. For example, the user-specified picture and the processing instructions entered by the user (such as "generate a comic effect image for me" and "make the eyes bigger") can be input into the model; alternatively, the style or filter template provided by the application can be used.
[0087] In some embodiments, the voice of the virtual avatar robot is generated based on a voice specified by the first user. The user can upload a voice from a local or cloud-based library, or select from a voice library provided by the application. Then, as needed, the user-specified voice can be voice-changed or mixed with multiple tones to create a richer sound effect.
[0088] The virtual avatar robot created in this way can be created or edited based on user-specified attributes. This improves the match between the virtual avatar robot and the first user's conversational intent, allowing the robot to better assist in the conversation.
[0089] After determining the robot to be called, the user can also be asked, and the robot can be called after the user confirms. In some embodiments, a first call control for the robot corresponding to the first conversation is displayed; in response to the first user triggering the first call control, auxiliary conversation information is triggered. Thus, the robot can be called after the first user confirms the robot recommended based on the conversation through the first call control. After the robot is called through the call control, it can provide at least one of the following as auxiliary conversation information: reference information associated with the first conversation, a summary of the first conversation, and a reply to the second user. Thus, in response to the user's confirmation, the robot can provide various forms of responses as a conversation assistant to the user.
[0090] Figures 5A and 5B illustrate diagrams of conversation interfaces according to some embodiments of the present disclosure. As shown in Figure 5A, the conversation interface 51 of this embodiment is an interface for a conversation between a first user AA and a second user BB. In the conversation, the two users are discussing the topic of traveling to Jiuzhaigou.
[0091] In response to detecting the topic, the application can generate an automatic reply message, such as the content shown in 511, so that user AA can directly send the content after triggering the control containing the automatic reply message to quickly reply. Control 511 can also include the identifier of the robot that provided the auxiliary conversation information. Control 511 is not visible to user BB.
[0092] In the conversation interface 51, a first call control 512 may also be displayed. Based on the conversation between the two users, it can be determined that the corresponding robot is "Travel Assistant." The first call control 512 may carry the identification of the travel assistant, such as its name or profile picture, to confirm whether to call the robot. The first call control 512 is also invisible to the user BB.
[0093] In some embodiments, the call control includes instruction information for the robot generated based on the semantic information of the first conversation. In response to the first user triggering the first call control, auxiliary conversation information is generated based on the instruction information and the first conversation. For example, in the conversation interface 51, the first call control 512 includes not only the name "Travel Assistant" but also the instruction information "Provide me a Jiuzhaigou itinerary plan." Therefore, when the user triggers the first call control 512, the instruction information can be sent to the travel assistant robot, so that the robot can generate a response based on the previous conversation and the instruction information.
[0094] As described in the previous embodiments, the robot can participate in the conversation or only provide visible reference information to the first user. In some embodiments, the dialogue interface between the first user and the second user displays auxiliary dialogue information sent by the robot participating in the dialogue between the first user and the second user. For example, the travel assistant can send a message as a participant in the dialogue, as shown in the message control 513 in Figure 5A. For another example, the travel assistant can also display the generated information on the first user's client. In some embodiments, the generated content can be directly filled into the first user's message input control, such as the input box 514 shown in Figure 5B, and then the first user can directly send the message or send the message after editing.
[0095] [Corrected 20.05.2024 in accordance with Rule 91] Following the first conversation, the first user can continue to send messages based on the auxiliary conversation information provided by the robot. In some embodiments, a second conversation sent by the first user is displayed, where the second conversation is determined based on the auxiliary conversation information, and the auxiliary conversation information is not visible to the second user. For example, following Figure 5B, the user can edit the content automatically filled into input box 514 by the robot, for example, changing "I put together a guide, day one... day two..." to "How about this arrangement, day one... day two..." and sending it. As a result, the message sent by user AA appears in the conversation interface 51, as shown in the message control 515 in Figure 5C.
[0096] The above embodiments provide methods for automatically recommending a robot to a user or directly and automatically calling a robot. In some embodiments, the user can also manually control whether to use a robot to assist in the conversation.
[0097] Figure 6 shows a diagram of a conversation interface according to some embodiments of the present disclosure. As shown in Figure 6, the conversation interface 60 of this embodiment is a conversation interface between a first user AA and a second user BB, wherein user BB sends a message 601 "What time will you be back in the evening?"
[0098] Because user AA hasn't replied for a long time, in some embodiments, user AA's virtual avatar robot aa can be automatically called based on the current conversation status. Robot aa then sends the message "He must be working overtime. He's been very busy lately" 602 to the conversation. In some embodiments, in response to users BB and AA being in a relationship, the conversation guidance robot "Atmosphere Group" can be automatically called. The "Atmosphere Group" robot then sends the message "Are you with your female colleague?" 603 to tease and liven up the conversation.
[0099] Some or all of the robots shown in Figure 6 can also be manually called by the user. In some embodiments, in response to the first user's operation of calling the robot, a robot determined based on the information of the first conversation, or a robot specified by the operation, is displayed. The robot can be displayed in the conversation (i.e., the robot is made to participate in the conversation) or the robot can be displayed for the first user (i.e., the robot is instructed to provide reference information). For example, the call 604 of robot aa and the call control 605 of the atmosphere group can be set in the interface 60. The first user can turn the robot on or off by triggering controls 604 and 605. The turned-on robot will participate in the conversation, and the turned-off robot will not participate in the conversation. In addition to the method of automatically calling the robot, the user can also manually call the robot as needed, which increases the user's freedom to use the robot. When the robot is turned on, the robot can decide whether to send a message and what content to send based on the content output by each subject in the current conversation.
[0100] In some embodiments, a robot can output a round of messages each time it responds to a call from the first user. For example, a second call control corresponding to a candidate robot can be displayed; in response to triggering the second call control, the candidate robot's reply to the first conversation can be displayed. For example, the conversation interface can include second call controls for one or more candidate robots. For a particular candidate robot, each time the first user clicks its second call control, the robot participates in the conversation and outputs a message. Afterward, the robot stops sending messages and instructs the user to trigger its second call control again. This allows the robot to participate in the conversation in response to user needs and provides stricter control over the robot's level of participation.
[0101] The above describes the dialogue method of the embodiment of the present disclosure. The following describes the related devices of the embodiment of the present disclosure in conjunction with the accompanying drawings.
[0102] Figure 7 shows a schematic diagram of the structure of a conversation device according to some embodiments of the present disclosure. As shown in Figure 7, the conversation device 70 of this embodiment includes: a first display module 701, configured to display a first conversation between the first user and a second user on the first user's client; a calling module 702, configured to call one or more robots for the first user based on information from the first conversation; a generating module 703, configured to generate auxiliary conversation information for the robots based on the first conversation; and a display module 704, configured to display the auxiliary conversation information.
[0103] In some embodiments, the auxiliary dialogue information is information sent by the robot during the robot's participation in the dialogue between the first user and the second user; or, the auxiliary dialogue information is information provided to the first user for reference.
[0104] In some embodiments, the calling module 702 is further configured to: determine the conversation scenario between the first user and the second user based on the information of the first conversation; in response to the conversation scenario belonging to the first type, call one or more robots for the first user to participate in the conversation between the first user and the second user; in response to the conversation scenario belonging to the second type, call one or more robots for the first user to provide reference information.
[0105] In some embodiments, the information of the first conversation includes semantic information, and the calling module 702 is further configured to: determine one or more robots corresponding to the first conversation based on the semantic information of the first conversation; and call one or more robots corresponding to the first conversation for the first user.
[0106] In some embodiments, the dialogue device 70 further includes a determination module 705 .
[0107] In some embodiments, the determination module 705 is configured to: take the last specified number of conversations generated in the first conversation as the target conversation, or take the last conversation generated in the first conversation within a specified time period as the target conversation, or take the conversations belonging to the same topic in the first conversation as the target conversation; based on the target conversation, determine the semantic information of the first conversation.
[0108] In some embodiments, the determination module 705 is configured to: determine the topics involved in one or more rounds of conversation in the first conversation; and in response to the number of topics involved in one or more rounds of conversation being not greater than a specified value, determine the summary information of the first conversation as the semantic information of the first conversation.
[0109] In some embodiments, the determination module 705 is configured to extract key information from the first conversation as semantic information of the first conversation.
[0110] In some embodiments, the calling module 702 is further configured to: in response to the number of topics involved in the semantic information being greater than a specified value, or the semantic information involved in the first conversation being unclear, determine the conversation guiding robot as the robot corresponding to the first conversation; in response to the number of topics involved in the semantic information being not greater than a specified value, match the semantic information of the first conversation with the attributes of one or more candidate robots, and determine one or more robots corresponding to the first conversation based on the matching results.
[0111] In some embodiments, the information of the first conversation includes the category of the second user, and the calling module 702 is further configured to: in response to the category of the second user being the specified category, call a conversation-guiding robot for the first user to participate in the conversation between the first user and the second user.
[0112] In some embodiments, the auxiliary dialogue information is a message sent by the dialogue guidance robot during the process of participating in the dialogue, and the generation module 703 is further configured to: generate association information of the semantic information based on the semantic information of the first dialogue and the second dialogue, and the information of the first user and the second user; based on the association information, generate a message sent by the dialogue guidance robot.
[0113] In some embodiments, the information of the first conversation includes the conversation time, and the calling module 702 is further configured to: in response to the last conversation in the first conversation being sent by the second user and the sending time of the last conversation being greater than a specified threshold from the current time, call the robot that is the virtual avatar of the first user to reply to the second user for the first user.
[0114] In some embodiments, the avatar of the virtual avatar robot is generated based on a picture specified by the first user, and the voice of the virtual avatar robot is generated based on a voice specified by the first user.
[0115] In some embodiments, the calling module 702 is further configured to: display a first calling control of the robot corresponding to the first conversation; and trigger generation of auxiliary conversation information in response to a triggering operation of the first user on the first calling control.
[0116] In some embodiments, the calling control includes instruction information for the robot generated based on the semantic information of the first conversation, and the calling module 702 is further configured to: in response to the first user's triggering operation on the first calling control, trigger the generation of auxiliary conversation information based on the instruction information and the first conversation.
[0117] In some embodiments, the auxiliary conversation information includes at least one of: reference information associated with the first conversation, a summary of the first conversation, and a reply to the second user.
[0118] In some embodiments, the display module 704 is further configured to: display a second conversation sent by the first user, wherein the second conversation is determined based on the auxiliary conversation information, and the auxiliary conversation information is not visible to the second user.
[0119] In some embodiments, the display module 704 is further configured to: display auxiliary dialogue information sent by the robot participating in the dialogue between the first user and the second user on the dialogue interface between the first user and the second user.
[0120] In some embodiments, the generation module 703 is further configured to: in response to the first conversation corresponding to the service intention, generate auxiliary conversation information including at least one of a product card and a service subscription card.
[0121] In some embodiments, the display module 704 is further configured to: in response to the first user's operation of calling the robot, display the robot determined based on the information of the first conversation, or operate the designated robot.
[0122] In some embodiments, the display module 704 is further configured to: display a second call control corresponding to the candidate robot; and display a reply of the candidate robot to the first conversation in response to a triggering operation on the second call control.
[0123] It should be noted that the above-mentioned units are merely logical modules divided according to the specific functions they implement, and are not intended to limit specific implementation methods. For example, they can be implemented in software, hardware, or a combination of software and hardware. In actual implementation, the above-mentioned units can be implemented as independent physical entities, or can also be implemented by a single entity (for example, a processor (CPU or DSP, etc.), an integrated circuit, etc.). In addition, the above-mentioned units are shown with dotted lines in the accompanying drawings to indicate that these units may not actually exist, and the operations / functions they implement can be implemented by the processing circuit itself.
[0124] In addition, although not shown, the device may also include a memory that can store various information generated by the device and the various units contained in the device during operation, programs and data used for operation, data to be sent by the communication unit, etc. The memory can be volatile memory and / or non-volatile memory. For example, the memory can include but is not limited to random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), read-only memory (ROM), and flash memory. Of course, the memory can also be located outside the device. Optionally, although not shown, the device may also include a communication unit that can be used to communicate with other devices. In one example, the communication unit can be implemented in an appropriate manner known in the art, for example, including communication components such as an antenna array and / or a radio frequency link, various types of interfaces, communication units, etc. This will not be described in detail here. In addition, the device may also include other components not shown, such as a radio frequency link, a baseband processing unit, a network interface, a processor, a controller, etc. This will not be described in detail here.
[0125] Some embodiments of the present disclosure also provide an electronic device. Figure 8 shows a schematic structural diagram of an electronic device according to some embodiments of the present disclosure. For example, in some embodiments, the electronic device 8 can be various types of devices, for example, including but not limited to mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. For example, the electronic device 8 may include a display panel for displaying data and / or execution results utilized in the scheme of the present disclosure. For example, the display panel can be of various shapes, such as a rectangular panel, an elliptical panel, or a polygonal panel. In addition, the display panel can be not only a flat panel, but also a curved panel or even a spherical panel.
[0126] As shown in FIG8 , the electronic device 8 of this embodiment includes a memory 81 and a processor 82 coupled to the memory 81. It should be noted that the components of the electronic device 8 shown in FIG8 are merely exemplary and non-limiting. The electronic device 8 may also include other components as required by actual applications. The processor 82 may control the other components in the electronic device 8 to perform desired functions.
[0127] In some embodiments, the memory 81 is configured to store one or more computer-readable instructions. When the processor 82 is configured to execute the computer-readable instructions, the computer-readable instructions, when executed by the processor 82, implement the method according to any of the above-described embodiments. The specific implementation and related explanations of each step of the method can be found in the above-described embodiments, and any repetitive details are omitted here.
[0128] For example, the processor 82 and the memory 81 may communicate with each other directly or indirectly. For example, the processor 82 and the memory 81 may communicate with each other via a network. The network may include a wireless network, a wired network, and / or any combination of wireless networks and wired networks. The processor 82 and the memory 81 may also communicate with each other via a system bus, which is not limited in this disclosure.
[0129] For example, the processor 82 can be embodied as various appropriate processors, processing devices, etc., such as a central processing unit (CPU), a graphics processing unit (GPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The central processing unit (CPU) can be an X86 or ARM architecture, etc. For example, the memory 81 can include any combination of various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The memory 81 can include, for example, a system memory, which stores, for example, an operating system, an application, a boot loader (Boot Loader), a database, and other programs. Various applications and various data can also be stored in the storage medium.
[0130] In addition, according to some embodiments of the present disclosure, when various operations / processes according to the present disclosure are implemented through software and / or firmware, the programs constituting the software can be installed from a storage medium or a network to a computer system having a dedicated hardware structure, such as the computer system 90 shown in Figure 9. When the various programs are installed, the computer system can perform various functions, including the functions described above. Figure 9 shows a schematic structural diagram of a computer system according to some embodiments of the present disclosure.
[0131] In Figure 9, a central processing unit (CPU) 901 performs various processes according to a program stored in a read-only memory (ROM) 902 or a program loaded from a storage part 908 to a random access memory (RAM) 903. In the RAM 903, data required when the CPU 901 performs various processes, etc., is also stored as needed. The central processing unit is merely exemplary and may also be other types of processors, such as the various processors described above. The ROM 902, RAM 903, and storage part 908 may be various forms of computer-readable storage media, as described below. It should be noted that although ROM 902, RAM 903, and storage device 908 are shown separately in Figure 9, one or more of them may be combined or located in the same or different memory or storage modules.
[0132] The CPU 901, the ROM 902, and the RAM 903 are connected to one another via a bus 904. An input / output interface 905 is also connected to the bus 904.
[0133] The following components are connected to the input / output interface 905: an input portion 909, such as a touch screen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; an output portion 907, including a display, such as a cathode ray tube (CRT), liquid crystal display (LCD), speaker, vibrator, etc.; a storage portion 908, including a hard disk, magnetic tape, etc.; and a communication portion 909, including a network interface card, such as a LAN card, modem, etc. The communication portion 909 allows communication processing to be performed via a network, such as the Internet. It will be readily understood that although FIG9 shows that the various devices or modules in the computer system 90 communicate via the bus 904, they may also communicate via a network or other means, where the network may include a wireless network, a wired network, and / or any combination of wireless and wired networks.
[0134] A drive 910 is also connected to the input / output interface 905 as needed. A removable medium 911 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory or the like is mounted on the drive 910 as needed so that a computer program read therefrom is installed in the storage portion 908 as needed.
[0135] In the case of realizing the above-described series of processing by software, a program constituting the software can be installed from a network such as the Internet or a storage medium such as the removable medium 911 .
[0136] According to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network via the communication device 909, or installed from the storage device 908, or installed from the ROM 902. When the computer program is executed by the CPU 901, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.
[0137] It should be noted that in the context of the present disclosure, a computer-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium or any combination thereof. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that may be used by or in conjunction with an instruction execution system, apparatus, or device. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries a computer-readable program code. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.
[0138] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0139] In some embodiments, a computer program is further provided, comprising: instructions, which, when executed by a processor, cause the processor to perform any of the methods of the above embodiments. For example, the instructions may be embodied as computer program codes.
[0140] In embodiments of the present disclosure, computer program code for performing the operations of the present disclosure may be written in one or more programming languages or combinations thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0141] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0142] The modules, components, or units described in the embodiments of the present disclosure may be implemented in software or hardware. The names of the modules, components, or units do not necessarily limit the modules, components, or units themselves.
[0143] The functions described above herein may be performed at least in part by one or more hardware logic components. For example, and without limitation, exemplary hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.
[0144] The above descriptions are merely some embodiments of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by a specific combination of the above-mentioned technical features, but also encompasses other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in the present disclosure.
[0145] In the description provided herein, numerous specific details are set forth. However, it is understood that embodiments of the present invention may be practiced without these specific details. In other cases, well-known methods, structures, and techniques are not presented in detail in order not to obscure the understanding of the description.
[0146] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details have been included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.
[0147] Although some specific embodiments of the present disclosure have been described in detail by way of examples, those skilled in the art will appreciate that the above examples are for illustrative purposes only and are not intended to limit the scope of the present disclosure. Those skilled in the art will appreciate that modifications may be made to the above embodiments without departing from the scope and spirit of the present disclosure. The scope of the present disclosure is defined by the appended claims.
Claims
1. A conversation method, comprising: On a client of a first user, displaying a first conversation between the first user and a second user; Invoking one or more robots for the first user based on information from the first conversation; generating auxiliary dialogue information of the robot based on the first dialogue; The auxiliary dialogue information is displayed.
2. The dialogue method according to claim 1, wherein: The auxiliary dialogue information is information sent by the robot during the process of the robot participating in the dialogue between the first user and the second user; or The auxiliary dialogue information is information provided to the first user for reference.
3. The dialogue method according to claim 1 or 2, wherein: The calling one or more robots for the first user based on the information of the first conversation includes: Determining a conversation scenario between the first user and the second user based on the information of the first conversation; In response to the conversation scenario belonging to the first type, invoking one or more robots for the first user to participate in the conversation between the first user and a second user; In response to the dialogue scenario belonging to the second type, one or more robots for providing reference information are called for the first user.
4. The dialogue method according to any one of claims 1 to 3, wherein: The information of the first conversation includes semantic information, and calling one or more robots for the first user based on the information of the first conversation includes: Determining one or more robots corresponding to the first conversation based on the semantic information of the first conversation; One or more robots corresponding to the first conversation are called for the first user.
5. The dialogue method according to claim 4, further comprising: The last conversations generated in the first conversation and a specified number of conversations are used as target conversations, or the last conversations generated in the first conversation and generated within a specified time period are used as target conversations, or the conversations in the first conversation that have the same topic are used as target conversations; Based on the target dialog, semantic information of the first dialog is determined.
6. The dialogue method according to claim 4 or 5, further comprising: determining a topic involved in one or more conversation rounds in the first conversation; In response to the number of topics involved in the one or more rounds of dialogue being no greater than a specified value, the summary information of the first dialogue is determined as the semantic information of the first dialogue.
7. The dialogue method according to any one of claims 4 to 6, wherein: Also includes: Key information is extracted from the first conversation as semantic information of the first conversation.
8. The dialogue method according to any one of claims 4 to 7, wherein: The determining, based on the semantic information of the first conversation, one or more robots corresponding to the first conversation includes: In response to the number of topics involved in the semantic information being greater than a specified value, or the semantic information involved in the first conversation being ambiguous, determining the conversation guiding robot as the robot corresponding to the first conversation; In response to the number of topics involved in the semantic information being no greater than the specified value, the semantic information of the first conversation is matched with attributes of one or more candidate robots, and one or more robots corresponding to the first conversation are determined based on the matching results.
9. The dialogue method according to any one of claims 1 to 8, wherein: The information of the first conversation includes a category of the second user, and calling one or more robots for the first user based on the information of the first conversation includes: In response to the second user's category being the specified category, a conversation guiding robot is called for the first user to participate in a conversation between the first user and the second user.
10. The dialogue method according to claim 8 or 9, wherein: The auxiliary dialogue information is a message sent by the dialogue guiding robot during the dialogue process, and the generating of the auxiliary dialogue information of the robot includes: generating association information of the semantic information based on the semantic information of the first conversation and the second conversation, and information of the first user and the second user; Based on the associated information, a message sent by the dialogue guiding robot is generated.
11. The dialogue method according to any one of claims 1 to 10, wherein: The information of the first conversation includes a conversation time, and calling one or more robots for the first user based on the information of the first conversation includes: In response to the last conversation in the first conversation being sent by the second user and the sending time of the last conversation exceeding a specified threshold from the current time, calling a robot that is the virtual avatar of the first user to reply to the second user.
12. The dialogue method according to claim 11, wherein: The avatar of the virtual avatar robot is generated according to the picture specified by the first user, and the voice of the virtual avatar robot is generated according to the voice specified by the first user.
13. The dialogue method according to any one of claims 1 to 12, wherein: The calling one or more robots for the first user includes: Displaying a first calling control of the robot corresponding to the first conversation; In response to the first user's triggering operation on the first calling control, the auxiliary dialogue information is triggered to be generated.
14. The dialogue method according to claim 13, wherein: The calling control includes instruction information for the robot generated according to the semantic information of the first dialogue, and the triggering of generating the auxiliary dialogue information includes: In response to the triggering operation of the first user on the first calling control, the auxiliary dialogue information is triggered to be generated according to the indication information and the first dialogue.
15. The dialogue method according to claim 13 or 14, wherein: The auxiliary conversation information includes at least one of: reference information associated with the first conversation, a summary of the first conversation, and a reply to the second user.
16. The dialogue method according to any one of claims 1 to 15, further comprising: A second conversation sent by the first user is displayed, wherein the second conversation is determined based on the auxiliary conversation information, and the auxiliary conversation information is not visible to the second user.
17. The dialogue method according to any one of claims 1 to 16, wherein: The displaying of the auxiliary dialogue information includes: On the conversation interface between the first user and the second user, the auxiliary conversation information sent by the robot participating in the conversation between the first user and the second user is displayed.
18. The dialogue method according to any one of claims 1 to 17, wherein: The generating of the robot's auxiliary dialogue information based on the first dialogue includes: In response to the first conversation corresponding to the service intention, the auxiliary conversation information including at least one of a product card and a service subscription card is generated.
19. The dialogue method according to any one of claims 1 to 18, further comprising: In response to the first user's operation of calling a robot, a robot determined based on the information of the first conversation or a robot specified by the operation is displayed.
20. The dialogue method according to any one of claims 1 to 19, further comprising: displaying a second call control corresponding to the candidate robot; In response to a triggering operation on the second call control, a reply of the candidate robot to the first dialogue is displayed.
21. A conversation device, comprising: A first display module is configured to display a first conversation between the first user and a second user on a client of the first user; a calling module, configured to call one or more robots for the first user based on information of the first conversation; a generating module configured to generate auxiliary dialogue information of the one or more robots based on the first dialogue; The second display module is configured to display the auxiliary dialogue information.
22. An electronic device comprising: Memory; as well as A processor coupled to the memory, wherein the processor is configured to execute the dialog method according to any one of claims 1 to 20 based on instructions stored in the memory.
23. A computer-readable storage medium having a computer program stored thereon, wherein when the program is executed by a processor, the dialogue method according to any one of claims 1 to 20 is implemented.
24. A computer program product, which, when running on a computer, enables the computer to implement the dialogue method according to any one of claims 1 to 20.
25. A computer program comprising: Instructions which, when executed by a processor, cause the processor to perform the dialog method according to any one of claims 1 to 20.
Citation Information
Patent Citations
Information processing method and device in user dialogue, electronic equipment and storage medium
CN112000781A
Message processing method, device and readable storage medium
CN113569037A
Intelligent dialogue method and device, electronic equipment and storage medium
CN115934901A
Contextual Help Recommendations for Conversational Interfaces Based on Interaction Patterns
US20210382925A1