Interaction method, electronic equipment, computer readable storage medium and program product
By displaying the trigger controls for additional objects in the dialogue interface between the agent and the user, the agent recommends suitable additional objects according to the user's intention, solving the problem that the agent cannot effectively explore the user's potential needs, and improving interaction efficiency and information acquisition methods.
Patent Information
- Application Number
- CN202510104024.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-09
AI Technical Summary
During the dialogue between the agent and the user, the agent mainly responds to the user through text, and cannot effectively explore the user's potential needs, resulting in a cumbersome interaction process and it is difficult for users to obtain more usage information of objects of type, which affects usage efficiency.
By displaying the trigger controls for additional objects in the dialogue interface, the agent recommends suitable additional objects based on the user's intention and reply to messages, such as the agent's skills, multimedia content, product or service cards, etc., to improve the user's access efficiency of more functions and information.
It improves the interaction efficiency between users and agents, helps users obtain information in more forms, enhances the utilization rate of functions and services provided by agents, and saves computing and network resources.
Smart Images

Figure CN119960634A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and in particular to an interaction method, an electronic device, a computer-readable storage medium, and a program product. Background Art
[0002] With the development of artificial intelligence technology, users can have conversations with intelligent agents. For example, users can send messages to intelligent agents in the form of text or voice, and the intelligent agents send reply texts based on the messages sent by users. The reply texts are used to respond to the messages sent by users. For example, when a user asks a question, the intelligent agent can answer the question through reply texts. Summary of the invention
[0003] According to some embodiments of the present disclosure, an interaction method is provided, comprising: displaying a conversation in a conversation interface between a user and an agent, wherein the conversation includes a first message sent by the user, and a second message sent by the agent, the second message being a response to the first message; determining the user's intention based on the first message; determining additional objects recommended by the agent to the user based on the user's intention and the second message; and displaying a trigger control for the additional object in the conversation interface.
[0004] According to some embodiments of the present disclosure, an electronic device is provided, comprising: a memory; and a processor coupled to the memory, wherein the processor is configured to execute the interaction method of any embodiment described in the present disclosure based on instructions stored in the memory.
[0005] According to some embodiments of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the program is executed by a processor, the processor executes the interactive method of any embodiment described in the present disclosure.
[0006] According to some embodiments of the present disclosure, a computer program product is provided. When the computer program product is executed on a computer, the computer is enabled to implement the interaction method of any embodiment described in the present disclosure.
[0007] Other features, aspects and advantages of the present disclosure will become apparent from the following detailed description of exemplary embodiments of the present disclosure with reference to the attached drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] The following is an explanation of the embodiments of the present disclosure with reference to the accompanying drawings. It should be understood that the drawings described below only relate to some embodiments of the present disclosure and do not constitute a limitation to the present disclosure. In the accompanying drawings:
[0009] Figure 1 A flowchart of an interaction method according to some embodiments of the present disclosure is shown.
[0010] Figure 2 A flowchart of a method for determining an additional object according to some embodiments of the present disclosure is shown.
[0011] Figure 3 A flowchart of a method for determining additional objects based on a user's creative intent according to some embodiments of the present disclosure is shown.
[0012] Figure 4 A flowchart of a method for determining additional objects based on a user's query intent according to some embodiments of the present disclosure is shown.
[0013] Figure 5 A flowchart of a method for determining additional objects based on a user's consumption decision intention according to some embodiments of the present disclosure is shown.
[0014] Figure 6 A flowchart of a method for determining additional objects based on ambiguous intent according to some embodiments of the present disclosure is shown.
[0015] Figure 7 A flowchart of a method for determining additional objects based on ambiguous intent according to other embodiments of the present disclosure is shown.
[0016] Figure 8 A schematic diagram of a dialogue interface according to some embodiments of the present disclosure is shown.
[0017] Fig. 9 A schematic structural diagram of an interaction device according to some embodiments of the present disclosure is shown.
[0018] Fig.10 A block diagram of an electronic device according to some embodiments of the present disclosure is shown.
[0019] Fig.11 A block diagram of an electronic device according to some other embodiments of the present disclosure is shown.
[0020] It should be understood that, for ease of description, the sizes of the various parts shown in the drawings are not necessarily drawn according to the actual proportional relationship. The same or similar reference numerals are used in the various drawings to represent the same or similar parts. Therefore, once an item is defined in one drawing, it may not be further discussed in subsequent drawings. DETAILED DESCRIPTION
[0021] The technical solutions in the embodiments of the present disclosure will be described clearly and completely below in conjunction with the accompanying drawings in the embodiments of the present disclosure. It should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments described herein.
[0022] It should be understood that the various steps described in the method embodiments of the present disclosure can be performed in different orders, and / or performed in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect. Unless otherwise specifically stated, the relative arrangement and numerical values of the steps set forth in these embodiments should be interpreted as being merely exemplary and not limiting the scope of the present disclosure.
[0023] The term “including” and its variations used in the present disclosure are intended to be open terms that include at least the following elements / features but do not exclude other elements / features, that is, “including but not limited to.” The term “based on” means “at least partly based on.”
[0024] It should be noted that the concepts of "first", "second", etc. mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units. Unless otherwise specified, the concepts of "first", "second", etc. are not intended to imply that the objects described in this way must be in a given order in time, space, ranking, or any other manner.
[0025] It should be noted that the modifications of "one" and "plurality" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, it should be understood as "one or more".
[0026] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.
[0027] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0028] The embodiments of the present disclosure are described in detail below in conjunction with the accompanying drawings, but the present disclosure is not limited to these specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. In addition, in one or more embodiments, specific features, structures or characteristics can be combined in any suitable manner that will be clear from the present disclosure by a person of ordinary skill in the art.
[0029] After analysis, it was found that in the current dialogue process between the agent and the user, the agent mainly replies to the user through text. However, the text can only reply to the message sent by the user, and cannot further explore the user's potential needs. Therefore, the user needs to express the actual needs through multiple rounds of messages or more specific descriptions, making the interaction process more cumbersome. In addition, the agent can provide users with more types of objects in addition to text, but in the above dialogue process, it is difficult for users to obtain usage information about more types of objects. Users who are not familiar with the application may not think of using more functions to assist the dialogue process, which also leads to low efficiency in the use of these objects.
[0030] The present disclosure provides an interactive method, which can recommend additional objects to users through an intelligent agent during a conversation between the user and the intelligent agent, so that the user can use various functions in the intelligent agent or application more efficiently, or obtain various information.
[0031] Figure 1 FIG. 1 is a flow chart showing an interactive method according to some embodiments of the present disclosure. Figure 1 As shown, the interactive method shown includes steps S11 to S14.
[0032] In step S11, a dialogue is displayed on a dialogue interface between a user and an agent, wherein the dialogue includes a first message sent by the user and a second message sent by the agent, wherein the second message is a response to the first message.
[0033] The dialogue interface between the user and the agent may include one or more messages sent by the user and one or more messages sent by the agent. The first message sent by the user may be voice or text, and the second message replied by the agent may also be voice or text. The dialogue between the user and the agent may include one or more rounds. That is, the dialogue interface may display one or more rounds of dialogue. Each round of dialogue includes one or more messages sent continuously by the user and one or more messages replied continuously by the agent.
[0034] In addition to messages, the dialogue interface can also include input controls for users to input text, voice, images, etc. In addition, the dialogue interface can also include an operation bar so that users can call various skills provided by the intelligent agent. The position of the operation bar is usually fixed near the input control (for example, above the input control).
[0035] In step S12, the user's intention is determined according to the first message.
[0036] The user's intention can be determined by semantically understanding the first message. For example, the first message can be processed using a natural language model to obtain a semantic understanding result of the first message. The user's intention is then extracted from the semantic understanding result. Alternatively, the first message can be directly processed using an intention analysis model to determine the user's intention. These models can be pre-trained using messages marked with user intentions to make the prediction results more accurate.
[0037] The user's intention refers to the core content or core instructions included in the message sent by the user. For example, if the first message is "Help me find an animated film, the genre of this film is action, and it tells a story of...", the user's intention can be "search" or "search for movies".
[0038] In step S13, based on the user's intention and the second message, the agent determines the additional objects to recommend to the user.
[0039] After the agent replies to the user, it can also display one or more prompt texts after the message. After these prompt texts are triggered by the user, a message sent by the user can be automatically displayed in the interface, and the message includes the triggered prompt information, but the content of this message is not actually input by the user, but sent automatically.
[0040] The additional object is different from the above-mentioned prompt text. That is, the additional object is not a simple text, voice or image, but an object presented based on a computer program. The additional object can be displayed in a specific interface, which can be independent of the dialogue interface or part of the dialogue interface. It can also be said that the additional object includes one or more controls, which are used to interact with the user.
[0041] Additional objects rely on specific skills and functions of the agent or programs that can be embedded in the dialogue interface. For example, additional objects can be skills provided by the agent, small programs that can be supported in the dialogue interface, multimedia content streams, product or service cards, specific functions of the agent, etc.
[0042] The additional object is determined not only based on the user's intention reflected in the first message, but also based on the second message. That is, the additional object is not a simple response to the first message, but a comprehensive recommendation to the user based on the first message and the second message. In this way, the second message generated and sent by the agent also provides a basis for recommending the additional object.
[0043] The second message can reply to the first message. That is, the second message can reply to the user's intention in the first message. However, the second message may only be able to respond to the first message itself sent by the user, but the user is likely to have further needs. Therefore, by further determining the additional objects recommended to the user by using the user's intention and the second message, it is possible to respond to the user in other forms other than the message based on the second message, in response to the user's possible extended needs.
[0044] In step S14, a trigger control for the additional object is displayed in the dialogue interface.
[0045] The trigger control and the message sent by the agent can be displayed in association. For example, the trigger control of the attached object is displayed after the message sent by the agent, so that the user can more clearly understand that the recommendation is made based on the dialogue between the user and the agent.
[0046] The trigger control of the additional object can be regarded as a quick entry of the additional object. In response to the user triggering the trigger control of the additional object, the interface of the additional object is displayed. The display interface of the additional object can be located in the dialogue interface or displayed as an independent interface.
[0047] In the above method, after the agent responds to the message sent by the user, by further recommending additional objects to the user, the potential needs of the user can be further responded to. And the user can be responded to in a way other than dialogue messages, thereby improving the interaction efficiency between the user and the agent. It can also help the user to obtain information in more forms. The improvement of the user's interaction efficiency can indirectly save computing resources and network resources. In addition, the user's use of more types of additional objects can also maximize the utilization of resources configured for these objects.
[0048] In response to the trigger control of the additional object being triggered, the interface of the additional object is displayed. For example, if the additional object is a media stream, the interface of the additional object is a playback interface of the media stream, and the user can switch different multimedia contents through sliding operations in the interface; if the additional object is a skill of an intelligent agent, the interface of the additional object is an input interface of the intelligent agent, including one or more input items and execution controls.
[0049] In the case where the attached object is a skill of an intelligent agent, after the user triggers its display, the user can manually input information and manually trigger the execution. In addition, the skill can also be executed automatically. For example, in response to the user triggering the trigger control, information matching the skill is extracted from the conversation; based on the information matching the skill, the skill is executed. Taking the skill of rewriting creative content as an example, in response to the trigger control of the skill being triggered by the user, the rewriting result can be automatically generated based on the first message and the second message. Thereby, the user's operating efficiency can be improved.
[0050] In the case where the additional object is at least one of a player, an item card, and a mini-program, displaying the trigger control of the additional object may include: determining the preview information of the additional object according to the second message; and displaying the trigger control including the preview information. That is, the second message will affect the display of the trigger control. For a player, the trigger control of the player may display the attribute information such as the cover, name, and singer of the song to be played; for an item card, the trigger control of the item card may display the name, price, picture, etc. of the item.
[0051] It can be seen that the additional objects recommended to the user in the present disclosure may have personalized information determined based on the conversation.
[0052] Since the agent can support multiple types of objects, in order to make more accurate recommendations to users, it is possible to first determine which type of object to recommend to the user. Then, a specific object is determined from the objects of that type as an additional object. Figure 2 An embodiment of the additional object determination method of the present disclosure is described.
[0053] Figure 2 FIG. 2 is a flow chart showing a method for determining an additional object according to some embodiments of the present disclosure. Figure 2 As shown, the method for determining the additional object in this embodiment includes steps S131 to S132.
[0054] In step S131 , according to the user's intention, the object type of the additional object to be recommended to the user is determined.
[0055] For example, when identifying the user's intention, the recognition result can be divided into one of several predetermined intentions, and then the object type that matches the user's intention is determined based on the correspondence between the intention and the object type. For example, several predetermined intentions can include generation, query, decision, etc. In addition, ambiguous intentions can also be used as a type. In this way, the output of the intention analysis model can be set to several fixed outputs.
[0056] For another example, when using the intent analysis model to identify the user's intent, the identification result can be expressed in any way. Then, classification is performed based on the intent, and each classification result corresponds to an object type. When classifying, a classification model based on natural language processing can be used.
[0057] In addition, the user's extended intent can be determined based on the user's intent in the first message, and then based on the extended intent, the type of object that matches the extended intent can be determined as the type of additional object recommended to the user. The extended intent refers to the intention that the user may further want to express after receiving the second message. For example, if the user instructs the agent to query a certain singer through the first message, the user's extended intent may be to play the singer's songs. For another example, if the user instructs the agent to generate a piece of content through the first message, the user's extended intent may be to modify the generated content.
[0058] The extended intent can be determined using a large language model. The extended intent refers to the intention that is further generated after the user's current intent (i.e., the intent in the first message) is satisfied. For example, the user intent in the first message is processed using a large language model to obtain the extended intent output by the model. Alternatively, the first message and the processing prompt can be input into the large language model, and the processing prompt instructs the model to determine the user's intent and extended intent based on the first message to obtain the extended intent output by the model.
[0059] In step S132, based on the second message, additional objects belonging to the object type recommended by the agent to the user are determined.
[0060] The second message provides detailed information of the additional object recommended to the user. That is, based on the object type of the additional object, it is further determined which additional object of the type is recommended. For example, the additional object recommended to the user can be determined from one or more alternative additional objects of the object type.
[0061] By matching the intent with various types of needs, and then specifically determining the object to be recommended to the user based on the object type and the second message, on the one hand, the range of objects to be selected can be narrowed when determining the recommended object to improve processing efficiency; on the other hand, the accuracy of the recommendation is improved, making the recommended object more likely to be triggered by the user. Therefore, the interaction efficiency can be improved.
[0062] The following uses creation intent, query intent, consumption decision intent, and ambiguous intent as examples to describe an exemplary method for recommending additional objects to a user after the user has a conversation with an intelligent agent.
[0063] In some embodiments, when the user's intention is to create, the agent's skills are recommended to the user as additional objects. The agent's skills can be used to further process the created content, such as editing, type conversion, etc.
[0064] Figure 3 FIG. 1 is a flow chart showing a method for determining an additional object based on a user's creative intent according to some embodiments of the present disclosure. Figure 3As shown, the determination method of this embodiment includes steps S1311 and S1321.
[0065] In step S1311, in response to the user's intention being a creative intention, the object type is determined to be a skill of the agent.
[0066] The creative intention is the intention to instruct the agent to generate creative content. The generated creative content includes text, images, music, video or multimedia content mixed in multiple modes. That is, the user instructs the agent to create through the first message, and the agent feeds back the created content through the second message. On this basis, after sending the created content, the agent can continue to recommend the agent's skills to the user so as to further process the creative content in the second message.
[0067] In step S1321, based on the second message, additional objects belonging to the agent's skills that the agent recommends to the user are determined.
[0068] For example, determining the additional object includes: determining a first content type of the creative content in the second message; and determining a skill for editing corresponding to the first content type as an additional object recommended by the agent to the user.
[0069] Content type can refer to the data type or genre type of the created content. Data types include text, images, music, etc., and the recommended skills are skills for processing the data type of the created content. Genre types include essays, scripts, news, etc., and the recommended skills are skills for processing the genre type of the created content. In addition, content types can also be divided in other ways, which will not be repeated here.
[0070] Taking text as an example, editing of creative content includes, for example, rewriting, expanding, abbreviating, continuing, etc., part or all of the creative content.
[0071] Therefore, after the intelligent agent sends the creative content through the second message, the editing skills matching the creative content can be sent to the user, so that the user can efficiently modify the content created by the intelligent agent, thereby improving the user's interaction efficiency.
[0072] In addition, the type of the creative content in the second message can be determined, and the second creative content can be converted into another type of creative content. For example, the creative content in the second message can be converted from a first content type to a second content type, and the first content type and the second content type are different. For example, determining the additional object includes: determining the first content type of the creative content in the second message; determining the skills used to generate the creative content of the second content type as the additional object recommended to the user by the intelligent agent, and the constituent elements of the creative content of the second content type match the first content type.
[0073] The constituent elements of the second content type's creative content match the first content type, which means that the first content type's content can be used as the generation material for the second content type's content. For example, the first content type is lyrics, and the second content type can be a song, that is, a song including lyrics can be generated.
[0074] The content of the second content type may not be a direct component of the content of the first content type, but an indirect component. For example, if the first content type is a story or scene description, pictures, videos or music can be generated based on the story or scene description, that is, the second content type can be pictures, videos or music.
[0075] Through this method, after the intelligent agent provides the user with the results of content creation, it can further assist the user in generating more different types of content, thereby efficiently generating richer creation results for the user.
[0076] In some embodiments, when the user's intention is a query intention, a player is recommended to the user as an additional object. The player is used to play one or more multimedia contents or play a media stream.
[0077] Figure 4 FIG. 1 is a flow chart showing a method for determining an additional object based on a user's query intent according to some embodiments of the present disclosure. Figure 4 As shown, the determination method of this embodiment includes steps S1312 and S1322.
[0078] In step S1312, in response to the user's intention being a query, the object type is determined to be a player, and the player is used to play one or more multimedia contents or a media stream.
[0079] In step S1322, based on the second message, the additional objects belonging to the player recommended by the agent to the user are determined.
[0080] For example, determining the additional object includes: in response to the second message including the operation step, determining a player for playing a video associated with the operation step as the additional content recommended to the user.
[0081] Whether the second message includes operation steps can be identified by a natural language processing model or by keyword matching, for example, whether the second message includes descriptions of step numbers or keywords such as "first...then...".
[0082] When determining the video associated with the operation step, the query information can be determined based on the second message, and the query information can be used to search for videos. The query can also be combined with the first message, that is, first search based on the first message, and then search based on the query information in the search results to narrow the scope of the query. For example, a user asks the agent how to make a dessert, but there are multiple ways to make the dessert. Based on the intention of the first message, multiple videos describing the dessert can be determined, and then a video matching the second message generated by the agent can be selected.
[0083] Another example of determining an additional object is: based on the key information in the second message, retrieving an existing work that matches the key information; in response to finding an existing work corresponding to the key information, determining a player for playing the existing work as additional content.
[0084] The key information in the second message may be a keyword, a key sentence, etc. in the second message. Information of a specified type may be determined as key information, for example, a named entity may be extracted from the second message. After obtaining the key information, an existing work may be searched in the media library.
[0085] In this way, the user can quickly have a deeper understanding of the content of the second message replied by the agent. For example, the user asks the agent for information about a singer A, and the agent introduces singer A in its answer, which includes the famous works of singer A. Then the additional object recommended by the agent to the user can be a player for playing the work, so that the user can play the work more quickly through the additional object based on the information about singer A, so as to learn more about the singer more efficiently.
[0086] In the case where the retrieved videos include multiple contents, the multiple contents can be carried by the multimedia stream to facilitate the user to quickly switch and browse. That is, in this case, the tool for playing the multimedia stream can be determined as an additional object.
[0087] In some embodiments, when the user's intention is a consumption decision intention, an object for displaying detailed information of an item is recommended to the user as an additional object. Such additional objects may include at least one of a player, an item card, and a mini-program.
[0088] Consumption decision means that the user requests the agent to help make a decision, and the object of the decision is a consumable object, such as goods, services, etc. For example, the user can ask the agent questions such as "What hairstyle is popular now" or "Which of the two models of mobile phones has better performance?" The intentions of these questions can be regarded as consumption decision intentions.
[0089] Figure 5FIG. 1 is a flow chart showing a method for determining an additional object based on a user's consumption decision intention according to some embodiments of the present disclosure. Figure 5 As shown, the determination method of this embodiment includes steps S1313 and S1323.
[0090] In step S1313, in response to the user's intention being a consumption decision, the object type is determined to be at least one of a player, an item card, and a mini-program. An item card is used to display item information. A mini-program can be used to display information of a single item, or to summarize or compare multiple items.
[0091] That is, for the consumer decision intention, additional objects that can further display the objects involved in the consumer decision can be recommended to the user. For example, a video of the object to be decided can be played through a player, so that the user can further understand it in a visual way. Or through the item card, the user can obtain more attribute information of the object to be decided, such as price, link, purchase location, etc. The mini program can be, for example, a mini program for purchasing, or a mini program for comparing performance parameters and prices.
[0092] In step S1323, based on the second message, the additional objects belonging to the object type recommended by the agent to the user are determined.
[0093] For example, after determining the object type, the player, item card or applet with the information involved in the consumption decision can be determined as the additional object recommended to the user. That is, based on the second message, a search can be performed in the database or data source corresponding to the additional object, and the additional object carrying the search result can be determined as the additional object recommended to the user.
[0094] In this way, after the intelligent agent sends a message to the user, it can further recommend additional objects such as players, item cards, and mini-programs for the user's consumption decision intention, so that the user can have a deeper understanding of the message sent by the intelligent agent, or perform further operations based on the content of the message sent by the intelligent agent, such as purchase, comparison, etc.
[0095] In some embodiments, when the user's intention is unclear, the multimodal input function is recommended to the user as additional content. For example, in response to the user's unclear intention, the object type is determined to be a multimodal input function. The user's unclear intention may mean that the intention cannot be extracted from the first message sent by the user, or that multiple intentions can be extracted from the first message and the probabilities of the multiple intentions are relatively small (for example, lower than a certain set probability threshold). At this time, the user needs to provide more information, or it is necessary to consider whether the user needs to provide information of multiple dimensions in other ways. Two recommendation methods are described below as examples.
[0096] Figure 6 A flowchart of a method for determining an additional object based on an ambiguous intention according to some embodiments of the present disclosure is shown. In this embodiment, the multimodal input function includes a voice call function, the conversation includes multiple rounds of messages, and the first message and the second message are the most recent round of messages. Figure 6 As shown, the determination method of this embodiment includes steps S1314 and S1324.
[0097] In step S1314, in response to the user's intention being unclear, the object type is determined to be a multimodal input function.
[0098] In step S1324, in response to the number of conversation turns reaching a first number within a specified time length, the function of talking to the agent is determined as an additional object recommended by the agent to the user.
[0099] If the user's intention is not clear, and the user and the agent have many conversations in a short period of time, it means that the user may find it difficult to express their needs clearly through text. At this time, by recommending the function of talking to the agent to the user, the interaction efficiency between the user and the agent can be improved, making it easier for the agent to guide the user to express his intention more clearly through voice.
[0100] Figure 7 FIG. 1 is a flow chart of a method for determining an additional object based on an ambiguous intention according to other embodiments of the present disclosure. In this embodiment, the additional object is a skill of an intelligent agent. Figure 7 As shown, the determination method of this embodiment includes steps S1315, S1325 and S1326.
[0101] In step S1315, an evaluation of the reply quality of the second message is determined.
[0102] The evaluation may be determined based on a model. For example, the second message may be input into a quality evaluation model to obtain an evaluation result output by the model. The evaluation result may be a score, a category, or any other form. The quality evaluation model may be trained using messages marked with evaluation results as training data. The quality evaluation model may further receive the first message to determine the evaluation based on the degree of match between the second message and the first message.
[0103] The evaluation result may also be determined by the user. For example, if the user expresses dissatisfaction with the reply result by "disliking" the message, it indicates that the reply quality is poor.
[0104] In step S1325, in response to the evaluation being lower than a designated level, based on the user's intention, a plurality of elements corresponding to the intention are determined.
[0105] If the rating is lower than the specified level, it means that the agent's reply does not respond well to the message sent by the user. One of the reasons is that the user did not clearly describe the demand. This may be caused by the user's poor language expression ability or because the user's intention is difficult to express through the text dimension.
[0106] The elements corresponding to the intent can be determined based on a large language model or a preset corresponding relationship. The multiple elements can be one or more of elements of multiple data types, elements of multiple semantic types, etc. For example, for the intention of asking for directions, in addition to providing a text description, it is best for the user to provide a picture of the actual environment so that the agent can more accurately guide the user; for the intention of searching for songs, in addition to describing the lyrics through text, the user can also describe the rhythm and atmosphere of the song through text, or can assist the agent in more accurately determining the song to be searched by inputting the audio of the user humming.
[0107] In step S1326, in response to the first message not covering the multiple elements, a skill satisfying the specified condition is determined as a skill recommended by the agent to the user, the input of the skill includes the multiple elements, and the output of the skill matches the intention. That is, the specified condition means that the input includes the multiple elements and the output matches the user's intention.
[0108] If the first message does not cover the multiple elements, it means that the information provided by the user through the first message is not comprehensive enough. At this time, the information covering the multiple elements can be provided by recommending skills to the user and guiding the user to input the skills as required.
[0109] In response to a skill being triggered, an input interface corresponding to the skill may be displayed. The input interface includes input controls for multiple elements and a description of each element, so that the user can efficiently input information according to the guidance of the interface.
[0110] Since the user has sent part of the information through the first message, the information can be extracted from the first message, and the extracted information does not need to be input again by the user, so as to save the user's input time and improve the operation efficiency. For example, in response to the skill being triggered, the information matching the skill is extracted from the first message; the information input by the user through the skill is received; and the skill is executed according to the information extracted from the first message and the information input by the user through the skill. In this example, the information input by the user through the skill may be an element not covered by the first message. As needed, the information extracted from the first message can be automatically filled into the input control of the skill, and the user can further modify it or not.
[0111] When recommending additional objects, the historical conversations between the user and the agent may also be referred to. For example, assuming that the conversation includes multiple rounds of messages, and the first message and the second message are the messages of the most recent round, then determining the additional objects recommended to the user includes: determining the additional objects recommended to the user based on the user's intention, the messages of each round and their weights, wherein the weight of the messages of each round is determined based on the time interval between the sending time of the messages of that round and the current time, and is negatively correlated with the time interval. That is, messages closer to the current time can have a greater impact on the determination of the additional objects, and the first message and the second message have the greatest impact. Thus, in the process of recommendation, the additional objects can be recommended to the user more accurately in combination with the user's historical conversations.
[0112] Figure 8 FIG. 2 shows a schematic diagram of a dialogue interface according to some embodiments of the present disclosure. Figure 8 As shown, in the interface 8, there are a first message 81 sent by the user and a second message 82 replied by the agent. The first message 81 is, for example, "Help me write a winter story", and the second message 82 is the creative content replied by the agent, for example: "On a snowy day, a little girl was making a snowman in the snow, when she discovered..." (For space considerations, the latter part of the creative content is omitted in this disclosure). Since the intention of the first message 81 is creation, and there are many descriptions of pictures and scenes in the second message 82, an additional object 83 for image generation can be recommended to the user. When the user triggers the additional object 83, the image creation result based on the second message 82 can be quickly obtained, or the interface of the image generation skill can be evoked, and the interface has been filled with some information extracted from the second message 82, so that the user can perform the generation process after fine-tuning.
[0113] The above describes the various methods of the present disclosure, and the following further describes the devices for executing the various methods.
[0114] Fig. 9 FIG. 2 shows a schematic diagram of the structure of an interactive device according to some embodiments of the present disclosure. Fig. 9 As shown, the interactive device 9 of this embodiment includes: a first display module 91, configured to display a dialogue in a dialogue interface between a user and an agent, wherein the dialogue includes a first message sent by the user and a second message sent by the agent, and the second message is a response to the first message; a first determination module 92, configured to determine the user's intention based on the first message; a second determination module 93, configured to determine the additional object recommended by the agent to the user based on the user's intention and the second message; and a second display module 94, configured to display a trigger control for the additional object in the dialogue interface.
[0115] In some embodiments, the second determination module 93 is further configured to: determine the object type of the additional object to be recommended to the user according to the user's intention; and determine the additional object belonging to the object type recommended to the user by the agent according to the second message.
[0116] In some embodiments, the second determination module 93 is further configured to: in response to the user's intention being a creative intention, determine the object type as a skill of the agent.
[0117] In some embodiments, the second determination module 93 is further configured to: determine a first content type of the creative content in the second message; and determine a skill for editing corresponding to the first content type as an additional object recommended by the agent to the user.
[0118] In some embodiments, the second determination module 93 is further configured to: determine the first content type of the creative content in the second message; determine the skills used to generate the creative content of the second content type as additional objects recommended by the intelligent agent to the user, and the constituent elements of the creative content of the second content type match the first content type.
[0119] In some embodiments, the second determination module 93 is further configured to: in response to the user's intention being a query, determine the object type as a player, where the player is used to play one or more multimedia contents or play a media stream.
[0120] In some embodiments, the second determination module 93 is further configured to: in response to the second message including the operation step, determine a player for playing a video associated with the operation step as the additional content.
[0121] In some embodiments, the second determination module 93 is further configured to: based on the key information in the second message, retrieve existing works that match the key information; in response to querying the existing works corresponding to the key information, determine the player used to play the existing works as additional content.
[0122] In some embodiments, the second determination module 93 is further configured to: in response to the user's intention being a consumption decision, determine the object type as at least one of a player, an item card, and a mini-program.
[0123] In some embodiments, the second determination module 93 is further configured to: in response to the user's intention being unclear, determine the object type as a multimodal input function.
[0124] In some embodiments, the multimodal input function includes a voice call function, the conversation includes multiple rounds of messages, the first message and the second message are messages of the most recent round, and the second determination module 93 is further configured to: in response to the number of rounds of the conversation reaching a first number within a specified time length, determine the function of talking to the agent as an additional object recommended by the agent to the user.
[0125] In some embodiments, the additional object is the skill of the intelligent agent, and the second determination module 93 is further configured to: determine an evaluation of the quality of the reply to the second message; in response to the evaluation being lower than a specified level, determine multiple elements corresponding to the intention based on the user's intention; in response to the first message not covering multiple elements, determine the skills that meet the specified conditions as the skills recommended by the intelligent agent to the user, the input of the skill includes multiple elements, and the output of the skill matches the intention.
[0126] In some embodiments, the interaction device 9 further includes an execution module 95 .
[0127] In some embodiments, the execution module 95 is configured to: in response to a skill being triggered, extract information matching the skill from the first message; receive information input by the user through the skill; and execute the skill based on the information extracted from the first message and the information input by the user through the skill.
[0128] In some embodiments, the additional object is a skill of the intelligent agent, and the execution module 95 is configured to: in response to the user triggering the trigger control, extract information matching the skill from the conversation; and execute the skill based on the information matching the skill.
[0129] In some embodiments, the additional object is at least one of a player, an item card, and a mini-program, and the second display module 94 is further configured to: determine preview information of the additional object according to the second message; and display a trigger control including the preview information.
[0130] In some embodiments, the conversation includes multiple rounds of messages, the first message and the second message are the messages of the most recent round, and the second determination module 93 is further configured to determine the additional objects recommended to the user based on the user's intention, the messages of each round and their weights, wherein the weight of the messages of each round is determined based on the time interval between the sending time of the messages of the round and the current time, and is negatively correlated with the time interval.
[0131] Fig.10 A block diagram of an electronic device according to some embodiments of the present disclosure is shown.
[0132] The memory 101 is used to store one or more computer-readable instructions. The memory 101 may include any combination of various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory, including but not limited to random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), read-only memory (ROM), flash memory. The memory 101 may store, for example, an operating system, an application, a boot loader (BootLoader), a database, and other programs, and may also store various applications and various data.
[0133] The processor 102 is used to run computer-readable instructions to implement the method described in any of the above embodiments. The specific implementation of each step of the method can be referred to the above embodiments, and the repeated parts are not repeated here.
[0134] The processor 102 may be configured to perform the steps of the embodiments of the present disclosure. The processor 102 may be embodied as various processing devices, such as a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The central processing unit (CPU) may be an X86 or ARM architecture, etc.
[0135] The processor 102 and the memory 101 may communicate with each other directly or indirectly. For example, the processor 102 and the memory 101 may communicate with each other via a network. The network may include a wireless network, a wired network, and / or any combination of a wireless network and a wired network. The processor 102 and the memory 101 may also communicate with each other via a system bus, which is not limited in the present disclosure.
[0136] It should be noted that Fig.10 The components of the electronic device 10 shown are only exemplary and non-limiting. The electronic device 10 may also have other components according to actual application requirements. The processor 102 may control other components in the electronic device 10 to perform desired functions.
[0137] The electronic device 10 may be implemented by software, firmware and / or hardware, and may be integrated into a device installed with relevant application programs.
[0138] Fig.11 A block diagram of an electronic device according to some other embodiments of the present disclosure is shown.
[0139] Fig.11 The electronic device 11 shown may be a computer system with a dedicated hardware structure, which can execute corresponding functions when a relevant application program is installed.
[0140] Electronic devices include, but are not limited to, mobile terminals such as smart phones, laptops, personal digital assistants (PDA), tablet personal computers (Tablet PC), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), wearable devices, etc., and fixed terminals such as digital televisions, desktop computers, etc.
[0141] like Fig.11 As shown, a central processing unit (CPU) 111 performs various processes according to a program stored in a read-only memory (ROM) 112 or a program loaded from a storage section 118 to a random access memory (RAM) 113. In the RAM 113, data required when the CPU 111 performs various processes, etc., is stored as needed. The central processing unit is merely exemplary, and it may also be other types of processors, such as the various processors described above. The ROM 112, the RAM 113, and the storage section 118 may be various forms of computer-readable storage media. It should be noted that although Fig.11 ROM 112, RAM 113 and storage section 118 are shown separately in FIG. 1 , but one or more of them may be combined or located in the same or different memory or storage modules.
[0142] The CPU 111, the ROM 112, and the RAM 113 are connected to each other via a bus 114. To the bus 114, an input / output interface 115 is also connected.
[0143] The following components are connected to the input / output interface 115: an input portion 116, such as a touch screen, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output portion 117, including a display, such as a cathode ray tube (CRT), a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage portion 118, including a hard disk, a magnetic tape, etc.; and a communication portion 119, including a network interface card such as a LAN card, a modem, etc. The communication portion 119 allows communication processing to be performed via a network such as the Internet. It is easy to understand that although Fig.11 Parts of the electronic device 11 are shown to communicate via the bus 114, but they may also communicate via a network or other means, wherein the network may include a wireless network, a wired network, and / or any combination of a wireless network and a wired network.
[0144] A drive 1110 is also connected to the input / output interface 115 as needed. A removable medium 1111 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is mounted on the drive 1110 as needed so that a computer program read therefrom is installed into the storage section 118 as needed.
[0145] When the series of processing described above is implemented by software, the program constituting the software can be installed from a network such as the Internet or a storage medium such as the removable medium 1111 .
[0146] According to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, some embodiments of the present disclosure include a computer program product, which, when the computer program product is run on a computer, enables the computer to implement the method described in any of the aforementioned embodiments. The computer program product includes a computer instruction carried on a computer-readable medium, containing a program code for executing the method shown in the flowchart. In such an embodiment, the computer instruction can be downloaded and installed from a network through the communication part 119, or installed from the storage part 118, or installed from the ROM 112. When the computer program is executed by the CPU 111, the method of the embodiment of the present disclosure is executed.
[0147] It should be noted that, in the context of the present disclosure, a computer-readable medium may be a tangible medium that may contain or store a program for use by an instruction execution system, apparatus, or device or for use in conjunction with an instruction execution system, apparatus, or device.
[0148] The computer readable medium may be a computer readable storage medium, or a computer readable signal medium, or any combination of the two.
[0149] Computer-readable storage media include, but are not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices or components, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, electrical connections with one or more wires, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination thereof. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that may be used by or in conjunction with an instruction execution system, device, or device. Computer instructions are stored on a computer-readable storage medium, and when the instructions are executed by a processor, the method described in any of the foregoing embodiments is implemented.
[0150] Computer readable signal media may include data signals propagated in baseband or as part of a carrier wave, which carry computer readable program code. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Computer readable signal media may also be any computer readable medium other than a computer readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0151] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0152] In some embodiments, a computer program is further provided, comprising: instructions, which, when executed by a processor, cause the processor to execute the method described in any of the above embodiments. For example, the instructions may be embodied as computer program codes.
[0153] In an embodiment of the present disclosure, computer program code for performing the operations of the present disclosure may be written in one or more programming languages or a combination thereof, including but not limited to object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" language or similar programming languages. The program code may be executed entirely on a user's computer, partially on a user's computer, as a separate software package, partially on a user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0154] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some implementations as replacements, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0155] The functions described above may be performed at least in part by one or more hardware logic components. For example, exemplary hardware logic components that may be used include, without limitation, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.
[0156] Although some specific embodiments of the present disclosure have been described in detail by way of example, it should be understood by those skilled in the art that the above examples are for illustration only and are not intended to limit the scope of the present disclosure. It should be understood by those skilled in the art that the above embodiments may be modified without departing from the scope and spirit of the present disclosure. The scope of the present disclosure is defined by the appended claims.
Claims
1. An interactive method, comprising: Displaying a conversation on a conversation interface between a user and an agent, wherein the conversation includes a first message sent by the user and a second message sent by the agent, the second message being a response to the first message; Determining the user's intention according to the first message; determining, based on the user's intention and the second message, an additional object recommended by the agent to the user; On the dialogue interface, a trigger control of the additional object is displayed.
2. The interactive method according to claim 1, wherein: The determining, according to the user's intention and the second message, the additional object recommended by the agent to the user comprises: determining, according to the user's intention, an object type of the additional object to be recommended to the user; Based on the second message, additional objects belonging to the object type that are recommended by the agent to the user are determined.
3. The interactive method according to claim 2, wherein: The determining, according to the user's intention, the object type of the additional object to be recommended to the user comprises: In response to the user's intention being a creative intention, the object type is determined to be a skill of an agent.
4. The interactive method according to claim 3, wherein: Determining, according to the second message, an additional object belonging to the object type recommended by the agent to the user comprises: determining a first content type of the creative content in the second message; The skill corresponding to the first content type and used for editing is determined as an additional object recommended by the agent to the user.
5. The interactive method according to claim 3, wherein: Determining, according to the second message, an additional object belonging to the object type recommended by the agent to the user comprises: determining a first content type of the creative content in the second message; A skill for generating creative content of a second content type, whose constituent elements match those of the first content type, is determined as an additional object recommended by the agent to the user.
6. The interactive method according to claim 2, wherein: The determining, according to the user's intention, the object type of the additional object to be recommended to the user comprises: In response to the user's intention being a query, the object type is determined to be a player, and the player is used to play one or more multimedia contents or a media stream.
7. The interactive method according to claim 6, wherein: Determining, according to the second message, an additional object belonging to the object type recommended by the agent to the user comprises: In response to the second message including an operation step, a player used to play a video associated with the operation step is determined as the additional content.
8. The interactive method according to claim 6, wherein: Determining, according to the second message, an additional object belonging to the object type recommended by the agent to the user comprises: Based on the key information in the second message, searching for an existing work matching the key information; In response to finding an existing work corresponding to the key information, a player used to play the existing work is determined as the additional content.
9. The interactive method according to claim 2, wherein: The determining, according to the user's intention, the object type of the additional object to be recommended to the user comprises: In response to the user's intention being a consumption decision, the object type is determined to be at least one of a player, an item card, and a mini-program.
10. The interactive method according to claim 2, wherein: The determining, according to the user's intention, the object type of the additional object to be recommended to the user comprises: In response to the user's intention being ambiguous, determining the object type as a multimodal input function.
11. The interactive method according to claim 10, wherein: The multimodal input function includes a voice call function, the conversation includes multiple rounds of messages, the first message and the second message are messages of the most recent round, and determining, based on the second message, the additional object that belongs to the object type and that is recommended by the agent to the user includes: In response to the number of turns of the conversation reaching a first number within a specified time length, a function of talking with an agent is determined as an additional object recommended by the agent to the user.
12. The interactive method according to claim 10, wherein: The additional object is a skill of the agent, and determining, according to the user's intention and the second message, the additional object recommended by the agent to the user comprises: determining a rating of a quality of a reply to the second message; In response to the evaluation being lower than a specified level, determining, based on the user's intention, a plurality of elements corresponding to the intention; In response to the first message not covering the multiple elements, a skill satisfying a specified condition is determined as a skill recommended by the agent to the user, the input of the skill includes the multiple elements, and the output of the skill matches the intention.
13. The interactive method according to claim 12, further comprising: In response to the skill being triggered, extracting information matching the skill from the first message; receiving information input by the user through the skill; The skill is executed based on the information extracted from the first message and the information input by the user through the skill.
14. The interactive method according to claim 1, wherein: The additional object is a skill of the agent, and the interaction method further comprises: In response to the user triggering the trigger control, extracting information matching the skill from the conversation; The skill is executed according to the information matching the skill.
15. The interactive method according to claim 1, wherein: The additional object is at least one of a player, an item card, and a mini-program, and the trigger control for displaying the additional object includes: Determining preview information of the additional object according to the second message; The trigger control including the preview information is displayed.
16. The interactive method according to claim 1, wherein: The conversation includes multiple rounds of messages, the first message and the second message are messages of the most recent round, and determining the additional object recommended by the agent to the user according to the user's intention and the second message includes: According to the user's intention, the messages of each round and their weights, the additional objects recommended to the user are determined, wherein the weight of the messages of each round is determined according to the time interval between the sending time of the messages of the round and the current time, and is negatively correlated with the time interval.
17. An electronic device comprising: Memory; as well as A processor coupled to the memory, wherein the processor is configured to execute the interaction method according to any one of claims 1 to 16 based on instructions stored in the memory.
18. A computer-readable storage medium having a computer program stored thereon, wherein when the program is executed by a processor, the processor implements the interactive method according to any one of claims 1 to 16.
19. A computer program product, when the computer program product is run on a computer, enables the computer to implement the interactive method according to any one of claims 1 to 16.
Citation Information
Patent Citations
Question and answer interaction method and electronic equipment
CN117235214A
Multimedia content generation method, electronic equipment, storage medium and product
CN118885628A
Content generation method and device, electronic equipment, storage medium and program product
CN118917897A
Interaction method and device, electronic equipment, storage medium and program product
CN119096224A
Cited By
Information interaction method and device, electronic equipment and computer readable medium
CN120560768A