Interaction method and device, electronic equipment, storage medium and program product

By receiving user information in the agent interactive interface and automatically determining the matching information platform, displaying corresponding multimedia resources, the problem that existing agents are difficult to meet user needs is solved, and more accurate and efficient information display is achieved.

CN120030178APending Publication Date: 2025-05-23FACE CUTE CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510122562.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

When existing dialogue agents provide information, it is difficult for users to meet their needs. Users need to decide on the information platform and search content on their own, resulting in inaccurate search results.

Method used

By receiving the information input by the user, a target information platform matching the information is determined. The platform includes multimedia resources matching the user's needs and directly displays these resources in the agent's interactive interface.

Benefits of technology

It improves the display accuracy and efficiency of multimedia resources, reduces the steps of users to choose information platforms and search by themselves, and improves the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030178A_ABST
    Figure CN120030178A_ABST
Patent Text Reader

Abstract

The invention relates to an interaction method and device, electronic equipment, a storage medium and a program product, and relates to the technical field of artificial intelligence and computers. The interaction method disclosed by the invention comprises the following steps: receiving first information input by a user in an interaction interface with an intelligent agent; according to the first information, a target information platform matched with the first information is determined, and the target information platform comprises target multimedia resources matched with the demand information corresponding to the first information; and displaying the target multimedia resource in an interactive interface of the user and the intelligent agent.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the fields of artificial intelligence and computer technology, and in particular to an interaction method, device, electronic device, storage medium and program product. Background Art

[0002] With the development of artificial intelligence (AI) technology, various types of agents have emerged. Agents can interact with users. At present, common conversational agents can be used in many scenarios.

[0003] Conversational agents usually provide information to users in the form of dialogue. For example, a user enters a question and the answer provided by the agent is displayed in the interaction interface between the user and the agent. Summary of the invention

[0004] According to some embodiments of the present disclosure, an interaction method is provided, including: receiving first information input by a user in an interaction interface with an intelligent agent; determining a target information platform matching the first information based on the first information, wherein the target information platform includes target multimedia resources matching the demand information corresponding to the first information; and displaying the target multimedia resources in the interaction interface between the user and the intelligent agent.

[0005] According to other embodiments of the present disclosure, an interactive device is provided, including: a receiving module, configured to receive first information input by a user in an interactive interface with an intelligent agent; a determining module, configured to determine a target information platform matching the first information based on the first information, wherein the target information platform includes target multimedia resources matching the demand information corresponding to the first information; and a display module, configured to display the target multimedia resources in the interactive interface between the user and the intelligent agent.

[0006] According to some further embodiments of the present disclosure, there is provided an electronic device, comprising: a processor; and a memory coupled to the processor, for storing instructions, wherein when the instructions are executed by the processor, the processor executes an interaction method as in any embodiment of the present disclosure.

[0007] According to some further embodiments of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the program is executed by a processor, the interactive method of any embodiment of the present disclosure is performed.

[0008] Other features, aspects and advantages of the present disclosure will become apparent from the following detailed description of exemplary embodiments of the present disclosure with reference to the attached drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] The following is an explanation of the embodiments of the present disclosure with reference to the accompanying drawings. It should be understood that the drawings described below only relate to some embodiments of the present disclosure and do not constitute a limitation to the present disclosure. In the accompanying drawings:

[0010] Figure 1 A schematic diagram showing a flow chart of an interaction method according to some embodiments of the present disclosure;

[0011] Figure 2 A schematic diagram showing a display interface of some embodiments of the present disclosure;

[0012] Figure 3 A schematic diagram showing a flow chart of an interaction method according to some other embodiments of the present disclosure;

[0013] Figure 4 A schematic diagram showing the structure of an interactive device in some embodiments of the present disclosure;

[0014] Figure 5 A schematic diagram showing the structure of an electronic device according to some embodiments of the present disclosure;

[0015] Figure 6 A schematic diagram showing the structure of an electronic device according to some other embodiments of the present disclosure. DETAILED DESCRIPTION

[0016] The technical solutions in the embodiments of the present disclosure will be described clearly and completely below in conjunction with the accompanying drawings in the embodiments of the present disclosure. It should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments described herein.

[0017] It should be understood that the various steps described in the method embodiments of the present disclosure can be performed in different orders and / or performed in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect. Unless otherwise specifically stated, the relative arrangement of the steps set forth in these embodiments should be interpreted as being merely exemplary and not limiting the scope of the present disclosure.

[0018] The term “including” and its variations used in the present disclosure are intended to be open terms that include at least the following elements / features but do not exclude other elements / features, that is, “including but not limited to.” The term “based on” means “at least partly based on.”

[0019] It should be noted that the concepts of "first", "second", etc. mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units. Unless otherwise specified, the concepts of "first", "second", etc. are not intended to imply that the objects described in this way must be in a given order in time, space, ranking, or any other manner.

[0020] It should be noted that the modifications of "one" and "plurality" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, it should be understood as "one or more".

[0021] The embodiments of the present disclosure are described in detail below in conjunction with the accompanying drawings, but the present disclosure is not limited to these specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. In addition, in one or more embodiments, specific features, structures or characteristics can be combined in any suitable manner that will be clear from the present disclosure by a person of ordinary skill in the art.

[0022] Conversational agents usually provide information to users through conversations. In some cases, agents also provide some search results to users. Usually, the corresponding information platform (for example, an application or website) is pre-configured for the agent. After receiving the information input by the user, it searches these information platforms based on the information input by the user to obtain search results. However, since different information platforms focus on different content types and features, in many cases, the search results in some pre-configured information platforms cannot meet the needs of users. If users want to obtain search results that better meet their needs, they need to decide which information platform to go to and search by themselves.

[0023] Based on the above reasons, the present disclosure provides an interactive method, which responds to receiving first information input by a user in an interactive interface with an intelligent agent, determines a target information platform that matches the first information based on the first information, and the target information platform includes target multimedia resources that match the demand information corresponding to the first information, and then displays the target multimedia resources in the interactive interface between the user and the intelligent agent. Since the target information platform matches the first information and includes target multimedia resources that match the demand information corresponding to the first information, it is not a pre-configured arbitrary information platform, so that the displayed target multimedia resources can meet the needs of the user, thereby improving the accuracy of the displayed target multimedia resources. The target multimedia resources are directly displayed in the interactive interface between the user and the intelligent agent, and the user does not need to determine the target information platform by himself, nor does he need to search in the target information platform by himself, thereby improving the display efficiency of the target multimedia resources.

[0024] Combine the following Figures 1 to 3 Describe the interactive method of the present disclosure.

[0025] Figure 1 Flowcharts of some embodiments of the interactive method disclosed herein. Figure 1As shown, the method of this embodiment includes: steps S102 to S106. The method of this embodiment can be executed by an intelligent agent, or an application or device including an intelligent agent. The intelligent agent can be implemented in software or a combination of software and hardware.

[0026] In step S102, first information input by a user in an interaction interface with an agent is received.

[0027] For example, the interactive interface with the agent can be displayed on the user's terminal device, or through other display devices, without limitation to the examples given. The user can input the first information through text, voice, image, video, etc.

[0028] In step S104, a target information platform matching the first information is determined according to the first information.

[0029] One or more target information platforms matching the first information may be determined from multiple information platforms. The target information platform may include at least one type of platform of an application and a website. The target information platform includes target multimedia resources matching the demand information corresponding to the first information. For example, a search may be performed in the target information platform according to the demand information corresponding to the first information to obtain matching target multimedia resources. The target multimedia resources may include at least one type of content of text, image, audio, and video. The target multimedia resources may be one or more.

[0030] In step S106, the target multimedia resource is displayed in the interaction interface between the user and the agent.

[0031] The display style of the target multimedia resource may be the same as or different from the style displayed in the target information platform. For example, a thumbnail of the target multimedia resource may be displayed, and in response to a user triggering operation on the thumbnail, the target multimedia resource may be enlarged for display or played. The page of the target information platform may be displayed in the interactive interface between the user and the agent, and the target multimedia resource may be displayed in the page of the target information platform, or the target multimedia resource may be directly displayed in the interactive interface between the user and the agent. There may be a variety of methods for displaying the target multimedia resource, not limited to the examples given.

[0032] In the method of the above embodiment, since the target information platform matches the first information and includes target multimedia resources that match the demand information corresponding to the first information, and is not a pre-configured arbitrary information platform, the displayed target multimedia resources can meet the needs of the user, thereby improving the accuracy of the displayed target multimedia resources. The target multimedia resources are directly displayed in the interactive interface between the user and the intelligent agent, and the user does not need to determine the target information platform by himself, nor does he need to search in the target information platform by himself, thereby improving the display efficiency of the target multimedia resources and enhancing the user experience.

[0033] The following describes how to determine the target information platform that matches the first information.

[0034] In some embodiments, determining a target information platform matching the first information according to the first information includes: determining demand information corresponding to the first information according to semantic information of the first information; and determining a target information platform matching the demand information according to the demand information.

[0035] For example, a machine learning model can be used to semantically understand the first information and determine the demand information corresponding to the first information. For example, the machine learning model can be an LLM (Large Language Model), which is not limited to the examples given. The demand information corresponding to the first information can also be determined in combination with the semantic information of the context of the first information, and then the target information platform that matches the demand information can be determined. For example, the first information input by the user is "I want to relax and listen to a song", it can be determined that the demand information corresponding to the first information is to find a relaxing song, and then the determined target information platform can be a music playback platform.

[0036] By semantically understanding the first information, the demand information is determined, and then the target information platform is determined, which better meets the needs of users and can display more accurate and demand-compliant target multimedia resources to users.

[0037] In some embodiments, determining the demand information corresponding to the first information based on the semantic information of the first information includes: determining the basic demand information corresponding to the first information based on the semantic information of the first information; expanding the basic demand information to obtain expanded demand information; and determining at least one of the basic demand information and the expanded demand information as the demand information.

[0038] The basic demand information can be determined based on the first information, and then the basic demand information can be expanded to obtain the expanded demand information. The basic demand information can be expanded in combination with the context of the first information. For example, the user mentioned before that "I like singer A", and the first information is "I want to relax and listen to a song", it can be determined that the basic demand information is to find a relaxing song, and the expanded demand information is to find a relaxing song by singer A. The basic demand information can be associated and expanded. For example, the basic demand information is the introduction information of the Great Wall, and the expanded demand information is the travel guide of the Great Wall. The basic demand information can be associated and expanded based on the keywords in the basic demand information. For example, the keyword in the basic demand information is "science fiction", and keywords such as "alien creatures" can be expanded to obtain the expanded demand information.

[0039] For example, a machine learning model is used to expand the basic demand information to obtain the expanded demand information. The machine learning model may be LLM, etc. The above are examples for facilitating understanding of the basic demand information and the expanded demand information. The logic inside the actual machine learning model may not be consistent with the above examples.

[0040] By determining basic demand information and expanded demand information, users can be matched with target information platforms with more diverse and accurate types, and then more accurate and rich target multimedia resources can be determined to enhance user experience.

[0041] The following describes how to determine the target information platform that matches the required information.

[0042] In some embodiments, the basic demand information corresponds to a first demand category, and the expanded demand information corresponds to a second demand category. Based on the demand information, determining a target information platform that matches the demand information includes: determining an information platform corresponding to at least one of the first demand category and the second demand category as the target information platform.

[0043] The basic demand information may include a first demand category, and the expanded demand information may include a second demand category. For example, the first information may be semantically understood and classified using a machine learning model to obtain a first demand category, and the second demand category may be obtained by expanding the first demand category. For example, one or more demand categories associated with the first demand category are determined as the second demand category.

[0044] The first demand category can also be determined based on the first information or basic demand information, and the second demand category can be determined based on the expanded demand information. For example, if the first information is "Please introduce the history of the Great Wall", the first demand category can be determined to be "knowledge category". The second demand category can then be expanded to be "travel guide category".

[0045] The demand category corresponding to each information platform can be determined in advance based on the interactive data of the resources of each information platform among the multiple information platforms. For example, among the multiple information platforms, the interactive volume of the resources of the "travel guide category" in information platform A is the largest or exceeds the threshold, and information platform A can be set to correspond to the "travel guide category". The interactive volume of resources can be measured based on at least one indicator of click volume, search volume, comment volume, and like volume. An information platform can correspond to one or more demand categories.

[0046] The correspondence between the information platform and the demand category may be stored in a database, and according to at least one of the first demand category and the second demand category, the corresponding information platform may be searched in the database as the target information platform.

[0047] Based on the method of the above embodiment, the first demand category and the second demand category are determined according to the first information, and then the corresponding target information platform is determined according to the first demand category and the second demand category. This can improve the accuracy of determining the target information platform, and according to the expanded second demand category, the types of matched target information platforms can be increased, providing users with target multimedia resources with richer content and more in line with user needs.

[0048] In some embodiments, determining a target information platform that matches the demand information based on the demand information includes: obtaining index information of multiple information platforms, wherein the index information includes summary information of multimedia resources of each of the multiple information platforms; determining a target information platform that matches the demand information from the multiple information platforms based on the demand information and semantic information of the summary information of the multimedia resources of each information platform.

[0049] An index library of information platforms can be established. For each information platform, semantic understanding can be performed on the multimedia resources in the information platform to generate summary information of the multimedia resources. For example, the title, content, etc. of each multimedia resource can be understood to generate summary information of each multimedia resource. Multimedia resources of the same type can also be summarized to obtain summary information of each type of multimedia resources. Interaction data of each multimedia resource can also be summarized to obtain summary information. For example, the summary information of multimedia resources of each information platform includes at least one of summary information of each multimedia resource in the information platform, summary information of each type of multimedia resources, summary information of interaction data of each multimedia resource, and summary information of interaction data of each type of multimedia resources.

[0050] The semantic information of the demand information can be matched with the semantic information of the summary information of the multimedia resources of each information platform, and a target information platform matching the demand information can be determined from multiple information platforms.

[0051] The method of the above embodiment constructs an index library of the information platform including summary information of the multimedia resources in each information platform, and then matches the demand information with the summary information to obtain the target information platform, which can improve the accuracy of determining the target information platform, and then improve the accuracy of determining the target multimedia resources, better meet the needs of users, and enhance the user experience.

[0052] The interactive interface between the user and the agent can also display reply information generated based on the first information. For example, a machine learning model is used to generate reply information based on the semantic information of the first information. For example, the first information is "Please introduce the history of the Great Wall", and the generated reply information can be the introduction information of the history of the Great Wall. For example, the reply information can be generated based on the semantic information of the first information and the semantic information of the target multimedia resource. The reply information may include summary information of the target multimedia resource. The reply information may also include guidance information corresponding to the target multimedia resource. The guidance information can be generated based on the target information platform and the target multimedia resource. For example, the guidance information is "I also found a video about the history of the Great Wall for you on information platform X, please take a look."

[0053] Generating reply information according to the target multimedia resource and the first information can make the displayed reply information and the target multimedia resource mutually associated, so that the user can better understand the reply information and the target multimedia resource, and improve the overall display effect.

[0054] Combine the following Figure 2 A scheme that describes how the response message and the target multimedia resource should be displayed.

[0055] In some embodiments, the interaction interface between the user and the agent includes a first column and a second column. In the first column, reply information generated according to the first information is displayed, and in the second column, the target multimedia resource is displayed.

[0056] A double-column display method can be used to display the reply information and the target multimedia resource. Different contents are displayed in the double columns, and different interactive operations can be performed. For example, the first column may include an input area and a display area for interaction between the user and the agent and display of interactive information, and the second column displays the target multimedia resource, and the user can interact with the target multimedia resource, for example, by triggering the operation of the target multimedia resource to enlarge and display the target multimedia resource, play the target multimedia resource, jump to the target information platform, etc.

[0057] For example, the target information platform page can be displayed in the second column, and the target multimedia resources can be displayed in the page. For example, if the target information platform is a website, the web page can be displayed in the second column, and the target multimedia resources can be displayed in the web page, and the display effect of the web page can be the same or similar to the display effect of opening the web page through a browser. For example, if the target information platform is an application (including applets, etc.), the application page can be displayed in the second column, and the target multimedia resources can be displayed in the application page.

[0058] like Figure 2 As shown, the interaction interface between the user and the agent includes a first column 201 and a second column 202. The first column can display a reply message 203, and the second column can display multiple target multimedia resources 204. The first column can also include an input area 205, and the user can input information through text, voice, etc. Each target multimedia resource can display the corresponding name, number of likes, author's logo, etc.

[0059] By displaying reply information and target multimedia resources in separate columns, and displaying different types of content in different columns, users can quickly locate the area where the required information is located according to their needs, thereby improving reading efficiency and accuracy. In addition, the reply information and target multimedia resources can be displayed at the same time without page jumps, which facilitates user reading and reduces user operations.

[0060] The target multimedia resources can be obtained by the following method.

[0061] In some embodiments, in the interaction interface between the user and the intelligent agent, displaying the target multimedia resources includes: determining a search term based on semantic information of the first information; generating second information corresponding to the target information platform based on the search term; sending the second information to the target information platform; and displaying the search results obtained based on the second information in the target information platform as the target multimedia resources in the interaction interface between the user and the intelligent agent.

[0062] Keywords can be extracted as search terms based on the semantic information of the first information, or search terms can be generated based on the demand information corresponding to the first information. The search terms can be directly used as the second information, or the search terms can be expanded to generate the second information. For example, a link corresponding to the target information platform is generated based on the second information, and a request is sent to the target information platform through the link, the request includes the second information, and then the search results returned by the target information platform are obtained and displayed as the target multimedia resource in the interactive interface between the user and the intelligent agent. For example, the link is a URL (Uniform Resource Locator), etc., and the second information can be spliced ​​into the search URL of the target information platform.

[0063] Through the method of the above embodiment, it is possible to search more accurately in the target information platform and obtain more accurate target multimedia resources that better meet user needs.

[0064] In some embodiments, in the interaction interface between the user and the intelligent agent, displaying the target multimedia resources includes: determining the search terms and the search result type based on the semantic information of the first information; generating second information corresponding to the target information platform based on the search terms and the search result type; sending the second information to the target information platform; and displaying the search results obtained based on the second information in the target information platform as the target multimedia resources in the interaction interface between the user and the intelligent agent.

[0065] In addition to determining the search term, the search result type may also be determined based on the semantic information of the first information or based on the demand information corresponding to the first information. For example, when the first information includes keywords corresponding to the search result type, the search result type is determined based on the keywords. For example, if the first information is "I want to relax and listen to a song", the search result type may be determined to be audio and / or video.

[0066] For example, the demand information includes at least one of basic demand information and extended demand information, the basic demand information corresponds to the first demand category, and the extended demand information corresponds to the second demand category. The search result type can be determined based on at least one of the first demand category and the second demand category. For example, the correspondence between the demand category and the search result type can be preconfigured. For example, the search result type corresponding to the travel guide category is video and / or graphic text, etc.

[0067] For example, a machine learning model is used to determine the search result type according to the semantic information of the first information. The machine learning model can perform semantic understanding on the first information and determine a suitable search result type.

[0068] like Figure 2 As shown, in the second column, search terms 206 and search result types 207 may also be displayed. For example, a page of the target information platform is displayed in the second column, search terms are displayed in the search area of ​​the page, a plurality of tabs may be included in the page, and target multimedia resources are displayed in the tab corresponding to the search result type.

[0069] The method of the above embodiment determines the search term and the search result type according to the first information, and further determines the type of the target multimedia resource, which can better meet the needs of the user and improve the accuracy of the displayed target multimedia resources.

[0070] The user can modify the second information to obtain new search results. For example, the target information platform also displays the second information, and the interaction method further includes: responding to the user's modification of the second information; sending the modified second information to the target information platform; and displaying the search results obtained in the target information platform according to the modified second information in the interaction interface between the user and the agent.

[0071] For example, users can modify search terms and search result types. Figure 2 As shown, the user can directly modify the search term in the search area in the second column, and can modify the search result type by switching tabs. The user can also enter prompt information for modifying the second information in the first column. In response to the user inputting the prompt information for modifying the second information, the modified second information is generated according to the prompt information. For example, the user enters "Please change to the picture and text of Beijing Food Guide".

[0072] Of course, the user may also re-enter the information. Based on the method of the aforementioned embodiment, the target information platform and the target multimedia resources may be re-determined and displayed, which will not be described in detail here.

[0073] Users can also further interact with target multimedia resources to generate new content.

[0074] In some embodiments, a generation control is displayed in the first column, and the interaction method further includes: in response to a user's selection operation on one or more multimedia resources in the target multimedia resources, and a user's triggering operation on the generation control, generating target content according to the one or more multimedia resources; and displaying the target content in the first column.

[0075] The user can select one or more multimedia resources in the second column by clicking or other preset operations, and further trigger the generation control in the first column to generate the target content. It is also possible to trigger the generation control first, and then select one or more multimedia resources. For example, a machine learning model can be used to semantically understand one or more multimedia resources to generate target content. The generated target content can be at least one type of content in text, audio, video, and image. The target content can be displayed in the upper layer of the second column, for example, in a floating layer, a mask, or a window in the upper layer of the second column, without limitation to the examples given.

[0076] The generated controls can also be displayed in the second column, or displayed on top of the second column.

[0077] In some embodiments, the interaction method also includes: in response to a user's selection operation on one or more multimedia resources in the target multimedia resources, displaying a generation control in the upper layer of the second column; in response to a user's triggering operation on the generation control, generating target content based on one or more multimedia resources; and displaying the target content in the first column.

[0078] For example, the generated control may be displayed in a floating layer or a mask layer on the upper layer of the second column, but is not limited to the examples given.

[0079] Based on the solution of the above embodiment, the user can generate new target content by selecting one or more multimedia resources through simple operations, thereby improving the efficiency of generating target content. The user does not need to fully view and understand each multimedia resource by himself, so that the user can obtain the desired content more quickly and accurately, better meet the user's needs, and improve the user experience.

[0080] The generation control may be one or more different types of generation controls. For example, the generation control may include at least one of a summary control and a comparison control.

[0081] In some embodiments, the generated control includes a control for summarizing one or more multimedia resources, and generating target content based on the one or more multimedia resources includes: summarizing the one or more multimedia resources based on semantic information of the one or more multimedia resources, and generating summary information as the target content.

[0082] A control that summarizes one or more multimedia resources is called a summary control. Figure 2 As shown, the summary control 208 may be displayed in the first column. The summary control 208 may also be displayed in the second column or in the upper layer of the second column. For example, the semantic information of each multimedia resource is understood by using a machine learning model, and then the summary information is summarized and generated.

[0083] Through the method of the above embodiment, the user can select one or more multimedia resources and let the intelligent agent summarize them, which saves the user's time in viewing all the content, facilitates the user's operation and reading, and improves efficiency.

[0084] In some embodiments, one or more multimedia resources include multiple multimedia resources, the generated control includes a control for comparing the multiple multimedia resources, and the generated target content based on the one or more multimedia resources includes: comparing the multiple multimedia resources based on semantic information of the multiple multimedia resources, and generating comparison information as the target content.

[0085] A control that compares multiple multimedia resources is called a comparison control. Figure 2As shown, the comparison control 209 can be displayed in the first column. The comparison control 209 can also be displayed in the second column or in the upper layer of the second column. For example, the semantic information of each multimedia resource is understood by using a machine learning model, and multiple multimedia resources are compared to generate comparison information. For example, two videos about Beijing travel guides can be compared to determine the differences in the attractions visited.

[0086] Through the method of the above embodiment, the user can select multiple multimedia resources and let the intelligent agent compare them, which saves the user's time in viewing all the content and allows the user to quickly and accurately determine the difference information between different multimedia resources.

[0087] The generation control may also include a type conversion control, for example, in response to a user triggering operation on the type conversion control, one or more multimedia resources may be converted into a preset type, for example, a video may be converted into a picture or text.

[0088] In addition to using the generation control to quickly and accurately obtain the desired target content, users can also generate the target content by entering prompt information.

[0089] In some embodiments, the interaction method also includes: receiving prompt information input by the user in the first column or the second column, wherein the prompt information includes selection information of one or more multimedia resources in the target multimedia resources; generating target content based on the prompt information and the one or more multimedia resources; and displaying the target content in the first column.

[0090] An input area can also be set in the second column, and the user can enter prompt information in the second column. The prompt information includes selection information of one or more multimedia resources. The prompt information can also include at least one of the subject and type of the target content. The machine learning model can be used to perform semantic understanding of the prompt information and semantic understanding of one or more multimedia resources to generate target content that conforms to at least one of the subject and the type. For example, the prompt information is "Please generate a summary text based on the first and second videos." If the user does not enter the type of the target content, the machine learning model can be used to determine the type of the target content based on the prompt information and the type of multimedia resources to match the user's needs.

[0091] According to the method of the above embodiment, the user can more flexibly control the generation of the target content by inputting prompt information, and can generate target content with more types and richer themes, which better meets the needs of the user and improves the user experience.

[0092] In other embodiments, in response to a user's selection operation of one or more multimedia resources in the target multimedia resources, and prompt information is input in the first column or the second column, wherein the prompt information includes at least one of the subject and type of the target content, the target content is generated based on the prompt information and the one or more multimedia resources; and the target content is displayed in the first column.

[0093] The user can directly select one or more multimedia resources without carrying the selection information in the prompt information, which is more convenient. The user only needs to input at least one of the subject and type of the target content, which reduces the input of information and improves the operation efficiency.

[0094] The user can also specify the reference section of each multimedia resource in the prompt information.

[0095] In some embodiments, the prompt information includes at least one of the subject and type of the target content, and information indicating a reference part of each of the one or more multimedia resources. Generating the target content based on the prompt information and the one or more multimedia resources includes: determining the reference part of each multimedia resource based on the prompt information and the semantic information of each multimedia resource; summarizing the one or more multimedia resources based on at least one of the subject and type of the target content and the semantic information of the reference part of each multimedia resource to generate the target content.

[0096] For example, the prompt information input by the user is "Please generate a travel guide based on the hotels in video 1 and the attractions in video 2". The machine learning model can be used to perform semantic understanding on the prompt information, determine the reference part of each multimedia resource specified by the user, perform semantic understanding on one or more multimedia resources, determine the content of the reference part of each multimedia resource, and then generate target content that meets at least one of the theme and type in the prompt information based on the content of the reference part of each multimedia resource.

[0097] Based on the method of the above embodiment, the user can generate richer and more diverse target content by inputting prompt information. The generated target content better meets the needs of the user, thereby improving the generation efficiency and accuracy of the target content.

[0098] The user may also ask further questions about one or more multimedia resources in the target multimedia resources.

[0099] In some embodiments, the interaction method further includes: responding to user input of question information regarding one or more multimedia resources of the target multimedia resource; generating an answer based on the question information and the one or more multimedia resources; and displaying the answer in an interaction interface between the user and the agent.

[0100] For example, a user can ask the agent "What hotel is the person staying in in video 1?" The machine learning model can be used to understand the semantics of one or more multimedia resources and generate corresponding answers based on the question information.

[0101] Based on the method of the above embodiment, the user can interact deeply with the intelligent agent for one or more multimedia resources, locate the content that the user wants to know more accurately and quickly, and improve reading efficiency.

[0102] like Figure 2 As shown, the first column may also include a control 210 corresponding to the target information platform. In response to the user triggering the control 210, it can jump to the target information platform and display the target multimedia resources. The first column may also display a control 211 for one or more candidate questions related to the first information. In response to the user triggering the control of the target question in one or more candidate questions, the answer generated according to the target question is displayed. For example, a machine learning model is used to generate one or more related questions based on the first information as one or more candidate questions. For example, the first information is "Beijing Travel Guide", and one or more candidate questions may be "Beijing Food Guide", "Beijing Attractions", etc.

[0103] In each of the above embodiments, the user can further interact with the intelligent agent with respect to the target multimedia resources. The intelligent agent can act as an assistant to provide the user with generation functions and help the user answer related questions, thereby facilitating the user's operation and assisting the user to quickly and accurately understand the target multimedia resources, thereby improving the user experience.

[0104] Figure 3 Flow charts of other embodiments of the interactive method disclosed herein. Figure 3 As shown, the method of this embodiment includes: steps S302 to S314.

[0105] In step S302, first information input by a user in an interaction interface with an agent is received.

[0106] In step S304, demand information corresponding to the first information is determined according to the semantic information of the first information.

[0107] In step S306, a target information platform matching the demand information is determined according to the demand information.

[0108] In step S308, target multimedia resources in the target information platform are determined according to the demand information.

[0109] In step S310, reply information is generated according to the first information.

[0110] In step S312, the reply information is displayed in the first column of the interaction interface between the user and the agent, and the target multimedia resource is displayed in the second column.

[0111] In step S314, in response to the user's selection operation on one or more multimedia resources in the target multimedia resources and the triggering of the generation instruction, target content is generated according to the one or more multimedia resources.

[0112] The user can select one or more multimedia resources and trigger the generation of instructions (summary, comparison, etc.) through controls, prompt information, etc., and can refer to the aforementioned embodiments, which will not be described in detail.

[0113] The method of the above embodiment can automatically determine the target information platform that meets the user's needs, determine the target multimedia resources in the target information platform, and generate reply information for the first information input by the user. The reply information and the target multimedia resources are displayed simultaneously in the interactive interface in a double-column form, and the user can also further interact with the intelligent agent for one or more multimedia resources. The method of the above embodiment improves the accuracy and efficiency of the display of the target multimedia resources. The intelligent agent can assist the user in understanding and generating content that meets the user's needs as an assistant, thereby improving accuracy and efficiency.

[0114] The present disclosure also provides an interactive device, Figure 4 Give a description.

[0115] Figure 4 FIG. 1 is a structural diagram of some embodiments of the interactive device disclosed in the present invention. Figure 4 As shown, the interaction device 40 of this embodiment includes: a receiving module 410 , a determining module 420 , and a display module 430 .

[0116] The receiving module 410 is configured to receive first information input by a user in an interaction interface with an agent.

[0117] The determination module 420 is configured to determine, based on the first information, a target information platform that matches the first information, wherein the target information platform includes target multimedia resources that match the demand information corresponding to the first information.

[0118] The display module 430 is configured to display the target multimedia resource in the interaction interface between the user and the agent.

[0119] In some embodiments, the determination module 420 is configured to determine, based on the first information, a target information platform that matches the first information, wherein the target information platform includes target multimedia resources that match the demand information corresponding to the first information.

[0120] In some embodiments, the determination module 420 is configured to determine basic demand information corresponding to the first information based on semantic information of the first information; expand the basic demand information to obtain expanded demand information; and determine at least one of the basic demand information and the expanded demand information as demand information.

[0121] In some embodiments, the basic demand information corresponds to a first demand category, the expanded demand information corresponds to a second demand category, and the determination module 420 is configured to determine an information platform corresponding to at least one of the first demand category and the second demand category as the target information platform.

[0122] In some embodiments, the determination module 420 is configured to obtain index information of multiple information platforms, wherein the index information includes summary information of multimedia resources of each of the multiple information platforms; and determine a target information platform that matches the demand information from the multiple information platforms based on the demand information and semantic information of the summary information of the multimedia resources of each information platform.

[0123] In some embodiments, the interaction interface between the user and the agent includes a first column and a second column. In the first column, reply information generated according to the first information is displayed, and in the second column, the target multimedia resource is displayed.

[0124] In some embodiments, a generation control is displayed in the first column, and the interaction device also includes: a generation module 440, which is configured to generate target content based on one or more multimedia resources in response to a user's selection operation on one or more multimedia resources in the target multimedia resources, and a user's triggering operation on the generation control; the display module 430 is also configured to display the target content in the first column.

[0125] In some embodiments, the display module 430 is also configured to display a generation control in the upper layer of the second column in response to a user's selection operation on one or more multimedia resources in the target multimedia resources; the interactive device also includes: a generation module 440, which is configured to generate target content based on one or more multimedia resources in response to a user's triggering operation on the generation control; the display module 430 is also configured to display the target content in the first column.

[0126] In some embodiments, the generation control includes a control for summarizing one or more multimedia resources, and the generation module 440 is configured to summarize the one or more multimedia resources according to the semantic information of the one or more multimedia resources and generate summary information as the target content.

[0127] In some embodiments, one or more multimedia resources include multiple multimedia resources, the generated control includes a control for comparing the multiple multimedia resources, and the generation module 440 is configured to compare the multiple multimedia resources according to the semantic information of the multiple multimedia resources and generate comparison information as the target content.

[0128] In some embodiments, the receiving module 410 is also configured to receive prompt information input by the user in the first column or the second column, wherein the prompt information includes selection information of one or more multimedia resources in the target multimedia resources; the interactive device also includes: a generating module 440, configured to generate target content based on the prompt information and one or more multimedia resources; the display module 430 is also configured to display the target content in the first column.

[0129] In some embodiments, the prompt information includes at least one of the subject and type of the target content, and information indicating the reference part of each multimedia resource in one or more multimedia resources. The generation module 440 is configured to determine the reference part of each multimedia resource based on the prompt information and the semantic information of each multimedia resource; summarize the one or more multimedia resources based on at least one of the subject and type of the target content and the semantic information of the reference part of each multimedia resource to generate the target content.

[0130] In some embodiments, the determination module 420 is also configured to determine the search terms and search result types based on the semantic information of the first information; generate second information corresponding to the target information platform based on the search terms and search result types; the display module 430 is also configured to send the second information to the target information platform; and display the search results obtained based on the second information in the target information platform as the target multimedia resources in the interaction interface between the user and the intelligent agent.

[0131] In some embodiments, the target information platform also displays the second information, and the receiving module 410 is further configured to receive user modifications to the second information; the display module 430 is further configured to send the modified second information to the target information platform; and in the interactive interface between the user and the agent, the search results obtained in the target information platform based on the modified second information are displayed.

[0132] In some embodiments, the receiving module 410 is also configured to receive question information input by the user regarding one or more multimedia resources of the target multimedia resource; the interactive device also includes: a generating module 440, configured to generate an answer based on the question information and one or more multimedia resources; the display module 430 is also configured to display the answer in the interactive interface between the user and the intelligent agent.

[0133] Figure 5A block diagram of an electronic device according to some embodiments of the present disclosure is shown.

[0134] The memory 51 is used to store one or more computer-readable instructions. The memory 51 may include any combination of various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory, including but not limited to random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), read-only memory (ROM), and flash memory. The memory 51 may store, for example, an operating system, an application, a boot loader (BootLoader), a database, and other programs, and may also store various applications and various data.

[0135] The processor 52 is used to run computer-readable instructions to implement the interaction method described in any of the above embodiments. The specific implementation of each step of the method can be referred to the above embodiments, and the repeated parts will not be repeated here.

[0136] The processor 52 may be configured to execute Figures 1 to 3 The processor 52 may be embodied as various processing devices, such as a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The central processing unit (CPU) may be an X86 or ARM architecture, etc.

[0137] The processor 52 and the memory 51 may communicate with each other directly or indirectly. For example, the processor 52 and the memory 51 may communicate with each other via a network. The network may include a wireless network, a wired network, and / or any combination of a wireless network and a wired network. The processor 52 and the memory 51 may also communicate with each other via a system bus, which is not limited in the present disclosure.

[0138] It should be noted that Figure 5 The components of the electronic device 5 shown are only exemplary and non-restrictive. The electronic device 5 may also have other components according to actual application requirements. The processor 52 may control other components in the electronic device 5 to perform desired functions.

[0139] The electronic device 5 may be implemented by software, firmware and / or hardware, and may be integrated into a device installed with relevant application programs.

[0140] Figure 6 A block diagram of an electronic device according to some other embodiments of the present disclosure is shown.

[0141] Figure 6The electronic device 6 shown may be a computer system with a dedicated hardware structure, which can execute corresponding functions when a relevant application program is installed.

[0142] Electronic devices include, but are not limited to, mobile terminals such as smart phones, laptops, personal digital assistants (PDA), tablet computers (Tablet Personal Computer, Tablet PC), PMP (portable multimedia player), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), wearable devices, etc., as well as fixed terminals such as digital televisions, desktop computers, etc.

[0143] like Figure 6 As shown, the central processing unit (CPU) 61 performs various processes according to the program stored in the read-only memory (ROM) 62 or the program loaded from the storage part 68 to the random access memory (RAM) 63. In the RAM 63, data required when the CPU 61 performs various processes, etc. is stored as needed. The central processing unit is only exemplary, and it can also be other types of processors, such as the various processors described above. The ROM 62, RAM 63 and the storage part 68 can be various forms of computer-readable storage media. It should be noted that although Figure 6 ROM 62, RAM 63 and storage section 68 are shown separately in FIG. 1 , but one or more of them may be combined or located in the same or different memory or storage modules.

[0144] The CPU 61, the ROM 62, and the RAM 63 are connected to one another via a bus 64. To the bus 64, an input / output interface 65 is also connected.

[0145] The following components are connected to the input / output interface 65: an input section 66, such as a touch screen, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output section 67, including a display, such as a cathode ray tube (CRT), a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage section 68, including a hard disk, a magnetic tape, etc.; and a communication section 69, including a network interface card such as a LAN card, a modem, etc. The communication section 69 allows communication processing to be performed via a network such as the Internet. It is easy to understand that although Figure 6 Some of the electronic devices 6 are shown to communicate via a bus 64, but they may also communicate via a network or other means, wherein the network may include a wireless network, a wired network, and / or any combination of a wireless network and a wired network.

[0146] A drive 610 is also connected to the input / output interface 65 as needed. A removable medium 611 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory or the like is mounted on the drive 610 as needed so that a computer program read therefrom is installed into the storage section 68 as needed.

[0147] When the above-described series of processing is implemented by software, a program constituting the software can be installed from a network such as the Internet or a storage medium such as the removable medium 611 .

[0148] According to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, some embodiments of the present disclosure include a computer program product, which, when the computer program product is run on a computer, enables the computer to implement the method described in any of the aforementioned embodiments. The computer program product includes a computer instruction carried on a computer-readable medium, containing a program code for executing the method shown in the flowchart. In such an embodiment, the computer instruction can be downloaded and installed from the network through the communication part 69, or installed from the storage part 68, or installed from the ROM 62. When the computer program is executed by the CPU 61, the method of the embodiment of the present disclosure is executed.

[0149] It should be noted that, in the context of the present disclosure, a computer-readable medium may be a tangible medium that may contain or store a program for use by an instruction execution system, apparatus, or device or for use in conjunction with an instruction execution system, apparatus, or device.

[0150] The computer readable medium may be a computer readable storage medium, or a computer readable signal medium, or any combination of the two.

[0151] Computer-readable storage media include, but are not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices or components, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, electrical connections with one or more wires, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination thereof. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that may be used by or in conjunction with an instruction execution system, device, or device. Computer instructions are stored on a computer-readable storage medium, and when the instructions are executed by a processor, the method described in any of the foregoing embodiments is implemented.

[0152] Computer readable signal media may include data signals propagated in baseband or as part of a carrier wave, which carry computer readable program codes. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Computer readable signal media may also be any computer readable medium other than a computer readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, device, or device. The program code contained on the computer readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0153] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.

[0154] In some embodiments, a computer program is further provided, comprising: instructions, which, when executed by a processor, cause the processor to execute the method described in any of the above embodiments. For example, the instructions may be embodied as computer program codes.

[0155] In embodiments of the present disclosure, computer program codes for performing the operations of the present disclosure may be written in one or more programming languages ​​or combinations thereof, including but not limited to object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" language or similar programming languages. The program code may be executed entirely on a user's computer, partially on a user's computer, as an independent software package, partially on a user's computer, partially on a remote computer, or entirely on a remote computer or server. In situations involving a remote computer, the remote computer may be connected to the user's computer via any type of network (including a local area network (LAN) or a wide area network (WAN)), or may be connected to an external computer (e.g., using an Internet service provider to connect via the Internet).

[0156] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some implementations as replacements, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0157] The functions described above may be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.

[0158] Although some specific embodiments of the present disclosure have been described in detail by way of example, it should be understood by those skilled in the art that the above examples are for illustration only and are not intended to limit the scope of the present disclosure. It should be understood by those skilled in the art that the above embodiments may be modified without departing from the scope and spirit of the present disclosure. The scope of the present disclosure is defined by the appended claims.

Claims

1. An interactive method, comprising: Receiving first information input by a user in an interaction interface with an agent; Determining, according to the first information, a target information platform matching the first information, wherein the target information platform includes target multimedia resources matching the demand information corresponding to the first information; The target multimedia resource is displayed in the interaction interface between the user and the agent.

2. The interactive method according to claim 1, wherein determining, based on the first information, a target information platform matching the first information comprises: Determining, according to the semantic information of the first information, demand information corresponding to the first information; According to the demand information, a target information platform matching the demand information is determined.

3. The interactive method according to claim 2, wherein: The determining, according to the semantic information of the first information, the requirement information corresponding to the first information includes: Determining basic requirement information corresponding to the first information according to the semantic information of the first information; Expanding the basic demand information to obtain expanded demand information; At least one of the basic requirement information and the expanded requirement information is determined as the requirement information.

4. The interactive method according to claim 3, wherein: The basic demand information corresponds to a first demand category, the expanded demand information corresponds to a second demand category, and determining, based on the demand information, a target information platform that matches the demand information includes: An information platform corresponding to at least one of the first demand category and the second demand category is determined as the target information platform.

5. The interactive method according to claim 2, wherein: The determining, according to the demand information, a target information platform matching the demand information comprises: Acquire index information of multiple information platforms, wherein the index information includes summary information of multimedia resources of each of the multiple information platforms; According to the demand information and the semantic information of the summary information of the multimedia resources of each information platform, a target information platform matching the demand information is determined from the multiple information platforms.

6. The interactive method according to claim 1, wherein: The interaction interface between the user and the agent includes a first column and a second column. The first column displays reply information generated according to the first information, and the second column displays the target multimedia resource.

7. The interactive method according to claim 6, wherein: The generation control is displayed in the first column, and the interaction method further includes: In response to the user's selection operation on one or more multimedia resources among the target multimedia resources and the user's triggering operation on the generation control, generating target content according to the one or more multimedia resources; The target content is displayed in the first column.

8. The interactive method according to claim 6, further comprising: In response to the user's selection operation on one or more multimedia resources in the target multimedia resources, displaying a generation control in an upper layer of the second column; In response to a triggering operation of the generating control by the user, generating target content according to the one or more multimedia resources; The target content is displayed in the first column.

9. The interactive method according to claim 7 or 8, wherein: The generating control includes a control for summarizing the one or more multimedia resources, and the generating target content according to the one or more multimedia resources includes: The one or more multimedia resources are summarized according to the semantic information of the one or more multimedia resources to generate summary information as the target content.

10. The interactive method according to claim 7 or 8, wherein: The one or more multimedia resources include multiple multimedia resources, the generating control includes a control for comparing the multiple multimedia resources, and the generating target content according to the one or more multimedia resources includes: The plurality of multimedia resources are compared according to the semantic information of the plurality of multimedia resources to generate comparison information as the target content.

11. The interactive method according to claim 6, further comprising: receiving prompt information input by the user in the first sub-column or the second sub-column, wherein the prompt information includes selection information of one or more multimedia resources among the target multimedia resources; generating target content according to the prompt information and the one or more multimedia resources; The target content is displayed in the first column.

12. The interactive method according to claim 11, wherein: The prompt information includes at least one of a subject and a type of the target content, and information indicating a reference portion of each of the one or more multimedia resources, and generating the target content according to the prompt information and the one or more multimedia resources includes: Determining a reference portion of each multimedia resource according to the prompt information and the semantic information of each multimedia resource; The one or more multimedia resources are summarized according to at least one of the subject and type of the target content and the semantic information of the reference part of each multimedia resource to generate the target content.

13. The interactive method according to any one of claims 1 to 12, wherein: In the interactive interface between the user and the agent, displaying the target multimedia resource includes: Determining a search term and a search result type according to the semantic information of the first information; Generate second information corresponding to the target information platform according to the search term and the search result type; Sending the second information to the target information platform; In the interactive interface between the user and the agent, the search results obtained in the target information platform based on the second information are displayed as the target multimedia resources.

14. The interactive method according to claim 13, wherein: The target information platform also displays the second information, and the interaction method further includes: receiving a modification of the second information by the user; Sending the modified second information to the target information platform; In the interactive interface between the user and the agent, the search results obtained in the target information platform according to the modified second information are displayed.

15. The interactive method according to any one of claims 1 to 12, further comprising: receiving question information input by the user regarding one or more multimedia resources of the target multimedia resource; generating an answer according to the question information and the one or more multimedia resources; The answer is displayed in the interaction interface between the user and the agent.

16. An interactive device, comprising: A receiving module, configured to receive first information input by a user in an interaction interface with an agent; a determination module configured to determine, based on the first information, a target information platform matching the first information, wherein the target information platform includes target multimedia resources matching the demand information corresponding to the first information; The display module is configured to display the target multimedia resource in the interaction interface between the user and the agent.

17. An electronic device comprising: processor; as well as A memory coupled to the processor, for storing instructions, wherein when the instructions are executed by the processor, the processor executes the interaction method according to any one of claims 1 to 15.

18. A computer-readable storage medium having a computer program stored thereon, wherein: When the program is executed by a processor, the interactive method according to any one of claims 1 to 15 is implemented.

19. A computer program product comprising instructions which, when executed by a processor, cause the processor to perform the interaction method according to any one of claims 1 to 15.