Intelligent agent information interaction method and device based on large language model, electronic equipment, medium and program product

Through the agent information interaction method based on the large language model, multimodal information is used to create an agent, and information interaction between agents is realized through context connection, the limitations of agent creation and interaction in the existing technology are solved, and more rich and personalized services and better user experience are achieved.

CN120046731APending Publication Date: 2025-05-27BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510104524.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

When creating an agent, the prior art relies on a large amount of information input and one-sided manual settings, and cannot cover all vertical fields, and the information interaction between the agents is not smooth enough.

Method used

Adopt the agent information interaction method based on the large language model to create an agent by obtaining multimodal information, and interacting with other related agents through context connection, expanding the agent's knowledge base and user experience.

Benefits of technology

It realizes richer and personalized services for the agent, enhances user experience, and improves the coherence and consistency of information interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120046731A_ABST
    Figure CN120046731A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent agent information interaction method based on a large language model, and relates to the field of artificial intelligence, in particular to the fields of large language models, intelligent agents, intelligent search and the like. According to the implementation scheme, the method comprises the steps of obtaining multi-modal information associated with a target object interested by a first user; creating a first agent for the target object based on the multi-modal information; at least one second agent associated with the first agent is determined, and the at least one second agent performs information interaction with at least one of the first user or the second user; performing context connection on the first intelligent agent and at least one second intelligent agent; and performing information interaction with the first user through the first intelligent agent.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, particularly to fields such as large language models, intelligent agents, intelligent search, etc. Specifically, it relates to an information interaction method for an intelligent agent based on a large language model, a multi-agent information interaction method, an information search method, a device, an electronic device, a computer-readable storage medium, and a computer program product. Background Art

[0002] Artificial intelligence is a discipline that studies how to make computers simulate certain human thinking processes and intelligent behaviors (such as learning, reasoning, thinking, planning, etc.). It has both hardware-level technologies and software-level technologies. Artificial intelligence hardware technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing; artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech recognition technology, natural language processing technology, and machine learning / deep learning, big data processing technology, and knowledge graph technology.

[0003] An intelligent agent is an embodiment form of artificial intelligence in a specific task execution scenario. Artificial intelligence provides the technology and methods for the intelligent agent to operate effectively. In recent years, with the development of large language model technology, intelligent agents implemented based on large language models have gradually transformed from a simple tool role to a role closer to a human partner, being able to deeply understand human intentions and needs and actively provide help and support. Currently, intelligent agents have become a capable assistant in human work and life.

[0004] The methods described in this section are not necessarily methods that have been previously envisioned or adopted. Unless otherwise specified, any method described in this section should not be considered prior art solely because it is included in this section. Similarly, unless otherwise specified, the problems mentioned in this section should not be considered to have been recognized in any prior art. Summary of the Invention

[0005] The present disclosure provides an information interaction method for an intelligent agent based on a large language model, a multi-agent information interaction method, an information search method, a device, an electronic device, a computer-readable storage medium, and a computer program product.

[0006] According to one aspect of the present disclosure, there is provided an information interaction method for an intelligent agent based on a large language model, including: obtaining multi-modal information associated with a target object of interest to a first user; creating a first intelligent agent for the target object based on the multi-modal information; determining at least one second intelligent agent associated with the first intelligent agent, wherein the at least one second intelligent agent has had information interaction with at least one of the first user or a second user; performing context connection between the first intelligent agent and the at least one second intelligent agent; and performing information interaction between the first intelligent agent and the first user via the first intelligent agent.

[0007] According to another aspect of the present disclosure, there is provided a multi-agent information interaction method, including: creating a chat group including a user and multiple agents, wherein the user and the multiple agents perform information interaction based on the information interaction method of the large language model-based agent as described above.

[0008] According to another aspect of the present disclosure, there is provided an information search method, including: receiving multi-modal information provided by a user and associated with a target object of interest; performing information interaction with the user based on the information interaction method of the large language model-based agent as described above to provide a search result for the target object to the user.

[0009] According to another aspect of the present disclosure, there is provided an information interaction device for a large language model-based agent, including: an information acquisition module configured to acquire multi-modal information associated with a target object of interest to a first user; an agent creation module configured to create a first agent for the target object based on the multi-modal information; an associated agent determination module configured to determine at least one second agent associated with the first agent, wherein the at least one second agent has performed information interaction with at least one of the first user or a second user; a context connection module configured to connect the context of the first agent and the at least one second agent; and a first information interaction module configured to perform information interaction with the first user via the first agent.

[0010] According to another aspect of the present disclosure, there is provided a multi-agent information interaction device, including: a chat group creation module configured to create a chat group including a user and multiple agents, wherein the user and the multiple agents perform information interaction based on the information interaction device of the large language model-based agent as described above.

[0011] According to another aspect of the present disclosure, there is provided an information search device, including: an information reception module configured to receive multi-modal information provided by a user and associated with a target object of interest; and a search result providing module configured to perform information interaction with the user based on the information interaction device of the large language model-based agent as described above to provide a search result for the target object to the user.

[0012] According to another aspect of the present disclosure, there is provided an electronic device, including at least one processor; and a memory communicatively connected to the at least one processor; the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the information interaction method of the large language model-based agent, the multi-agent information interaction method, and the information search method as described above in the present disclosure.

[0013] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the information interaction method of the large language model-based agent, the multi-agent information interaction method, and the information search method as described above in the present disclosure.

[0014] According to another aspect of the present disclosure, there is provided a computer program product including a computer program which, when executed by a processor, implements the information interaction method of the large language model-based agent, the multi-agent information interaction method, and the information search method as described above in the present disclosure.

[0015] According to one or more embodiments of the present disclosure, the means for creating agents can be enriched and multi-agent interaction can be realized, thereby enhancing the user experience.

[0016] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The drawings exemplarily illustrate embodiments and form a part of the specification, and are used together with the written description of the specification to explain the exemplary embodiments. The illustrated embodiments are for illustrative purposes only and do not limit the scope of the claims. In all the drawings, the same reference numerals refer to similar but not necessarily identical elements.

[0018] Figure 1 A schematic diagram of an exemplary system in which the various methods described herein can be implemented according to an embodiment of the present disclosure;

[0019] Figure 2 A flowchart of an information interaction method of a large language model-based agent according to an embodiment of the present disclosure;

[0020] Figure 3 A schematic diagram of a process for determining a second agent associated with a first agent according to an embodiment of the present disclosure;

[0021] Figure 4 A schematic diagram of modifying historical context information based on retrieved updated information according to an embodiment of the present disclosure;

[0022] Figure 5 A schematic diagram of creating a first agent according to an embodiment of the present disclosure;

[0023] Figure 6 A schematic diagram of creating a chat group of a user and multiple agents according to an embodiment of the present disclosure;

[0024] Figure 7 shows a flowchart of an information search method according to an embodiment of the present disclosure;

[0025] Figure 8 shows a structural block diagram of an information interaction device of an agent based on a large language model according to an embodiment of the present disclosure;

[0026] Figure 9 shows a structural block diagram of an information interaction device of an agent based on a large language model according to another embodiment of the present disclosure;

[0027] Figure 10 shows a structural block diagram of a multi-agent information interaction device according to an embodiment of the present disclosure;

[0028] Figure 11 shows a structural block diagram of an information search device according to an embodiment of the present disclosure;

[0029] Figure 12 shows a structural block diagram of an exemplary electronic device that can be used to implement the embodiments of the present disclosure. Detailed implementation manners

[0030] The following makes an explanation of the exemplary embodiments of the present disclosure with reference to the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for the sake of clarity and conciseness, the description of well-known functions and structures is omitted below.

[0031] In the present disclosure, unless otherwise specified, the terms "first", "second", etc. are used to describe various elements and do not intend to limit the positional relationship, timing relationship or importance relationship of these elements. Such terms are only used to distinguish one element from another. In some examples, the first element and the second element may refer to the same instance of the element, and in certain cases, based on the context description, they may also refer to different instances.

[0032] In the description of various examples in the present disclosure, the terms used are only for the purpose of describing specific examples and are not intended to be restrictive. Unless the context clearly indicates otherwise, if the number of elements is not specifically limited, the element may be one or more. In addition, the term "and / or" used in the present disclosure covers any one of the listed items and all possible combinations.

[0033] In the related art, the traditional method of creating an agent is through text input, which requires a clear appeal for the agent, and then continuously optimizes rules and logic to improve the thinking path and answering effect of the agent. This method requires a large amount of information input, and the image and model of the created agent are very one-sided and cannot cover all vertical fields.

[0034] To this end, embodiments of the present disclosure provide an effective information interaction technology for agents based on large language models.

[0035] Embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.

[0036] Figure 1 FIG. shows a schematic diagram of an exemplary system 100 in which the various methods and apparatuses described herein can be implemented according to embodiments of the present disclosure. Referring to Figure 1 , the system 100 includes one or more client devices 101, 102, 103, 104, 105 and 106, a server 120, and one or more communication networks 110 that couple the one or more client devices to the server 120. The client devices 101, 102, 103, 104, 105 and 106 can be configured to execute one or more applications.

[0037] In embodiments of the present disclosure, the server 120 can run one or more services or software applications that enable the execution of information interaction methods for agents based on large language models, multi-agent information interaction methods, and information search methods.

[0038] In certain embodiments, the server 120 can also provide other services or software applications, which can include non-virtual environments and virtual environments. In certain embodiments, these services can be provided as web-based services or cloud services, for example, provided to users of the client devices 101, 102, 103, 104, 105 and / or 106 under a software as a service (SaaS) model.

[0039] In Figure 1 the configuration shown, the server 120 can include one or more components that implement the functions performed by the server 120. These components can include software components, hardware components, or a combination thereof that can be executed by one or more processors. Users operating the client devices 101, 102, 103, 104, 105 and / or 106 can in turn utilize one or more client applications to interact with the server 120 to utilize the services provided by these components. It should be understood that various different system configurations are possible, which can be different from the system 100. Therefore, Figure 1 is an example of a system for implementing the various methods described herein and is not intended to be limiting.

[0040] Users can use client devices 101, 102, 103, 104, 105, and / or 106 to input multimodal information such as text, pictures, audio, and / or video. The client devices can provide an interface that enables the users of the client devices to interact with the client devices. The client devices can also output information to the users via this interface. Although Figure 1 only six client devices are depicted, those skilled in the art will be able to understand that the present disclosure can support any number of client devices.

[0041] Client devices 101, 102, 103, 104, 105, and / or 106 can include various types of computer devices, such as portable handheld devices, general-purpose computers (such as personal computers and laptop computers), workstation computers, wearable devices, smart screen devices, self-service terminal devices, service robots, gaming systems, thin clients, various messaging devices, sensors, or other sensing devices, etc. These computer devices can run various types and versions of software applications and operating systems, such as MICROSOFT Windows, APPLE iOS, UNIX-like operating systems, Linux, or Linux-like operating systems (such as GOOGLE Chrome OS); or include various mobile operating systems, such as MICROSOFT WindowsMobile OS, iOS, Windows Phone, Android. Portable handheld devices can include cellular phones, smartphones, tablets, personal digital assistants (PDAs), etc. Wearable devices can include head-mounted displays (such as smart glasses) and other devices. Gaming systems can include various handheld gaming devices, Internet-enabled gaming devices, etc. The client devices are capable of executing various different applications, such as various Internet-related applications, communication applications (such as email applications), short message service (SMS) applications, and can use various communication protocols.

[0042] Network 110 can be any type of network well-known to those skilled in the art, which can support data communication using any one of a variety of available protocols (including but not limited to TCP / IP, SNA, IPX, etc.). By way of example only, one or more networks 110 can be a local area network (LAN), an Ethernet-based network, token ring, wide area network (WAN), the Internet, virtual network, virtual private network (VPN), intranet, extranet, blockchain network, public switched telephone network (PSTN), infrared network, wireless network (such as Bluetooth, WIFI), and / or any combination of these and / or other networks.

[0043] Server 120 may include one or more general-purpose computers, dedicated server computers (such as PC (Personal Computer) servers, UNIX servers, midrange servers), blade servers, mainframes, server clusters, or any other suitable arrangement and / or combination. Server 120 may include one or more virtual machines running a virtual operating system, or other computing architectures involving virtualization (such as one or more flexible pools of logical storage devices that can be virtualized to maintain virtual storage devices of the server). In various embodiments, server 120 may run one or more services or software applications that provide the functions described below.

[0044] The computing units in server 120 may run one or more operating systems including any of the above operating systems and any commercially available server operating systems. Server 120 may also run any one of a variety of additional server applications and / or middleware applications, including HTTP servers, FTP servers, CGI servers, JAVA servers, database servers, etc.

[0045] In some embodiments, server 120 may include one or more applications to analyze and merge data feeds and / or event updates received from users of client devices 101, 102, 103, 104, 105, and / or 106. Server 120 may also include one or more applications to display data feeds and / or real-time events via one or more display devices of client devices 101, 102, 103, 104, 105, and / or 106.

[0046] In some embodiments, server 120 may be a server of a distributed system, or a server incorporating a blockchain. Server 120 may also be a cloud server, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology. A cloud server is a host product in the cloud computing service system, which addresses the defects of difficult management and weak business scalability existing in traditional physical hosts and virtual private server (VPS, Virtual Private Server) services.

[0047] System 100 may also include one or more databases 130. In some embodiments, these databases can be used to store data and other information. For example, one or more of databases 130 can be used to store information such as audio files and video files. Databases 130 can reside in various locations. For example, the databases used by server 120 can be local to server 120, or can be remote from server 120 and can communicate with server 120 via a network-based or dedicated connection. Databases 130 can be of different types. In some embodiments, the databases used by server 120 can be, for example, relational databases. One or more of these databases can store, update, and retrieve data to and from the databases in response to commands.

[0048] In some embodiments, one or more of databases 130 can also be used by applications to store application data. The databases used by applications can be different types of databases, such as key-value stores, object stores, or conventional stores supported by a file system.

[0049] Figure 1 System 100 can be configured and operated in various ways to enable the application of the various methods and apparatuses described according to the present disclosure.

[0050] Aspects of an information interaction method of an agent based on a large language model according to embodiments of the present disclosure will be described in detail below.

[0051] Figure 2 A flowchart of an information interaction method 200 of an agent based on a large language model according to an embodiment of the present disclosure is shown.

[0052] As Figure 2 shown, method 200 includes step S201, step S202, step S203, step S204, and step S205.

[0053] In step S201, multimodal information associated with a target object of interest to a first user is obtained.

[0054] In the example, the multimodal information can involve different forms or types of information, such as image information, video information, audio information, text information, etc. The first user can refer to the user who currently intends to create an agent, and the target object of interest can refer to the thing or topic that the user is concerned about. For example, the user may be interested in a specific product, object, person, or other content. After determining the target object, the user can take a photo or video of the target object (such as a 30-second video) through a camera or other image capture device, or record audio related to the target object through an audio recording device, or also collect text information related to the target object through a text editing device, so as to obtain multimodal information associated with the target object.

[0055] In step S202, based on the multimodal information, a first agent for the target object is created.

[0056] In the example, according to the obtained multimodal information of the target object, such as a video information of the target object being "Bronze Divine Tree" obtained in the application scenario of a historical museum, a first agent for the "Bronze Divine Tree" can be created. This agent can be, for example, a smart assistant specifically serving the first user for the "Bronze Divine Tree", and can answer various questions about the "Bronze Divine Tree", such as the construction time and historical background of the "Bronze Divine Tree".

[0057] In step S203, at least one second agent associated with the first agent is determined. The at least one second agent has had information interaction with at least one of the first user or the second user.

[0058] In the example, the second agent being associated with the first agent can mean that there is a direct or indirect connection or influence between the second agent and the first agent. For example, in the historical museum scenario, after creating an agent for the "Bronze Divine Tree" based on the obtained multimodal information of the "Bronze Divine Tree", the multimodal information of the "Bronze Divine Tree" can be analyzed, and then other agents associated with the agent of the "Bronze Divine Tree" can be determined, such as agents for the "Bronze Mask" and "Gold Scepter" located in the same museum. These associated agents, that is, the second agents, have had conversation interactions with the current first user or other users other than the first user (i.e., the second user). For example, the agent for the "Bronze Mask" may have interacted with other users, so it has a certain understanding and knowledge of the "Bronze Mask". Similarly, the agent for the "Gold Scepter" may have interacted with the current first user before, so it has a certain understanding and knowledge of the "Gold Scepter". Such understanding and knowledge can be memorized and stored through the interaction information, and thus can become the basis for the interaction between multiple agents, enabling more rich and accurate services to be provided through multi-agent interaction.

[0059] In step S204, context connection is performed between the first agent and the at least one second agent.

[0060] In an example, context connection may involve the agent being able to understand and maintain the context of a conversation, enabling the conversation to smoothly transition from one topic to another, or remaining coherent and consistent when switching between different agents. Additionally, context connection may also involve emotional continuity, for example, the agent being able to maintain the emotional tone in the conversation, such as remaining humorous in a relaxed conversation and being serious in a serious topic.

[0061] In an example, as mentioned above, since the second agent has interacted with the current first user or another user, i.e., the second user, relevant historical interaction information has been accumulated. This historical interaction information can be sent to the first agent after being integrated, for example, as the basis for the first agent to converse with the current first user. For example, the historical interaction information generated by the "bronze mask" agent in the example of step S203 and another user, i.e., the second user, is sent to the "bronze divine tree" agent currently conversing with the first user, enabling the "bronze divine tree" agent to also provide the first user with relevant information about the "bronze mask". This can not only enrich the knowledge reserve of the agent but also enhance the user experience.

[0062] In step S205, information interaction is performed between the first agent and the first user.

[0063] In an example, the current first user can converse with the first agent created based on the multimodal information associated with the target object provided by the user to obtain relevant information and services. In this process, the first agent can provide a more comprehensive and coherent answer by receiving the sharing of historical interaction information from other relevant agents, i.e., the second agent.

[0064] Therefore, in the information interaction method 200 of the agent based on the large language model according to the embodiments of the present disclosure, first, a first agent is created based on the obtained multimodal information, then other agents associated with the first agent, i.e., the second agent, are determined. Furthermore, by using the historical interaction information between the second agent and the current first user or other second users except the first user, through context connection, the knowledge base of the first agent is extended, enabling its cognition and understanding to extend to the associated second agent. Thus, the cognition and understanding of the first agent are correspondingly extended, and it can provide a more rich and comprehensive answer for the current first user, thereby ensuring the coherence and consistency of information interaction. This method 200 can not only enrich the means of creating agents but also provide more comprehensive and personalized services, thereby significantly enhancing the user experience.

[0065] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0066] In some embodiments, as combined with Figure 2 Determining at least one second agent associated with the first agent as shown in step S203 may include: determining at least one associated object associated with the target object based on the multi-modal information corresponding to the target object; and determining at least one second agent respectively for the at least one associated object.

[0067] In the example, in the historical museum scenario, the target object is, for example, the "Bronze Standing Human Figure of the Shang Dynasty", so the multi-modal information of the "Bronze Standing Human Figure of the Shang Dynasty" can be obtained. Furthermore, the associated objects associated with the "Bronze Standing Human Figure of the Shang Dynasty" within the same historical museum can be determined, such as cultural relics like the "Bronze Tiger of the Shang Dynasty". These associated objects can be, for example, other bronzes manufactured during the same historical period, or other cultural relics near the exhibition location of the "Bronze Standing Human Figure of the Shang Dynasty". For each associated object, an agent for that associated object can be determined accordingly. For example, an agent for the associated object "Bronze Tiger of the Shang Dynasty" of the above-mentioned target object "Bronze Standing Human Figure of the Shang Dynasty" can be determined.

[0068] Therefore, by virtue of the inherent correlation between the multi-modal information of multiple objects, it is possible to start from the target object and lead to the associated objects related to it, thereby facilitating indirectly determining which or which agents are the second agents associated with the first agent.

[0069] Figure 3 Shows a schematic diagram of the process of determining a second agent associated with a first agent according to an embodiment of the present disclosure.

[0070] As Figure 3As shown, still taking the application scenario of the aforementioned historical museum as an example, the first agent can be created based on a target object 301 such as the "Bronze Standing Statue of a Human Figure". In this case, according to the multi-modal information corresponding to the target object 301, the "Bronze Standing Statue of a Human Figure", three associated objects 311, 312, and 313 associated with the target object 301, the "Bronze Standing Statue of a Human Figure", can be determined. These three associated objects can be, for example, the "Bronze Tiger", the "Bronze Mask with Vertical Pupils", and the "Gold Mask" that are the same as or close to the manufacturing time of the "Bronze Standing Statue of a Human Figure". The agents of these associated objects 311, 312, and 313 may have been created by the current user or other users, and then the agent 321 of the associated object 311, the "Bronze Tiger", the agent 322 of the associated object 312, the "Bronze Mask with Vertical Pupils", and the agent 323 of the associated object 313, the "Gold Mask", can be determined, that is, the second agents 321, 322, and 323.

[0071] It can be understood that Figure 3 only three associated objects are taken as examples for illustration. However, the embodiments of the present disclosure are not limited thereto, and the number of associated objects can be one or more.

[0072] In some embodiments, as combined with Figure 2 the step S204 shown, the context connection between the first agent and at least one second agent may include: sharing the historical context information generated by the information interaction between the at least one second agent and at least one of the first user or the second user with the first agent.

[0073] In the example, the second agent has had a conversation with the current user or other users. The historical context information of these conversations can include, for example, questions raised by the current user or other users, answers of the agent, etc. Therefore, sharing this historical context information with the first agent can enable the first agent to refer to or quote this information in the conversation with the current user.

[0074] For example, the first agent used for the "Bronze Standing Statue of a Human Figure" is having a conversation with the current user, and the second agent associated with this first agent is the agent used for the "Bronze Tiger". Among them, the second agent used for the "Bronze Tiger" has had a conversation with other users, then the generated historical context information about the "Bronze Tiger" can be shared with the first agent used for the "Bronze Standing Statue of a Human Figure". In this way, when the current user raises a relevant question about the "Bronze Tiger" or mentions content related to the "Bronze Tiger" during the conversation with the first agent used for the "Bronze Standing Statue of a Human Figure", the first agent used for the "Bronze Standing Statue of a Human Figure" can also give a corresponding answer or make a corresponding feedback.

[0075] Therefore, by sharing historical context information among multiple agents, the agents can reference a wide range of knowledge in the conversation with the user, provide more comprehensive and coherent answers, thereby significantly enhancing the user experience and making the interaction process more natural and efficient.

[0076] In some embodiments, in response to at least one second agent having interacted with a first user, context connection of the first agent with the at least one second agent, such as in accordance with Figure 2 step S204 shown, may further include: determining whether the historical context information contains a question associated with the current target object, where the at least one second agent has not provided an answer to the question to the first user; and in response to determining that there is a question, determining an answer to the question based on multi-modal information associated with the target object.

[0077] In an example, if the second agent has interacted with the current user, the generated historical context information can be shared with the first agent, and the first agent can analyze the historical context information and extract a question associated with the target object of interest provided by the current user but not answered by the second agent. Then, relevant content can be obtained from the multi-modal information of the target object to answer the question regarding the target object.

[0078] For example, in the application scenario of shopping, the user currently takes an image or video of a "raincoat" of a certain brand and thus creates a first agent for the "raincoat" to interact with it. At this time, assuming that it can be determined that the user had a conversation interaction with a second agent for the "umbrella" of this brand in the previous time period, and in this conversation, the user asked a question about the "material" of the "raincoat", but the second agent for the "umbrella" could not provide an answer to this question to the user at that time. Therefore, when the user currently converses with the first agent for the "raincoat", the historical context information generated by the user's conversation with the second agent for the "umbrella" in the previous time period can be shared with the current first agent for the "raincoat", and from this historical context information, it can be determined that the user asked a question about the "material" of the "raincoat" but was not answered by the second agent for the "umbrella". Therefore, currently, based on the image or video of the "raincoat" taken by the user, the "material" of the "raincoat" can be determined and the answer can be provided to the user by the current first agent for the "raincoat".

[0079] Therefore, by analyzing the historical interaction records between the second agent and the current user, determining the questions related to the current target object that have not been answered yet, and then providing answers to these questions based on the multimodal information of the target object, it can be ensured that the questions raised by the user will not be ignored, enabling the "pending" questions to be answered in a timely manner. Thereby, the integrity of question answering can be improved, and further the user experience and satisfaction can be enhanced.

[0080] In some embodiments, before performing context connection on the first agent and at least one second agent as shown in Figure 2 step S204, it may further include: determining the time when at least one second agent interacts with information of at least one of the first user or the second user; retrieving updated information associated with at least one associated object that appears after this time; and modifying the historical context information based on the updated information.

[0081] In an example, before the second agent shares historical context information with the first agent, additional information accuracy verification can be performed to ensure the timeliness and accuracy of the content in the historical context information. For this purpose, for the second agent used for the associated object, when it has previously had a conversation interaction with any user (such as the current user or other users), the latest information associated with the associated object can be retrieved on the network. These latest information can come from various information channels. For example, in the application scenario of a historical museum, it can come from new archaeological discoveries, academic research results, news reports, etc. Based on the retrieved latest information, the historical context information generated by the conversation between the second agent and the user can be updated and corrected, and then the modified historical context information can be shared with the first agent.

[0082] Therefore, this embodiment provides an agent data refresh mechanism. By retrieving and updating the historical context information of the second agent, the timeliness and accuracy of the historical context information can be ensured. Furthermore, the extended cognition and understanding of the first agent can be ensured to be timely and accurate. Thereby, when the first agent converses with the current user, it can provide the latest and comprehensive answers, avoiding the interference of outdated information, and thus improving the quality of answering questions.

[0083] Figure 4 The figure shows a schematic diagram of modifying historical context information based on the retrieved updated information according to an embodiment of the present disclosure.

[0084] As Figure 4As shown, it is assumed that the first user is currently interacting with the first agent for the target object at time t3, and the second agent associated with the first agent interacted with the first user at time t1 before time t3. At this time, the historical context information 401 generated by the interaction between the second agent and the first user at time t1 can be shared with the first agent. In the historical context information 401, the first user asks the second agent "What is the carving time of Cultural Relic A?", and the second agent answers "Its carving time is XX years". In this embodiment, before the sharing operation, it is possible to retrieve on a news website or the like whether there is updated information about the historical context information 401 after time t1. Assuming that updated information 402 about the actual carving time of Cultural Relic A being YY years is published on a certain news website at time t2 after time t1, then based on the updated information 402, the historical context information 401 can be modified to generate modified historical context information 401', where the answer to the carving time of Cultural Relic A is modified to YY years.

[0085] In some embodiments, as combined with Figure 2 As shown, step S202 of creating a first agent for a target object based on multimodal information may include: determining first label information and second label information based on multimodal information, where the first label information includes a subject guess word for characterizing the target object, and the second label information includes user data of the first user; and generating a first agent based on the first label information and the second label information.

[0086] In the example, the subject guess word included in the first label information can be used to indicate what kind of entity the target object belongs to. For example, the subject guess word "Bronze Standing Statue of the Shang Dynasty" can be extracted from a video about the Bronze Standing Statue of the Shang Dynasty obtained. Therefore, the first label information may include key information for depicting the target object, and can also be referred to as main label information. The user data included in the second label information may be personalized data related to the user, such as user search behavior, such as user cognition, user preferences, and user habits. Therefore, the second label information can also be referred to as secondary label information.

[0087] Based on the acquired multi-modal information associated with the target object, such as images, videos, audio, etc., the determination of the first tag information and the second tag information can be carried out synchronously. Taking the application scenario of the aforementioned historical museum as an example, the target object can be, for example, a certain cultural relic. On the one hand, the main intention recognition can be carried out based on the multi-modal information of the cultural relic to determine whether the main body of the cultural relic is recognized. If the main body of the cultural relic is not recognized, it may be due to unclear videos taken by the user or other reasons. In this case, the user can be prompted to take pictures again, or the traditional search results of "searching by image" can be directly provided to the user without creating an agent. If the main body of the cultural relic is recognized, the multi-modal language model can be used to refine the topic guessing words. For example, the topic guessing word "Shang Dynasty Bronze Standing Statue of a Man" can be refined from the video taken by the user. On the other hand, user data collection can be carried out. For example, in addition to the aforementioned user search behavior, the positioning features and historical records related to the user can also be collected. The positioning features can include, for example, other museums near the current historical museum, or the historical and cultural features of the location of the current historical museum, such as whether it is the ancient capital of a certain dynasty, etc. The historical records can include, for example, the tourism information and scenic spot reservation information of the location of the current historical museum, etc.

[0088] In the example, when generating an agent, a creation link can be added to pour the first tag information and the second tag information into the platform library, and generate the desired agent in the necessary format for agent creation. For example, the creation process can involve the determination of the basic persona, such as the main body name, personality, image, and language pack, etc., and can also involve the role setting, target task, and detailed background of the agent. In addition, it can also involve the ability invocation of the agent, such as the ability of full-network search, multi-round dialogue, and historical records.

[0089] Therefore, by extracting the topic guessing words of the target object and the personalized data of the user to generate an agent specifically for the target object, the generated agent can provide accurate and personalized services for the user, significantly improving the user experience.

[0090] Figure 5 The schematic diagram of creating the first agent according to an embodiment of the present disclosure is shown.

[0091] As Figure 5As shown, based on the multimodal information 501 associated with the target object, the first tag information 510 and the second tag information 520 associated with the target object can be determined. The first tag information 510 may include a subject guess 511 for the target object. The second tag information 520 may include the user's behavior portrait 521 and the user's location information 522. For example, the user's behavior portrait 521 may include the user search behavior as described above, and the user's location information 522 may include location features and historical records related to the user. The subject guess 511 of the target object, the user's behavior portrait 521, and the user's location information 522 can be integrated to create an agent 530 specifically for the target object.

[0092] In some embodiments, as combined with Figure 2 shown in step S202 of creating a first agent for the target object based on multimodal information may further include: before generating the first agent, determining whether the first agent already exists; and in response to determining that the first agent already exists, updating the second tag information determined when the first agent was previously generated to the currently determined second tag information.

[0093] In the example, after extracting the subject guess representing the target object from the multimodal information, it can be checked in the agent database whether there is an agent for the same target object according to the subject guess. If there is no agent for the same target object, a new agent for the target object can be created; if there is an agent for the same target object, the personalized data of the previous user obtained when the existing agent was previously generated can be updated to the personalized data of the current user to meet the needs of the current user. For example, currently, the subject guess "Bronze Standing Statue of a Shang Dynasty Person" is extracted from the multimodal information, and it is searched in the agent database that there is already an agent for the "Bronze Standing Statue of a Shang Dynasty Person", then the data of the previous user obtained during the previous generation process of the "Bronze Standing Statue of a Shang Dynasty Person" agent can be updated to the data of the current user, so as to obtain an agent specifically for the current user based on the updated user data.

[0094] Therefore, in this way, not only can the computing resources be saved by directly processing the recalled agent and avoid repeated creation of agents, but also it can be ensured that the information and services provided by the agent are adapted to the current user's needs.

[0095] In some embodiments, as combined with Figure 2 the information interaction method 200 of the agent based on the large language model as described may further include: sharing the current context information generated by the interaction between the first agent and the first user with at least one second agent; and interacting with the first user via the at least one second agent.

[0096] In the example, when the current user interacts with the first agent, context information may be generated, such as questions asked by the current user, answers from the agent, feedback from the current user, etc. In this case, the context information generated by the current conversation may be shared with the second agents associated with the first agent, so that these associated agents can also learn the real-time interactive content for the current user. In this sharing manner, when the current user subsequently interacts with these associated second agents, the second agents can use the context information to give more accurate and comprehensive answers.

[0097] For example, the current user is interacting with the first agent for "Shang Bronze Standing Figure". After the interaction is completed, a piece of context information can be generated, and the context information contains relevant information about "Shang Bronze Standing Figure". The context information can then be shared with the second agent associated with the first agent, for example, with the second agent for "Shang Bronze Tiger". In the subsequent process of information interaction with the second agent of "Shang Bronze Tiger", the user may also ask questions about "Shang Bronze Standing Figure", such as "What is the difference in construction between Shang Bronze Standing Figure and Shang Bronze Tiger?" At this time, the second agent for "Shang Bronze Tiger" can combine the shared context information for "Shang Bronze Standing Figure" for analysis to provide the user with an accurate answer.

[0098] Therefore, by sharing the real-time contextual information generated by the interaction between the first agent and the user with the second agent associated with the first agent, the cognition and understanding of each other between the first agent and the second agent can be connected, so that these associated agents can have a continuous and relevant dialogue with the user based on the real-time interactive information, thereby providing users with more accurate and coherent question-and-answer services by ensuring information sharing and synchronization between different agents.

[0099] According to an embodiment of the present disclosure, a multi-agent information interaction method is also provided, comprising: creating a chat group including a user and a plurality of agents, wherein the user and the plurality of agents interact with each other based on the information interaction method as described above (e.g., in combination with Figure 2 The information interaction method 200) performs information interaction.

[0100] In the example, multi-agent interaction can be performed by creating a chat group, so that the group includes a user and multiple associated agents. The group can simulate a multi-person chat environment, where users can ask questions or express needs in the group, and multiple agents in the group can provide corresponding answers based on their respective areas of expertise. The various agents can share historical context information generated by the dialogue with the user to ensure the accuracy and coherence of the answers and avoid repeated or contradictory answers.

[0101] In the example, the user can indicate in the chat group to have a conversation only with a certain or some specific agents among the multiple agents. The conversation can be one-round or multi-round.

[0102] Therefore, by creating a chat group for the user and multiple agents, the user can have an efficient conversation with multiple agents. This method can not only answer the user's questions accurately and completely, but also enhance the user experience and make the interaction process more natural.

[0103] Figure 6 A schematic diagram showing the creation of a chat group for the user and multiple agents according to an embodiment of the present disclosure is shown.

[0104] As Figure 6 shown, in the chat group 600 created for the user and multiple agents, the user 610 can ask a question 611 "What cultural relics are there in this history museum?". The agent 620 represented by the triangle symbol in the figure can be an agent for the bronze sacred tree, so it can provide the user with knowledge about the bronze sacred tree in the history museum, such as providing a reply answer 621 "The bronze sacred tree, the sacred tree consists of a base, a tree, and a dragon, and is cast by the segmented casting method, with a total height of 396 cm". Similarly, the agent 630 represented by the rectangular symbol in the figure can be an agent for the bronze vertical-eyed mask, so it can provide the user with knowledge about the bronze vertical-eyed mask in the history museum, such as providing a reply answer 631 "The bronze vertical-eyed mask is the most peculiar in shape among the many bronze masks unearthed locally". Next, the user 610 may continue to ask a question 612 "Are there any similar to humans?", and then the agent 640 for the bronze standing figure represented by the circular symbol in the figure can provide a reply answer 641 "The bronze standing figure, the figure is cast by the segmented casting method and inlaid, the clothing pattern is complex and delicate, the hands are held in a ring and hollow, and it stands barefoot on a square monster seat". In this way, the user 610 can quickly obtain a detailed answer about the cultural relics in the history museum through information interaction with multiple agents in a chat group.

[0105] Figure 7 A flowchart of an information search method 700 according to an embodiment of the present disclosure is shown.

[0106] As Figure 7 shown, the method 700 includes step S701 and step S702.

[0107] In step S701, receive multimodal information provided by the user and associated with the target object of interest.

[0108] In the example, the target object that the user is interested in can refer to a specific thing or entity that the user pays special attention to in a specific scenario. In the application scenario of a historical museum, the multimodal information of the target object can be, for example, one or more photos or videos of cultural relics that the user uploads, a voice explanation provided, or detailed descriptive text input. Key information in the multimodal information can be extracted through image recognition technology, speech recognition technology, etc., and this key information can be used to indicate what kind of thing or entity the target object that the user is interested in is.

[0109] In step S702, based on the information interaction method described above (for example, combined with Figure 2 the information interaction method 200)) to interact with the user to provide the user with search results for the target object.

[0110] In the example, according to the multimodal information of the target object provided by the user, an agent for the target object can be created, which can answer questions about the target object, thereby providing search results for the user. For example, in the scenario of taking a photo and recognizing the object, the user takes and uploads a picture of the "Bronze Divine Tree". Accordingly, an agent about the "Bronze Divine Tree" can be created based on this picture. This agent can introduce the object in the picture to the user in detail, such as the name, historical background, unearthed time, etc. of the "Bronze Divine Tree". At the same time, this agent can also be shared with the historical context information generated by the conversation between the associated agent and any user to enrich its knowledge base, so as to provide more comprehensive search results for the user.

[0111] Therefore, by searching in the way of creating an agent based on the multimodal information provided by the user, it can not only provide the user with a more natural and efficient real-time search experience, but also increase the real-time feedback and emotional interaction of the user during the object recognition search process.

[0112] According to an embodiment of the present disclosure, there is also provided an information interaction device for an agent based on a large language model.

[0113] Figure 8 The structural block diagram of an information interaction device 800 for an agent based on a large language model according to an embodiment of the present disclosure is shown.

[0114] As Figure 8As shown, the apparatus 800 includes an information acquisition module 801, an agent creation module 802, an associated agent determination module 803, a context connection module 804, and a first information interaction module 805. The information acquisition module 801 is configured to acquire multimodal information associated with a target object of interest to a first user. The agent creation module 802 is configured to create a first agent for the target object based on the multimodal information. The associated agent determination module 803 is configured to determine at least one second agent associated with the first agent, where at least one second agent has interacted with at least one of the first user or a second user. The context connection module 804 is configured to connect the context of the first agent and the at least one second agent. The first information interaction module 805 is configured to interact with the first user via the first agent.

[0115] The operations of the above-mentioned information acquisition module 801, agent creation module 802, associated agent determination module 803, context connection module 804, and first information interaction module 805 can respectively correspond to the operations of steps S201, S202, S203, S204, and S205 as Figure 2 shown. Therefore, the details of each aspect will not be elaborated here.

[0116] Figure 9 The structural block diagram of an information interaction apparatus 900 for an agent based on a large language model according to another embodiment of the present disclosure is shown.

[0117] As Figure 9 shown, the apparatus 900 may include an information acquisition module 901, an agent creation module 902, an associated agent determination module 903, a context connection module 904, and a first information interaction module 905. The operations of the above-mentioned modules may be the same as those of the information acquisition module 801, agent creation module 802, associated agent determination module 803, context connection module 804, and first information interaction module 805 as Figure 8 shown. In addition, the above-mentioned modules may further include further sub-modules.

[0118] In some embodiments, the associated agent determination module 903 may include an associated object determination module 9031 and a result determination module 9032. The associated object determination module 9031 may be configured to determine at least one associated object associated with the target object based on the multimodal information corresponding to the target object. The result determination module 9032 may be configured to determine at least one second agent respectively for the at least one associated object.

[0119] In some embodiments, the context connection module 904 may include a first information sharing module 9041. The first information sharing module 9041 may be configured to share historical context information generated by information interaction between at least one second agent and at least one of the first user or the second user with the first agent.

[0120] In some embodiments, the context connection module 904 may further include a question determination module 9042 and an answer determination module 9043. The question determination module 9042 may be configured to determine whether the historical context information contains a question associated with the current target object, where at least one second agent has not provided an answer to the first user for the question. The answer determination module 9043 may be configured to, in response to determining that there is a question, determine an answer to the question based on multimodal information associated with the target object.

[0121] In some embodiments, the apparatus 900 may further include the following modules arranged before the context connection module 904: an interaction time determination module 906, an updated information retrieval module 907, and a context information modification module 908. The interaction time determination module 906 may be configured to determine the time when at least one second agent interacts with at least one of the first user or the second user. The updated information retrieval module 907 may be configured to retrieve updated information associated with at least one associated object that appears after this time. The context information modification module 908 may be configured to modify the historical context information based on the updated information.

[0122] In some embodiments, the agent creation module 902 may include a tag information determination module 9021 and an agent generation module 9022. The tag information determination module 9021 may be configured to determine first tag information and second tag information based on multimodal information, where the first tag information includes a guessed word for the subject characterizing the target object, and the second tag information includes user data of the first user. The agent generation module 9022 may be configured to generate a first agent based on the first tag information and the second tag information.

[0123] In some embodiments, the agent creation module 902 may further include a recall determination module 9023 and a tag information update module 9024. The recall determination module 9023 may be configured to determine whether a first agent already exists before generating the first agent. The tag information update module 9024 may be configured to, in response to determining that a first agent already exists, update the second tag information determined during the previous generation of the first agent to the currently determined second tag information.

[0124] In some embodiments, the apparatus 900 may further include a second information sharing module 909 and a second information interaction module 910. The second information sharing module 909 may be configured to share the current context information generated by the information interaction between the first agent and the first user with at least one second agent. The second information interaction module 910 may be configured to interact with the first user via at least one second agent.

[0125] According to an embodiment of the present disclosure, there is also provided a multi-agent information interaction apparatus.

[0126] Figure 10 FIG. shows a structural block diagram of a multi-agent information interaction apparatus 1000 according to an embodiment of the present disclosure.

[0127] As Figure 10 shown, the apparatus 1000 includes a chat group creation module 1001. The chat group creation module 1001 may be configured to create a chat group including a user and a plurality of agents, wherein the user and the plurality of agents perform information interaction based on the information interaction apparatus as described above (for example, in combination with Figure 8 the apparatus 800 described above or Figure 9 the apparatus 900 described above).

[0128] According to an embodiment of the present disclosure, there is also provided an information search apparatus.

[0129] Figure 11 FIG. shows a structural block diagram of an information search apparatus 1100 according to an embodiment of the present disclosure.

[0130] As Figure 11 shown, the apparatus 1100 includes an information receiving module 1101 and a search result providing module 1102. The information receiving module 1101 may be configured to receive multi-modal information provided by the user and associated with a target object of interest. The search result providing module 1102 may be configured to interact with the user based on the information interaction apparatus as described above (for example, in combination with Figure 8 the apparatus 800 described above or Figure 9 the apparatus 900 described above) to provide the user with search results for the target object.

[0131] According to an embodiment of the present disclosure, there is also provided an electronic device, a readable storage medium, and a computer program product.

[0132] According to an embodiment of the present disclosure, there is also provided an electronic device, including at least one processor; and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method as described above.

[0133] According to an embodiment of the present disclosure, there is also provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the method described above.

[0134] According to an embodiment of the present disclosure, there is also provided a computer program product including a computer program, wherein the computer program, when executed by a processor, implements the method described above.

[0135] Referring Figure 12 , a block diagram of an electronic device 1200 that can be a server or a client of the present disclosure will now be described. It is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device is intended to represent various forms of digital electronic computer devices, such as, for example, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0136] As Figure 12 shown, the electronic device 1200 includes a computing unit 1201, which can execute various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1202 or a computer program loaded from a storage unit 1208 into a random access memory (RAM) 1203. In the RAM 1203, various programs and data required for the operation of the electronic device 1200 can also be stored. The computing unit 1201, the ROM 1202, and the RAM 1203 are connected to each other via a bus 1204. An input / output (I / O) interface 1205 is also connected to the bus 1204.

[0137] Multiple components in the electronic device 1200 are connected to the I / O interface 1205, including: an input unit 1206, an output unit 1207, a storage unit 1208, and a communication unit 1209. The input unit 1206 can be any type of device capable of inputting information into the electronic device 1200. The input unit 1206 can receive input digital or character information, and generate key signal inputs related to the user settings and / or function controls of the electronic device, and can include but are not limited to a mouse, a keyboard, a touch screen, a trackpad, a trackball, a joystick, a microphone, and / or a remote control. The output unit 1207 can be any type of device capable of presenting information, and can include but are not limited to a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The storage unit 1208 can include but are not limited to a magnetic disk and an optical disk. The communication unit 1209 allows the electronic device 1200 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks, and can include but are not limited to a modem, a network card, an infrared communication device, a wireless communication transceiver, and / or a chipset, such as a Bluetooth device, an 802.11 device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.

[0138] The computing unit 1201 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1201 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1201 executes the various methods and processes described above. For example, in some embodiments, the method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 1208. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 1200 via the ROM 1202 and / or the communication unit 1209. When the computer program is loaded into the RAM 1203 and executed by the computing unit 1201, one or more steps of the method described above can be executed. Alternatively, in other embodiments, the computing unit 1201 can be configured to execute the method in any other suitable way (e.g., by means of firmware).

[0139] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.

[0140] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code can be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine or entirely on the remote machine or server.

[0141] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0142] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).

[0143] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), the Internet, and blockchain networks.

[0144] A computer system can include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The client - server relationship is created by computer programs running on the respective computers and having a client - server relationship with each other. The server can be a cloud server, or a server of a distributed system, or a server incorporating blockchain.

[0145] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this is not limited herein.

[0146] Although embodiments or examples of the present disclosure have been described with reference to the accompanying drawings, it should be understood that the above methods, systems, and devices are merely exemplary embodiments or examples, and the scope of the present invention is not limited by these embodiments or examples, but is only defined by the authorized claims and their equivalent scope. Various elements in the embodiments or examples may be omitted or replaced by their equivalent elements. In addition, the steps may be executed in an order different from that described in the present disclosure. Further, various elements in the embodiments or examples may be combined in various ways. Importantly, with the evolution of technology, many of the elements described herein may be replaced by equivalent elements that emerge after the present disclosure.

Claims

1. An information interaction method of an intelligent agent based on a large language model, comprising: Acquire multimodal information associated with a target object of interest to the first user; Based on the multimodal information, creating a first agent for the target object; Determining at least one second agent associated with the first agent, wherein the at least one second agent has interacted with at least one of the first user or the second user; Contextually connecting the first agent with the at least one second agent; and Information is exchanged with the first user via the first agent.

2. The method according to claim 1, wherein: The determining of at least one second agent associated with the first agent comprises: Determining at least one associated object associated with the target object based on the multimodal information corresponding to the target object; and The at least one second agent is determined for each of the at least one associated objects.

3. The method according to claim 1 or 2, wherein: The contextual connection between the first agent and the at least one second agent includes: The historical context information generated by the information interaction between the at least one second agent and the first user or at least one of the second users is shared with the first agent.

4. The method according to claim 3, wherein: In response to the at least one second agent having interacted with the first user, the contextual connection between the first agent and the at least one second agent further includes: determining whether the historical context information includes a question associated with the current target object, wherein the at least one second agent has not provided an answer to the first user for the question; and In response to determining that the question exists, the answer to the question is determined based on the multimodal information associated with the target object.

5. The method according to any one of claims 2 to 4, wherein: Before contextually connecting the first agent with the at least one second agent, the method further includes: Determining a time when the at least one second agent interacts with at least one of the first user or the second user; retrieving update information associated with the at least one associated object that occurs after the time; and Based on the update information, the historical context information is modified.

6. The method according to any one of claims 1 to 5, wherein: The step of creating a first agent for the target object based on the multimodal information comprises: Based on the multimodal information, determining first label information and second label information, wherein the first label information includes a subject guess word for characterizing the target object, and the second label information includes user data of the first user; and Based on the first tag information and the second tag information, the first agent is generated.

7. The method according to claim 6, wherein: The method of creating a first agent for the target object based on the multimodal information further includes: Before generating the first agent, determining whether the first agent already exists; and In response to determining that the first agent already exists, the second tag information determined when the first agent was last generated is updated to the currently determined second tag information.

8. The method according to any one of claims 1 to 7, wherein: The method further comprises: Sharing the current context information generated by the information interaction between the first agent and the first user with the at least one second agent; and Information is exchanged with the first user via the at least one second agent.

9. A multi-agent information interaction method, comprising: A chat group including a user and a plurality of agents is created, wherein the user and the plurality of agents perform information interaction based on the method according to any one of claims 1 to 8.

10. An information search method, comprising: receiving multimodal information associated with a target object of interest provided by a user; as well as Information interaction is performed with the user based on the method according to any one of claims 1 to 8 to provide the user with search results for the target object.

11. An information interaction device of an intelligent agent based on a large language model, comprising: an information acquisition module, configured to acquire multimodal information associated with a target object of interest to the first user; An agent creation module, configured to create a first agent for the target object based on the multimodal information; an associated agent determination module, configured to determine at least one second agent associated with the first agent, wherein the at least one second agent has interacted with at least one of the first user or the second user; a context connection module, configured to perform context connection between the first agent and the at least one second agent; and The first information interaction module is configured to perform information interaction with the first user via the first agent.

12. The device according to claim 11, wherein The associated agent determination module includes: an associated object determining module, configured to determine at least one associated object associated with the target object based on the multimodal information corresponding to the target object; and The result determination module is configured to determine the at least one second agent respectively used for the at least one associated object.

13. The device according to claim 11 or 12, wherein: The context connection module includes: The first information sharing module is configured to share historical context information generated by information interaction between the at least one second agent and the first user or at least one of the second users with the first agent.

14. The device according to claim 13, wherein: In response to the at least one second agent having interacted with the first user, the context connection module further includes: a question determination module configured to determine whether the historical context information contains a question associated with the current target object, wherein the at least one second agent has not provided an answer to the first user for the question; and The answer determination module is configured to determine the answer to the question based on the multimodal information associated with the target object in response to determining that the question exists.

15. The device according to any one of claims 12 to 14, wherein: The device also includes the following modules arranged before the context connection module: An interaction time determination module, configured to determine a time for information interaction between the at least one second agent and at least one of the first user or the second user; an update information retrieval module, configured to retrieve update information associated with the at least one associated object that occurs after the time; as well as The context information modification module is configured to modify the historical context information based on the update information.

16. The device according to any one of claims 11 to 15, wherein: The agent creation module includes: a tag information determination module, configured to determine first tag information and second tag information based on the multimodal information, wherein the first tag information includes a subject guess word for characterizing the target object, and the second tag information includes user data of the first user; and The agent generation module is configured to generate the first agent based on the first tag information and the second tag information.

17. The device according to claim 16, wherein: The agent creation module also includes: a recall determination module, configured to determine whether the first agent already exists before generating the first agent; and The label information updating module is configured to update the second label information determined when the first agent was generated last time to the currently determined second label information in response to determining that the first agent already exists.

18. The device according to any one of claims 11 to 17, wherein: The device also includes: A second information sharing module is configured to share the current context information generated by the information interaction between the first agent and the first user with the at least one second agent; and The second information interaction module is configured to interact with the first user via the at least one second agent.

19. A multi-agent information interaction device, comprising: A chat group creation module is configured to create a chat group including a user and a plurality of agents, wherein the user and the plurality of agents interact with each other based on the device according to any one of claims 11 to 18.

20. An information search device, comprising: An information receiving module, configured to receive multimodal information associated with a target object of interest provided by a user; as well as A search result providing module is configured to perform information interaction with the user based on the apparatus according to any one of claims 11 to 18 to provide the user with search results for the target object.

21. An electronic device, comprising: at least one processor; as well as a memory communicatively coupled to the at least one processor; The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1-10.

22. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to make a computer execute the method according to any one of claims 1-10.

23. A computer program product comprising a computer program, wherein: The computer program implements the method according to any one of claims 1 to 10 when executed by a processor.