Immersive digital human interaction method and related device based on multi-agent collaboration

Through the multi-agent collaboration, the candidate digital person is determined using the user's real-time state information and initiates interaction requests, and the target digital person is controlled to perform real-time interaction behavior, solving the problem of lack of anthropomorphism in the interaction mode of the generative large language model, and achieving higher immersion and user acceptance.

CN119558341BActive Publication Date: 2025-08-26BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411621038.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-13
Publication Date
2025-08-26
Estimated Expiration
2044-11-13

AI Technical Summary

Technical Problem

In the prior art, the interaction mode of the generative large language model lacks anthropomorphic feeling, and it is difficult to provide a highly immersive interactive experience with users, and the user acceptance is low.

Method used

The multi-agent collaboration method is adopted to determine the candidate digital person through the main agent, initiate attempted interaction requests based on the user's real-time status information, and control the target digital person to present matching real-time interaction behaviors, including content and posture feedback, to enhance the anthropomorphism and immersion of the interaction.

Benefits of technology

It improves the acceptance of user interaction with agents, provides a higher degree of immersive interactive experience, and enhances the anthropomorphicity and real-time matching of the interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119558341B_ABST
    Figure CN119558341B_ABST
Patent Text Reader

Abstract

The present disclosure provides an immersive digital human interaction method and related devices based on multi-agent collaboration, involving the fields of generative large language models, agents, digital humans, and other artificial intelligence technologies. The method includes: determining multiple candidate digital humans based on a user's real-time status information; using different digital humans pre-built using agent technology to provide interactive services to users in different states; controlling each candidate digital human to initiate tentative interaction requests to the user at different times, with the interaction requests containing the first round of dialogue content matching the real-time status information; and controlling the target digital human to present real-time interactive behaviors to the user using a target image matching the real-time status information; the target digital human being is the candidate digital human for whom the user has accepted the corresponding interaction request. This solution can provide users with real-time interactive services that are more closely matched to their real-time status information and more immersive.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of data processing technology, specifically to the field of artificial intelligence technology such as generative large language models, intelligent agents, digital humans, and more particularly to an immersive digital human interaction method, device, electronic device, computer-readable storage medium, and computer program product based on multi-agent collaboration. Background Art

[0002] With the rapid development and iteration of generative large language models, they have a better understanding of user input requirements and the ability to give corresponding results.

[0003] To tailor the output of the generative large language model to specific needs, an intelligent agent was constructed using the generative large language model as a base model and combining it with pre-set character parameters. Furthermore, to enhance the sense of anthropomorphism and increase user acceptance of interacting with the intelligent agent, a digital human with an anthropomorphic image was developed based on the intelligent agent.

[0004] How to use digital humans to provide users with better interaction methods is still a key research topic that needs to be urgently addressed by technical personnel in this field. Summary of the Invention

[0005] The embodiments of the present disclosure propose an immersive digital human interaction method, device, electronic device, computer-readable storage medium and computer program product based on multi-agent collaboration.

[0006] In the first aspect, an embodiment of the present disclosure proposes an immersive digital human interaction method based on multi-agent collaboration, including: determining multiple candidate digital humans based on the user's real-time status information; wherein, different digital humans pre-constructed based on agent technology are used to provide interactive services to users in different states; controlling each candidate digital human to initiate a tentative interaction request to the user at different times; wherein, the interaction request includes the first round of dialogue content that matches the real-time status information; controlling the target digital human to present to the user a real-time interactive behavior with a target image that matches the real-time status information; wherein, the target digital human is a candidate digital human whose corresponding interaction request has been accepted by the user, and the real-time interactive behavior includes real-time interactive content feedback and real-time interactive posture feedback for the real-time information transmitted by the user.

[0007] On the second aspect, an embodiment of the present disclosure proposes an immersive digital human interaction device based on multi-agent collaboration, comprising: a candidate digital human determination unit, configured to determine multiple candidate digital humans based on the user's real-time status information; wherein, different digital humans pre-constructed based on agent technology are used to provide interactive services to users in different states; a tentative interaction request initiation control unit, configured to control each candidate digital human to initiate a tentative interaction request to the user at different times; wherein, the interaction request includes the first round of dialogue content matching the real-time status information; a real-time interaction control unit, configured to control the target digital human to present to the user a real-time interaction behavior with a target image matching the real-time status information; wherein, the target digital human is a candidate digital human whose corresponding interaction request has been accepted by the user, and the real-time interaction behavior includes real-time interaction content feedback and real-time interaction posture feedback for the real-time information transmitted by the user.

[0008] In a third aspect, an embodiment of the present disclosure provides an electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can implement the immersive digital human interaction method based on multi-agent collaboration as described in the first aspect when executing the instructions.

[0009] In a fourth aspect, an embodiment of the present disclosure provides a non-transitory computer-readable storage medium storing computer instructions, which are used to enable a computer to implement the immersive digital human interaction method based on multi-agent collaboration as described in the first aspect when executed.

[0010] In a fifth aspect, an embodiment of the present disclosure provides a computer program product comprising a computer program, which, when executed by a processor, can implement the various steps of the immersive digital human interaction method based on multi-agent collaboration as described in the first aspect.

[0011] The immersive digital human interaction solution based on multi-agent collaboration provided by the present disclosure first determines multiple candidate digital humans based on the user's real-time status information. The digital humans are pre-built based on agent technology and are designed to provide interactive services to users in different states. Then, under the control of the master agent, each candidate digital human initiates a trial interaction request to the user at different times. The interaction request includes the content of the first round of dialogue that matches the real-time status information. By matching the real-time status information, the user's desire to interact with the candidate digital human that initiated the first round of dialogue is maximized. When the user accepts the interaction request initiated by a candidate digital human and wishes to engage in formal interaction, the target digital human, under the control of the master agent, will also present to the user real-time interaction behavior using a target image that matches the real-time status information. This real-time interaction behavior includes real-time interaction content feedback and real-time interaction posture feedback in response to the real-time information transmitted by the user. Due to the provision of instant generation and rendering of content and images, the digital human can provide users with a more immersive interactive experience, improving users' acceptance of interaction with the agents.

[0012] That is, the present disclosure adopts an intelligent agent cluster formed by a pre-built main intelligent agent and multiple digital humans based on the intelligent agent to provide users with interactive services that match their real-time status information. That is, through the collaboration between the main intelligent agent and the digital humans built based on intelligent agent technology, the appropriate digital humans can fully exert their own performance and capabilities under the control of the main intelligent agent. Compared with the original related technologies with fixed interactive postures and only real-time interactive content feedback, this solution can not only provide users with real-time interactive services that are more matched with their real-time status information and more immersive, but also indirectly improve users' acceptance of interacting with intelligent agents.

[0013] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Other features, objects and advantages of the present disclosure will become more apparent from a reading of the detailed description of non-limiting embodiments made with reference to the following drawings:

[0015] Figure 1 is an exemplary system architecture in which the present disclosure may be applied;

[0016] Figure 2 A flowchart of an immersive digital human interaction method based on multi-agent collaboration provided in an embodiment of the present disclosure;

[0017] Figure 3 A flowchart of a method for determining a target domain based on natural language input and real-time status information provided by an embodiment of the present disclosure;

[0018] Figure 4 A flowchart of a method for creating a digital human provided in an embodiment of the present disclosure;

[0019] Figure 5 A flowchart of a method for establishing further interaction channels based on interaction heat provided by an embodiment of the present disclosure;

[0020] Figure 6 A flowchart of a method for controlling a candidate digital human to generate an interaction request provided by an embodiment of the present disclosure;

[0021] Figure 7-1 to Figure 7-2 These are all example diagrams of an immersive digital human interaction method based on multi-agent collaboration in an application scenario provided by an embodiment of the present disclosure;

[0022] Figure 8 A structural block diagram of an immersive digital human interaction device based on multi-agent collaboration provided by an embodiment of the present disclosure;

[0023] Figure 9 A schematic structural diagram of an electronic device suitable for executing an immersive digital human interaction method based on multi-agent collaboration provided in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0024] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those skilled in the art that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description. It should be noted that the embodiments in the present disclosure and the features in the embodiments can be combined with each other unless there is a conflict.

[0025] In the technical solutions disclosed herein, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0026] Figure 1 An exemplary system architecture 100 is shown to which embodiments of the immersive digital human interaction method, apparatus, electronic device, and computer-readable storage medium based on multi-agent collaboration of the present disclosure can be applied.

[0027] like Figure 1As shown, system architecture 100 may include terminal devices 101, 102, 103, a network 104, and a server 105. Network 104 is a medium for providing communication links between terminal devices 101, 102, 103 and server 105. Network 104 may include various connection types, such as wired or wireless communication links or fiber optic cables.

[0028] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various applications for enabling information communication between the terminal devices 101, 102, and 103 and server 105 can be installed, such as complex task processing applications, browser applications, and instant messaging applications.

[0029] Terminal devices 101, 102, 103 and server 105 can be either hardware or software. When terminal devices 101, 102, 103 are hardware, they can be various electronic devices with display screens, including but not limited to smartphones, tablet computers, laptop computers, and desktop computers. When terminal devices 101, 102, 103 are software, they can be installed in the electronic devices listed above. They can be implemented as multiple software or software modules, or as a single software or software module, and are not specifically limited here. When server 105 is hardware, it can be implemented as a distributed server cluster consisting of multiple servers, or as a single server. When the server is software, it can be implemented as multiple software or software modules, or as a single software or software module, and are not specifically limited here.

[0030] The server 105 can provide various services through various built-in applications. Taking a complex task processing application that can provide one-click processing services for complex tasks as an example, the server 105 can achieve the following effects when running the complex task processing application: first, it can receive real-time status information transmitted by the user using the terminal devices 101, 102, and 103 through the network 104, and then determine multiple candidate digital humans based on the real-time status information. The different digital humans pre-built based on intelligent agent technology are used to provide interactive services to users in different states; then, each candidate digital human is controlled to initiate a tentative interaction request to the user's current browsing behavior at different times. The interaction request contains the first round of dialogue content that matches the real-time status information; finally, the target digital human is controlled to present a real-time interaction behavior to the user with a target image that matches the real-time status information. The target digital human is a candidate digital human whose corresponding interaction request has been accepted by the user. The real-time interaction behavior includes real-time interaction content feedback and real-time interaction posture feedback made to the real-time information transmitted by the user.

[0031] Furthermore, the server 105 may also transmit the real-time data corresponding to the real-time interactive behavior back to the terminal devices 101, 102, 103 via the network 104 in real time, so that the terminal devices 101, 102, 103 display the received real-time data to the user in real time.

[0032] It should be noted that, in addition to being temporarily acquired from terminal devices 101, 102, and 103 via network 104, real-time status information can also be pre-stored or recorded locally on server 105 in various ways. Therefore, when server 105 detects that such data is already locally stored (e.g., when starting to process a previously pending task), it can choose to directly acquire such data locally. In this case, exemplary system architecture 100 may also exclude terminal devices 101, 102, and 103 and network 104.

[0033] Since providing real-time interactive behaviors based on digital humans requires a relatively large amount of computing resources and strong computing power, the immersive digital human interaction methods based on multi-agent collaboration provided in the subsequent embodiments of this disclosure are generally executed by a server 105 with relatively strong computing power and resources. Correspondingly, the immersive digital human interaction device based on multi-agent collaboration is also generally provided in the server 105. However, it should also be pointed out that when the terminal devices 101, 102, and 103 also have computing power and computing resources that meet the requirements, the terminal devices 101, 102, and 103 can also complete the various calculations assigned to the server 105 through the complex task processing applications installed thereon, and then output the same results as the server 105. In particular, when there are multiple terminal devices with different computing capabilities, if a complex task processing application determines that the terminal device it is located on has stronger computing capabilities and more remaining computing resources, it can allow the terminal device to perform the aforementioned operations, thereby appropriately reducing the computing pressure on the server 105. Accordingly, the immersive digital human interaction device based on multi-agent collaboration can also be set up in the terminal devices 101, 102, and 103. In this case, the exemplary system architecture 100 may also not include the server 105 and the network 104.

[0034] It should be noted that the main agent and each sub-target agent can be installed on the server 105 at the same time, and each sub-target agent called and controlled by the main agent can also be installed on other servers or terminal devices different from the server 105. No specific limitation is made here.

[0035] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.

[0036] Please refer to Figure 2 , Figure 2 This is a flowchart of an immersive digital human interaction method based on multi-agent collaboration provided by an embodiment of the present disclosure, wherein process 200 includes the following steps:

[0037] Step 201: Determine multiple candidate digital humans based on the user's real-time status information;

[0038] This step is intended to be carried out by the execution subject of the immersive digital human interaction method based on multi-agent collaboration (e.g. Figure 1 The server 105 (shown as hosting a master agent) determines multiple candidate digital humans based on the user's real-time status information. The candidate digital humans are selected and matched by the execution entity from multiple pre-built digital humans based on the user's real-time status information. All digital humans are constructed based on agent technology, and the different constructed digital humans are designed to provide interactive services to users in different states. By subdividing the states, many different specific states can be identified, such as working, resting, angry, happy, shopping, meeting, engaged, disappointed, injured, and so on.

[0039] The real-time status information may include at least one of the following:

[0040] Short-term interests expressed within a first preset time period from the current moment, long-term interests expressed within a second preset time period from the current moment, real-time emotional polarity, historical fluctuations in emotional polarity, and browsing preferences corresponding to current information browsing behavior, where the second preset time period is much longer than the first preset time period. In other words, this real-time status information can reflect not only the user's long-term and short-term interests, but also the user's real-time emotional tendencies, historical emotional fluctuations, and information browsing preferences. It can also include real-time consumption information, historical consumption information, and consumption preferences.

[0041] In order to determine multiple candidate digital humans based on the user's real-time status information, one implementation method is to select suitable candidate digital humans based solely on the intelligence recorded or reflected in the real-time status information. Another implementation method is to combine the real-time status information with other information from the user as much as possible, such as the user's current natural language input or current information browsing behavior, application usage behavior, etc.

[0042] In some other implementations, the aforementioned execution entity may first determine the user's target field of interest based on the user's real-time status information, and then select multiple candidate digital humans from multiple digital humans belonging to the target field. In other words, the digital humans in this case should be used to provide interactive services to users interested in different fields. "Field" is a relatively broad concept, and a field of interest can specifically be at least one of a topic of interest, an object of interest, a person of interest, science and technology of interest, knowledge of interest, or a region of interest.

[0043] Step 202: Control each candidate digital human to initiate a trial interaction request to the user at different times;

[0044] Based on step 201, this step aims to control each candidate digital human to initiate a trial interaction request to the user at different times. The interaction request includes the first round of conversation content that matches the real-time status information. In other words, by using the first round of conversation content that matches the real-time status information, the user who sees the interaction request is expected to be motivated to interact with the corresponding candidate digital human.

[0045] Specifically, the timing for initiating a trial interaction request may include at least one of the following:

[0046] The total time the user browses all information during this browsing behavior (i.e., an interaction request is inserted at the corresponding browsing position when the total browsing time reaches a certain set time threshold), the duration of the user's browsing of any information during this browsing behavior (i.e., an interaction request is inserted at a suitable browsing position in the information when the browsing time of a single information reaches a certain set time threshold), the total number of information items browsed during this browsing behavior (i.e., an interaction request is inserted at a suitable browsing position in the information browsing interface when the total number of items browsed reaches a certain set number threshold), the preset position of the waterfall information flow presented by the browser (i.e., an interaction request can be inserted at certain preset dynamic insertion positions in the waterfall information flow), the end position of the information body presented by the browser, etc. These are just a limited list of possible ways, but are not limited to the above-mentioned situations.

[0047] Step 203: Control the target digital human to present to the user a real-time interactive behavior using the target image that matches the real-time status information.

[0048] Based on step 202, this step aims to control the target digital human to present to the user a real-time interactive behavior using a target image that matches the real-time status information. The target digital human is a candidate digital human whose corresponding interaction request has been accepted by the user. The real-time interactive behavior includes real-time interactive content feedback and real-time interactive posture feedback for the real-time information passed in by the user. The real-time interactive content feedback is used to ensure that the dialogue content feedback matches the real-time information passed in by the user, and the real-time interactive posture feedback is used to ensure that the overall digital human image matches the real-time information passed in by the user, rather than being unchanging, to enhance the sense of immersion.

[0049] Specifically, the real-time content of the real-time interactive content feedback includes: text content, content expression style, modal particles, dialect types, etc.; the real-time posture of the real-time interactive posture feedback includes: real-time body movements, real-time facial expressions, real-time clothing, object interaction actions with real-time objects, etc.

[0050] Specifically, the target image can be matched to the real-time status information by at least one of the following image parameters:

[0051] This includes physical attribute parameters such as gender, height, and body shape, personality parameters, clothing parameters, posture parameters, dialogue expression style parameters, and other parameter categories that can also affect the adaptation to user needs.

[0052] It should be understood that each candidate digital human and target digital human should actually generate interaction requests and real-time interaction behaviors under the call or arrangement of the main intelligent entity. Specifically, the main intelligent entity can actively send relevant information on how to generate interaction requests and real-time interaction behaviors to each digital human, so that the digital human can actively generate interaction requests, select appropriate timing, and create appropriate images according to the received relevant information. The main intelligent entity can also actively call each candidate digital human to generate interaction requests as required when it determines that there is a suitable time for each candidate digital human to issue an interaction request, and the main intelligent entity will issue it at the appropriate time selected by the main intelligent entity. In other words, each candidate digital human only passively performs the corresponding operation according to the call instruction.

[0053] Both the active execution mechanism and the passive execution mechanism mentioned above belong to different forms of collaboration between the main intelligent agent and multiple digital humans to jointly process tasks. The specific choice can be based on the actual needs of the actual application scenario and is not specifically limited here.

[0054] In an immersive digital human interaction method based on multi-agent collaboration, a master agent first determines multiple candidate digital humans based on a user's real-time status information. These digital humans are pre-built based on agent technology and are designed to provide interactive services to users in different states. Then, under the control of the master agent, each candidate digital human initiates a tentative interaction request to the user at different times. The interaction request includes the content of a first round of dialogue that matches the real-time status information. This is done to maximize the user's desire to interact with the candidate digital human that initiated the first round of dialogue by matching the real-time status information. When the user accepts the interaction request initiated by a candidate digital human and wishes to formally interact, the target digital human, under the control of the master agent, will present to the user real-time interaction behavior using a target image that matches the real-time status information. This real-time interaction behavior includes real-time interaction content feedback and real-time interaction posture feedback in response to the real-time information transmitted by the user. By providing instant generation and rendering of content and images, the digital human provides the user with a more immersive interactive experience, improving the user's acceptance of interaction with the agent.

[0055] That is, this embodiment uses an intelligent agent cluster formed by a pre-built main intelligent agent and multiple digital humans based on intelligent agents to provide users with interactive services that match their real-time status information. That is, through the collaboration between the main intelligent agent and the digital humans built based on intelligent agent technology, the appropriate digital humans can fully exert their own performance and capabilities under the control of the main intelligent agent. Compared with the original related technologies with fixed interactive postures and only real-time interactive content feedback, the solution provided by this embodiment can not only provide users with real-time interactive services that are more matched with their real-time status information and more immersive, but also indirectly improve users' acceptance of interacting with intelligent agents.

[0056] To further understand how to identify target topics, please refer to Figure 3 , Figure 3 This is a flowchart of a method for determining a target domain based on natural language input and real-time status information provided by an embodiment of the present disclosure, wherein process 300 includes the following steps:

[0057] Step 301: Convert the user's natural language input into natural language text;

[0058] This step aims to convert the natural language input in various forms into natural language text that is easy to recognize and output by the above-mentioned execution subject, that is, to uniformly convert various forms into text form. For example, speech-to-text technology can be used to convert natural speech signals in speech form into natural language text.

[0059] Step 302: Perform semantic understanding and domain entity extraction on the natural language text to obtain suspected domains of interest;

[0060] Based on step 301, this step aims to perform semantic understanding and domain entity extraction on the natural language text by the above-mentioned execution subject to obtain suspected domains of interest.

[0061] This step can usually be implemented within the technical framework of natural language processing, and the implementation process can be broken down into the following steps:

[0062] 1) Semantic Understanding

[0063] Semantic understanding refers to extracting deeper meaning from text, not just the literal meaning. It typically involves the following key technologies: word sense disambiguation: Different words can have multiple meanings (e.g., "bank" can mean "bank" or "river bank"), and the correct meaning of a word needs to be determined through context; syntactic analysis: understanding the grammatical structure within a sentence and determining the relationship between words. This is typically achieved through a parse tree, which helps the model identify components such as the subject, predicate, and object; semantic role labeling (SRL): identifying different components in a sentence and their semantic roles, such as who did what, what the goal was, and how it was done; and contextual understanding: inferring the specific semantics of words and phrases through context. For example, "apple" can refer to different things (fruit vs. company) in different contexts.

[0064] 2) Entity Recognition

[0065] It refers to identifying entities or information related to a specific field from text. These entities can be names of people, places, companies, products, time, etc.

[0066] Commonly used techniques include: Named Entity Recognition (NER): a technique for automatically identifying entities with specific meanings in text; Entity Linking: accurately annotating identified entities to ensure they match entities in external knowledge bases. For example, identifying whether "apple" refers to a fruit or a technology company and linking it to the corresponding database entry; and Relation Extraction: extracting the relationships between entities from text. For example, extracting the relationship between "Apple" and "founded in" from the sentence "Apple was founded in 1976."

[0067] A specific step to extract domain entities can be:

[0068] ① Data preprocessing: perform preprocessing such as word segmentation, stop word removal, and part-of-speech tagging on the text;

[0069] ② Model selection and training: Use methods such as conditional random fields, long short-term memory networks, or pre-trained language models to identify domain-related entities;

[0070] ③Entity recognition: The model will label the entities in the text and classify them (such as "person's name", "place name", "date", etc.).

[0071] 3) Analysis of suspected areas of interest

[0072] After semantic understanding and entity extraction of the text, the next step is to determine whether the text involves a specific field or topic based on the extracted entities and context.

[0073] This can be achieved through the following methods: Domain classification: Based on the extracted entities and semantic information, the text is assigned to predefined domain labels. For example, the medical field, the financial field, the technology field, etc. Machine learning models (such as SVM, Random Forest) or deep learning models can be used for training; Keyword matching: By building a domain-related keyword library, key entities in the text are matched with these keywords to infer possible fields of interest. For example, if words such as "heart disease" and "surgery" appear frequently in the text, it can be inferred that the text may belong to the medical field; Topic model-based analysis: Using topic modeling techniques such as LDA (Latent Dirichlet Allocation), implicit topics in the text are analyzed to determine the field of interest. The distribution of topics can reveal the core topics discussed in the text; Cluster analysis: Using unsupervised learning methods, texts are clustered based on semantic similarity. Similar texts are clustered in the same cluster, and the field or topic covered by the text can be inferred.

[0074] A specific step may be:

[0075] ① Preprocessing and feature extraction: Convert the text into a vector representation (such as TF-IDF, Word2Vec, or BERT vectors) and extract key features. ② Model training: Use a classification model to train the text to determine which areas or topics are of interest. You can use existing annotated datasets to train a multi-category classification model. ③ Inference and prediction: Input the text to be analyzed into the trained model to predict its likely areas of interest.

[0076] Step 303: Using the real-time status information, modify the suspected area of ​​interest to obtain the target area;

[0077] On the basis of step 302, in order to improve the accuracy, this step aims to have the execution subject perform interest correction on the suspected area of ​​interest using the real-time status information to obtain a corrected target area.

[0078] That is, the implementation method then uses the user's real-time status information to modify the suspected area of ​​interest, such as screening, excluding, supplementing implicit or missing information, etc., so that the final target area is more consistent with the user's actual situation.

[0079] Step 304: Select multiple digital humans belonging to the target field as multiple candidate digital humans.

[0080] This embodiment provides a more specific implementation method for determining the target domain based on natural language input and real-time status information through steps 301 to 304, which includes key processing steps such as conversion of expression form, semantic understanding, real-time extraction, and interest correction based on real-time status information, so as to provide an implementation solution with high feasibility and more accurate determination of the target domain as much as possible.

[0081] To further understand how to pre-create different digital people, see Figure 4 , Figure 4 This is a flowchart of a method for creating a digital human provided in an embodiment of the present disclosure, wherein process 400 includes the following steps:

[0082] Step 401: Determine a state set for which interactive services need to be provided;

[0083] This step determines a state set by analyzing and identifying the current environment, user needs, and interaction context. This state set includes multiple interaction contexts or states, which may arise from changes in user behavior, needs, or the environment. For example, states may include "Requesting Help," "Information Query," "Emotional Support," etc., each representing a different interaction need. By monitoring the system or user, the system can dynamically determine which interaction state is present and respond accordingly.

[0084] Step 402: For each state in the state set, a digital human with corresponding state processing capabilities and an image adapted to various user preferences is constructed based on intelligent agent technology.

[0085] This step can be divided into multiple sub-steps, such as building a digital human based on intelligent agent technology, providing it with a digital human image that adapts to user preferences, and enabling it to have corresponding state processing capabilities. The following will be explained in detail:

[0086] 1) Building digital humans based on intelligent agent technology

[0087] For each identified interaction state, a digital human adapted to that state can be constructed based on intelligent agent technologies. These technologies typically include natural language processing, machine learning, sentiment analysis, speech recognition, and computer vision, enabling dynamic responses based on user input, behavior, and preferences. Specifically, when a specific state is identified, the digital human will possess the necessary capabilities to handle it.

[0088] For example, if the current status is "requesting help", the digital human can appear friendly and patient, and have the ability to handle problem solving and guidance; if the status is "emotional support", the digital human can express care and comfort through tone and body language to meet the user's emotional needs.

[0089] 2) Adapt digital human images to user preferences

[0090] In addition to generating processing capabilities based on state, digital humans also need to be able to adapt to different user preferences. User preferences may include voice style, appearance, interaction methods, and more. Through technologies such as machine learning, user preferences can be modeled and customized digital human avatars can be generated based on their needs. This adaptability can significantly enhance the user experience, ensuring that each user receives interactive services that better meet their individual needs.

[0091] For example, some users may prefer a concise and direct interaction style, while others may prefer a warm and emotional interaction style.

[0092] 3) Enable it to have corresponding state processing capabilities

[0093] Each digital human will have corresponding state processing capabilities based on the state it is in. The formation of this capability relies on a variety of intelligent algorithms and models in intelligent agent technology, including:

[0094] Natural language understanding: Identifying the user's intentions and emotions based on the sentences they input, thereby determining how the digital human responds; emotional computing: Selecting appropriate responses based on the user's emotional state, such as tone, expression, body language, etc.; contextual awareness: Considering the user's historical behavior, current environment, and interaction history to select appropriate interaction methods, as well as other special types of state processing capabilities, which are not listed here one by one.

[0095] This embodiment provides a specific implementation method for constructing a digital human with corresponding state processing capabilities and an image adapted to various user preferences based on the intelligent agent technology for the determined multiple states through steps 401 and 402.

[0096] In order to further enhance the closeness between digital humans and real people, digital humans can also include simulated digital humans that imitate experts and scholars in the real world, representative figures, or human objects with a popularity exceeding a preset level, that is, when constructing the simulated digital humans, full reference is made to the various attribute parameters of real people.

[0097] See also Figure 5 , Figure 5 This is a flow chart of a method for establishing further interaction channels based on interaction popularity provided by an embodiment of the present disclosure, wherein process 500 includes the following steps:

[0098] Step 501: controlling the simulated digital human to determine the interaction heat with the user;

[0099] Among them, the interaction heat can be reflected and calculated through multiple dimensions, such as the number of conversations, the depth of the conversation content, the length of the conversation reflection, the emotional tendency of the conversation content, the emotional information conveyed by the user's conversation content, etc.

[0100] Step 502: In response to the interaction heat exceeding the preset heat, an interaction channel is established between the user and the target real person imitated by the simulated digital human.

[0101] This step is based on the premise that the interaction heat exceeds the preset heat, and aims to enable the above-mentioned execution entity to establish an interaction channel between the user and the target real person imitated by the simulated digital human, so that the user can interact with the target real person, thereby facilitating the user's interaction from the level of virtual human to the level of real person, and also facilitating the reception of some information conveyed by the target real person.

[0102] Based on any of the above embodiments, please refer to Figure 6 , Figure 6 This is a flow chart of a method for controlling a candidate digital human to generate an interaction request provided by an embodiment of the present disclosure. The process 600 includes the following steps:

[0103] Step 601: Control each candidate digital human to generate a first round of interactive dialogue that matches the real-time status information;

[0104] In order to avoid repetition and learn from failure, the content of the first round of interactive dialogues generated by different candidate digital humans can also be controlled to be not completely consistent.

[0105] Step 602: Control each candidate digital human to generate an interactive guidance image including a corresponding first-round interactive dialogue and a target image matching the real-time status information;

[0106] Step 603: Control each candidate digital human to initiate a trial interaction request to the user at different times, with the interaction guide image as the request content.

[0107] This embodiment further specifically provides a method of generating an interaction guidance image together with the target image on the basis of the first round of interactive dialogue through steps 601 to 603, and then using the interaction guidance image as the interaction request content. In this way, the interaction request contains both content information of the target field of interest to the user and anthropomorphic image information that matches the user's real-time status information, thereby further increasing the user's tendency to accept the interaction request.

[0108] To deepen understanding, the present disclosure also provides a specific implementation solution that attempts to eliminate the existing technical defects and overcome the existing technical problems in combination with the actual existing technical defects in specific application scenarios:

[0109] Current content consumption is based on the distribution of recommendations for content already produced by creators. This approach uses personalized recommendation technology to match user interests with content, and a multi-layered ranking funnel filters out the final content for distribution. The advantage of this static content model is that it allows ecosystem creators to create content in multiple fields, with quality content included through content evaluation. Production and consumption are separated, allowing for independent optimization. However, the drawbacks are also significant. Content cannot be flexibly modified after production, and content is not interactive. The matching of content with user interests is distributed through the serial connection of various modules in the recommendation system, resulting in significant error accumulation, high precision requirements for each module, and a limited ceiling.

[0110] Another form of real-time content consumption is live streaming recommendations. Live streams on various topics are distributed through a recommendation system, and users can interact with the streamers through the comment section. This format produces content in real time and offers a certain degree of interactivity. However, this model creates a one-to-many relationship between streamers and users, making it impossible to provide one-to-one real-time service. Furthermore, one-to-one live streaming for users is costly and only feasible in specialized consumption scenarios.

[0111] To address these difficulties, a digital human consumption satisfaction technology based on intelligent agent dialogue is proposed for immersive real-time content consumption scenarios. The characteristics of this solution are as follows:

[0112] 1) It is consumable content that is built purely automatically and does not rely on the content ecosystem or creator operations. It has the ability to quickly generate content and iterate on results.

[0113] 2) This solution is built on a large model. As the performance of the large model improves and the training cost of the large model decreases, the overall conversation experience and commercial value will continue to improve.

[0114] 3) Production based on the user's real-time status can greatly capture the dynamic changes in the user's status, including changes in the user's short-term consumption interests and consumption emotions, and can better meet the user's current consumption needs.

[0115] This embodiment proposes an agent-based immersive digital human dialogue technology that can be applied to information flow consumption scenarios. It can generate immersive and consumable topics based on user interests and simultaneously output interactive consumption through digital human dialogue with users. The main tasks include the following:

[0116] 1) Content planning based on user status. Given a user's interest profile, real-time consumption dynamics, and real-time consumption sentiment, a plan is generated to recommend content of interest to the user in real time.

[0117] 2) Generate a list of digital human agents. Based on the planning of interest content, a large model is used to generate a list of digital human immersive consumption, supporting different role types, such as emotional companionship roles, sales and service roles, etc.

[0118] 3) Character content generation: After the digital human immersion consumption list is determined, a digital image and the first round of conversation content that matches the user's interests and status are generated for each specific digital human.

[0119] The implementation block diagram of this solution is as follows Figure 7-1 As shown, the input is the user's profile and short-term user status, as well as their short-term consumer interests, consumer satisfaction, and other emotional states. The output is a real-time, immersive streaming list of digital human consumption. The entire system is built based on multi-agent collaboration and mainly consists of two types of agents: 1) Main Agent: This agent performs overall task understanding, decomposition, and concatenation, providing the final result. 2) Content Agent: This agent generates a list of recommended agents, and each sub-agent generates corresponding consumption content based on the task understanding input.

[0120] The solution provided by this embodiment can be widely applied to various information-satisfying applications or independent products. The following uses the scenario of the BaiX application as an example to illustrate the specific application form of this embodiment:

[0121] In the user's consumption recommendation flow, the real-time generated immersive content flow of the invention is inserted, such as Figure 7-2 As shown in the figure, users click on the corresponding cover to enter the immersive consumption page, where the dynamic digital human image is displayed and the first round of conversation content is generated.

[0122] When the user slides the screen up and down, other digital humans of the same type will be displayed, showing the full myopic consumption form. The user can interact with them by speaking or inputting text, and the digital human image of different users will be personalized.

[0123] The above examples show that this embodiment has the following improvements and technical effects compared to the prior art:

[0124] 1) The user's consumption content is matched and generated in real time with the user's interests and status, which greatly improves satisfaction.

[0125] Traditional content recommendation algorithms are often based on users' historical interests and feedback, but they are unable to capture users' real-time consumption status or make good use of their real-time status. There is no precise connection between users' consumption status and content production. The present invention can take the user's current real-time status, including highly real-time signals such as interests, consumption content, and consumption mood, as input, and generate content that matches the immediate status in real time. This can greatly improve users' acceptance of content and truly understand users and provide consumption content that can satisfy them.

[0126] 2) Based on multi-agent collaboration, the recommendation generation and content generation are organically combined.

[0127] Because current mainstream content streaming consumption products are created by ecosystems, content consumption technologies can only achieve generative recommendations through recommendation algorithms or real-time content generation, but not immersive consumption. This invention integrates these two aspects into a holistic model, creating an integrated generation system that significantly increases the ceiling for optimizing the immersive consumption experience.

[0128] 3) The adoption of an immersive digital human consumption model can maintain consistency in consumer perception with the current mainstream content model, and naturally connect with the content of ecosystem creators.

[0129] In the technology provided in this embodiment, digital human conversational consumption, due to its efficient and accurate matching, serves as the first connection point. After meeting basic user needs and building trust, it can be further connected to real-person services and content, achieving deeper satisfaction and commercial conversion. This has significant application value in scenarios such as emotional counseling, legal consultation, and medical consultation.

[0130] Further references Figure 8 As an implementation of the methods shown in the above figures, the present disclosure provides an embodiment of an immersive digital human interaction device based on multi-agent collaboration. Figure 2 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.

[0131] like Figure 8As shown, the immersive digital human interaction device 900 based on multi-agent collaboration in this embodiment may include: a candidate digital human determination unit 801, a tentative interaction request initiation control unit 802, and a real-time interaction control unit 803. The candidate digital human determination unit 801 is configured to determine multiple candidate digital humans based on the user's real-time status information. The different digital humans pre-built based on agent technology are used to provide interactive services to users in different states. The tentative interaction request initiation control unit 802 is configured to control each candidate digital human to initiate a tentative interaction request to the user at different times. The interaction request includes the first round of dialogue content that matches the real-time status information. The real-time interaction control unit 803 is configured to control the target digital human to present a real-time interaction behavior to the user using a target image that matches the real-time status information. The target digital human is the candidate digital human whose corresponding interaction request has been accepted by the user. The real-time interaction behavior includes real-time interaction content feedback and real-time interaction gesture feedback in response to the real-time information input by the user.

[0132] In this embodiment, in the immersive digital human interaction device 800 based on multi-agent collaboration, the specific processing of the candidate digital human determination unit 801, the tentative interaction request initiation control unit 802, and the real-time interaction control unit 803 and the technical effects thereof can be referred to in the respective Figure 2 The relevant descriptions of steps 201-203 in the corresponding embodiment are not repeated here.

[0133] In some optional implementations of this embodiment, the real-time status information includes at least one of the following:

[0134] Short-term interests expressed within a first preset time period from the current moment, long-term interests expressed within a second preset time period from the current moment, real-time emotional polarity, historical fluctuations in emotional polarity, and browsing preferences corresponding to current information browsing behavior; wherein the second preset time period is much longer than the first preset time period.

[0135] In some optional implementations of this embodiment, the target image is matched to the real-time status information through at least one of the following image parameters:

[0136] Including physical attribute parameters such as gender, height, and body shape, personality parameters, clothing parameters, posture parameters, and dialogue expression style parameters.

[0137] In some optional implementations of this embodiment, the candidate digital human determining unit 801 may include:

[0138] an interest area determination subunit, configured to determine a target area of ​​interest to the user based on the real-time status information;

[0139] The candidate digital human selection subunit is configured to select multiple digital humans belonging to the target field as multiple candidate digital humans; among which, different digital humans pre-built based on intelligent agent technology are used to provide interactive services to users interested in different fields.

[0140] In some optional implementations of this embodiment, the field of interest includes at least one of the following:

[0141] Topics of interest, objects of interest, people of interest, science and technology of interest, knowledge of interest, areas of interest.

[0142] In some optional implementations of this embodiment, the area of ​​interest determination subunit includes:

[0143] The interest domain determination module is configured to determine the target domain of interest to the user based on the user's natural language input and real-time status information.

[0144] In some optional implementations of this embodiment, the area of ​​interest determination module is further configured to:

[0145] Convert user's natural language input into natural language text;

[0146] Perform semantic understanding and domain entity extraction on natural language text to obtain suspected areas of interest;

[0147] The real-time status information is used to modify the suspected areas of interest and obtain the target areas.

[0148] In some optional implementations of this embodiment, the immersive digital human interaction device 800 based on multi-agent collaboration may further include a digital human construction unit, which is further configured to:

[0149] Determine the state set that needs to provide interactive services;

[0150] For each state in the state set, a digital human with corresponding state processing capabilities and an image adapted to various user preferences is built based on intelligent agent technology.

[0151] In some optional implementations of this embodiment, the digital human includes a simulated digital human that imitates experts and scholars, representative figures, or human objects with a temperature exceeding a preset value in the real world.

[0152] In some optional implementations of this embodiment, the immersive digital human interaction device 800 based on multi-agent collaboration may further include:

[0153] An interaction heat determination control unit configured to control the simulated digital human to determine the interaction heat of the interaction with the user;

[0154] The interaction channel establishing unit is configured to establish an interaction channel between the user and the target real person imitated by the simulated digital human in response to the interaction heat exceeding the preset heat, so that the user can interact with the target real person.

[0155] In some optional implementations of this embodiment, the timing includes at least one of the following:

[0156] The total time the user browses all information during this information browsing behavior, the duration the user browses any information during this information browsing behavior, the total number of information items browsed by the user during this information browsing behavior, the preset position of the browser when presenting the waterfall-style information flow, and the end position of the body of the browsed information presented by the browser.

[0157] In some optional implementations of this embodiment, the tentative interaction request initiation control unit 802 is further configured to:

[0158] Control each candidate digital human to generate a first round of interactive dialogue that matches the real-time status information;

[0159] Controlling each candidate digital human to generate an interactive guidance image including a corresponding first-round interactive dialogue and a target image matching the real-time status information;

[0160] Control each candidate digital human to initiate a tentative interaction request to the user at different times, with the interaction guide image as the request content.

[0161] In some optional implementations of this embodiment, the immersive digital human interaction device 800 based on multi-agent collaboration may further include:

[0162] The incomplete consistency control unit is configured to control the dialogue contents of the first round of interactive dialogues generated by different candidate digital humans to be incompletely consistent.

[0163] This embodiment exists as an apparatus embodiment corresponding to the above-mentioned method embodiment. This embodiment provides an immersive digital human interaction device based on multi-agent collaboration. A master agent first determines multiple candidate digital humans based on a user's real-time status information. These digital humans are pre-built based on agent technology and are designed to provide interactive services to users in different states. Then, under the control of the master agent, each candidate digital human initiates a tentative interaction request to the user at different times. This interaction request includes the content of a first round of conversation that matches the real-time status information. By matching this real-time status information, the user's desire to interact with the candidate digital human that initiated the first round of conversation is maximized. When the user accepts the interaction request initiated by a candidate digital human and wishes to formally interact, the target digital human, under the control of the master agent, will also present real-time interactive behavior to the user using a target image that matches the real-time status information. This real-time interactive behavior includes real-time interactive content feedback and real-time interactive gesture feedback in response to the real-time information transmitted by the user. By providing instant generation and rendering of content and images, the digital human provides the user with a more immersive interactive experience, improving the user's acceptance of interacting with the agent.

[0164] That is, this solution uses an agent cluster formed by a pre-built main agent and multiple digital humans based on the agent to provide users with interactive services that match their real-time status information. That is, through the collaboration between the main agent and the digital humans built based on agent technology, the appropriate digital humans can fully exert their performance and capabilities under the control of the main agent. Compared with the original related technologies with fixed interactive postures and only real-time interactive content feedback, this solution can not only provide users with real-time interactive services that are more matched with their real-time status information and more immersive, but also indirectly improve users' acceptance of interacting with agents.

[0165] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can implement the immersive digital human interaction method based on multi-agent collaboration described in any of the above embodiments when executing.

[0166] According to an embodiment of the present disclosure, the present disclosure also provides a readable storage medium, which stores computer instructions, and the computer instructions are used to enable a computer to implement the immersive digital human interaction method based on multi-agent collaboration described in any of the above embodiments when executed.

[0167] According to an embodiment of the present disclosure, the present disclosure also provides a computer program product, which, when executed by a processor, can implement the immersive digital human interaction method based on multi-agent collaboration described in any of the above embodiments.

[0168] Figure 9 A schematic block diagram of an example electronic device 900 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0169] like Figure 9 As shown, device 900 includes a computing unit 901, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 902 or a computer program loaded from a storage unit 908 into a random access memory (RAM) 903. RAM 903 may also store various programs and data required for the operation of device 900. Computing unit 901, ROM 902, and RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to bus 904.

[0170] Various components in the device 900 are connected to the I / O interface 905, including an input unit 906, such as a keyboard, a mouse, etc.; an output unit 907, such as various types of displays, speakers, etc.; a storage unit 908, such as a magnetic disk, an optical disk, etc.; and a communication unit 909, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 909 allows the device 900 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0171] The computing unit 901 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 901 performs the various methods and processes described above, such as the immersive digital human interaction method based on multi-agent collaboration. For example, in some embodiments, the immersive digital human interaction method based on multi-agent collaboration can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded into the RAM 903 and executed by the computing unit 901, one or more steps of the immersive digital human interaction method based on multi-agent collaboration described above can be performed. Alternatively, in other embodiments, the computing unit 901 may be configured in any other appropriate manner (e.g., by means of firmware) to execute an immersive digital human interaction method based on multi-agent collaboration.

[0172] Various embodiments of the systems and techniques described above can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0173] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0174] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fibers, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0175] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0176] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0177] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact via a communication network. This client-server relationship arises through computer programs running on the respective computers and establishing a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or cloud host. This is a hosting product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosting and virtual private server (VPS) services.

[0178] The technical solution of the embodiment of the present disclosure,

[0179] The main intelligent agent first determines multiple candidate digital humans based on the user's real-time status information. These digital humans are pre-built based on intelligent agent technology and are designed to provide interactive services to users in different states. Then, under the control of the main intelligent agent, each candidate digital human initiates a tentative interaction request to the user at different times. This interaction request contains the content of the first round of dialogue that matches the real-time status information. By matching the real-time status information, the user's desire to interact with the candidate digital human who initiated the first round of dialogue content is maximized. When the user accepts the interaction request initiated by a candidate digital human and wants to formally interact, the target digital human will also present to the user in real-time interaction behavior under the control of the main intelligent agent, using the target image that matches the real-time status information. This real-time interaction behavior includes real-time interactive content feedback and real-time interactive posture feedback based on the real-time information transmitted by the user. Due to the provision of instant generation and rendering of content and images, the digital human can provide users with a more immersive interactive experience, improving users' acceptance of interaction with the intelligent agent.

[0180] That is, this solution uses an agent cluster formed by a pre-built main agent and multiple digital humans based on the agent to provide users with interactive services that match their real-time status information. That is, through the collaboration between the main agent and the digital humans built based on agent technology, the appropriate digital humans can fully exert their performance and capabilities under the control of the main agent. Compared with the original related technologies with fixed interactive postures and only real-time interactive content feedback, this solution can not only provide users with real-time interactive services that are more matched with their real-time status information and more immersive, but also indirectly improve users' acceptance of interacting with agents.

[0181] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not a limitation herein.

[0182] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.

Claims

1. An immersive digital human interaction method based on multi-agent collaboration, applied to a master agent, comprising: Determine multiple candidate digital humans based on the user's real-time status information; among them, different digital humans pre-built based on intelligent agent technology are used to provide interactive services to users in different states; Controlling each candidate digital person to initiate a tentative interaction request to the user at different times; wherein the interaction request includes the first round of dialogue content that matches the real-time status information; The target digital human is controlled to present to the user a real-time interactive behavior using a target image that matches the real-time status information; wherein the target digital human is a candidate digital human whose corresponding interactive request has been accepted by the user, and the real-time interactive behavior includes real-time interactive content feedback and real-time interactive posture feedback in response to the real-time information transmitted by the user.

2. The method according to claim 1, wherein The real-time status information includes at least one of the following: Short-term interests expressed within a first preset time period from the current moment, long-term interests expressed within a second preset time period from the current moment, real-time emotional polarity, historical fluctuations in emotional polarity, and browsing preferences corresponding to current information browsing behavior; wherein the second preset time period is much longer than the first preset time period.

3. The method according to claim 1, wherein The target image is matched to the real-time status information by at least one of the following image parameters: Including physical attribute parameters such as gender, height, and body shape, personality parameters, clothing parameters, posture parameters, and dialogue expression style parameters.

4. The method according to claim 1, wherein The step of determining a plurality of candidate digital humans based on the user's real-time status information includes: determining a target area of ​​interest to the user based on the real-time status information; A plurality of digital humans belonging to the target field are selected as a plurality of candidate digital humans; wherein the different digital humans pre-built based on the intelligent agent technology are used to provide interactive services to users interested in different fields.

5. The method according to claim 4, wherein Areas of interest include at least one of the following: Topics of interest, objects of interest, people of interest, science and technology of interest, knowledge of interest, areas of interest.

6. The method according to claim 4, wherein: The determining the target area of ​​interest to the user based on the real-time status information includes: The target area of ​​interest to the user is determined based on the user's natural language input and the real-time status information.

7. The method according to claim 6, wherein: The determining, based on the user's natural language input and the real-time status information, the target area of ​​interest to the user includes: Converting the user's natural language input into natural language text; Performing semantic understanding and domain entity extraction on the natural language text to obtain suspected domains of interest; The real-time status information is used to perform interest correction on the suspected area of ​​interest to obtain the target area.

8. The method according to claim 1, wherein The process of constructing the digital human includes: Determine the state set that needs to provide interactive services; For each state in the state set, a digital human with corresponding state processing capabilities and an image adapted to various user preferences is constructed based on the intelligent agent technology.

9. The method according to claim 8, wherein The digital human includes a simulated digital human that is modeled after experts and scholars, representative figures, or human objects with a temperature exceeding a preset value in the real world.

10. The method according to claim 9, further comprising: Controlling the simulated digital human to determine the degree of interaction with the user; In response to the interaction heat exceeding a preset heat, an interaction channel between the user and the target real person imitated by the simulated digital human is established, so that the user can interact with the target real person.

11. The method according to claim 1, wherein The timing includes at least one of the following: The total time the user browses all information during this information browsing behavior, the duration of the user's browsing of any information during this information browsing behavior, the total number of information browsed by the user during this information browsing behavior, the preset position of the browser presenting the waterfall-style information flow, and the end position of the information body presented by the browser.

12. The method according to any one of claims 1 to 11, wherein: The controlling each candidate digital person to initiate a tentative interaction request to the user at different times includes: Controlling each candidate digital person to generate a first round of interactive dialogue matching the real-time status information; Controlling each candidate digital human to generate an interactive guidance image including a corresponding first-round interactive dialogue and a target image matching the real-time status information; Each candidate digital human is controlled to initiate a tentative interaction request to the user at different times, with the interaction guide image as the request content.

13. The method according to claim 12, further comprising: The contents of the first round of interactive dialogues generated by controlling different candidate digital humans are not completely consistent.

14. An immersive digital human interaction device based on multi-agent collaboration, applied to a master agent, comprising: A candidate digital human determining unit is configured to determine a plurality of candidate digital humans based on the real-time status information of the user; wherein different digital humans pre-built based on the intelligent agent technology are used to provide interactive services to users in different states; a tentative interaction request initiation control unit configured to control each candidate digital human to initiate a tentative interaction request to the user at different times; wherein the interaction request includes the first round of dialogue content that matches the real-time status information; A real-time interaction control unit is configured to control the target digital human to present to the user a real-time interaction behavior using a target image that matches the real-time status information; wherein the target digital human is a candidate digital human whose corresponding interaction request has been accepted by the user, and the real-time interaction behavior includes real-time interaction content feedback and real-time interaction posture feedback in response to the real-time information transmitted by the user.

15. The device according to claim 14, wherein The real-time status information includes at least one of the following: Short-term interests expressed within a first preset time period from the current moment, long-term interests expressed within a second preset time period from the current moment, real-time emotional polarity, historical fluctuations in emotional polarity, and browsing preferences corresponding to current information browsing behavior; wherein the second preset time period is much longer than the first preset time period.

16. The device according to claim 14, wherein The target image is matched to the real-time status information by at least one of the following image parameters: Including physical attribute parameters such as gender, height, and body shape, personality parameters, clothing parameters, posture parameters, and dialogue expression style parameters.

17. The device according to claim 14, wherein The candidate digital human determination unit includes: an interest domain determining subunit, configured to determine a target domain of interest to the user based on the real-time status information; The candidate digital human selection subunit is configured to select multiple digital humans belonging to the target field as multiple candidate digital humans; wherein, different digital humans pre-built based on the intelligent agent technology are used to provide interactive services to users interested in different fields.

18. The device according to claim 17, wherein Areas of interest include at least one of the following: Topics of interest, objects of interest, people of interest, science and technology of interest, knowledge of interest, areas of interest.

19. The device according to claim 17, wherein The area of ​​interest determination subunit includes: The interest domain determination module is configured to determine the target domain that the user is interested in based on the user's natural language input and the real-time status information.

20. The device according to claim 19, wherein The area of ​​interest determination module is further configured to: Converting the user's natural language input into natural language text; Performing semantic understanding and domain entity extraction on the natural language text to obtain suspected domains of interest; The real-time status information is used to perform interest correction on the suspected area of ​​interest to obtain the target area.

21. The apparatus according to claim 14, further comprising: A digital human construction unit, wherein the digital human construction unit is further configured to: Determine the state set that needs to provide interactive services; For each state in the state set, a digital human with corresponding state processing capabilities and an image adapted to various user preferences is constructed based on the intelligent agent technology.

22. The device according to claim 21, wherein The digital human includes a simulated digital human that is modeled after experts and scholars, representative figures, or human objects with a temperature exceeding a preset value in the real world.

23. The apparatus according to claim 22, further comprising: An interaction heat determination control unit, configured to control the simulated digital human to determine the interaction heat of the interaction with the user; The interaction channel establishing unit is configured to establish an interaction channel between the user and the target real person imitated by the simulated digital human in response to the interaction heat exceeding a preset heat, so that the user can interact with the target real person.

24. The apparatus according to claim 14, wherein The timing includes at least one of the following: The total time the user browses all information during this information browsing behavior, the duration of the user's browsing of any information during this information browsing behavior, the total number of information browsed by the user during this information browsing behavior, the preset position of the browser presenting the waterfall-style information flow, and the end position of the information body presented by the browser.

25. The device according to any one of claims 14 to 24, wherein: The tentative interaction request initiation control unit is further configured to: Controlling each candidate digital person to generate a first round of interactive dialogue matching the real-time status information; Controlling each candidate digital human to generate an interactive guidance image including a corresponding first-round interactive dialogue and a target image matching the real-time status information; Each candidate digital human is controlled to initiate a tentative interaction request to the user at different times, with the interaction guide image as the request content.

26. The apparatus according to claim 25, further comprising: The incomplete consistency control unit is configured to control the dialogue contents of the first round of interactive dialogues generated by different candidate digital humans to be incompletely consistent.

27. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the immersive digital human interaction method based on multi-agent collaboration described in any one of claims 1-13.

28. A non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable the computer to execute the immersive digital human interaction method based on multi-agent collaboration as described in any one of claims 1-13.

29. A computer program product, comprising a computer program, which, when executed by a processor, implements the steps of the immersive digital human interaction method based on multi-agent collaboration according to any one of claims 1 to 13.

Citation Information

Patent Citations

  • Multi-modal interactive virtual digital human generation method and device, storage medium and terminal

    CN114495927A

  • Avatar selection of applications hosted on remote service environment

    CN116348861A