Multi-agent collaboration method, device and equipment based on large model

By performing intention recognition and task allocation in a multi-agent system, parallel computing and modular collaboration are adopted, the complexity and efficiency problems of the multi-agent collaboration system are solved, efficient and flexible multi-agent collaboration is achieved, and user interaction experience and information processing capabilities are improved.

CN120297319APending Publication Date: 2025-07-11ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 0 Cites 6 Cited by

Patent Information

Application Number
CN202510447018.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing multi-agent collaboration system framework is complex, has a high usage threshold, loose structure, poor interface stability, difficult to handle multi-type output, low programming language efficiency, and difficult to achieve efficient multi-agent collaboration.

Method used

通过获取用户输入信息,进行意图识别,选取适当的智能体执行计算任务,并进行结果整合,采用并行计算架构和智能体池,实现模块化协作,动态资源优化和知识集成,支持多模态信息处理。

Benefits of technology

Reduce computing power consumption, realize precise task allocation, shorten processing delay, reduce system restructuring costs, improve knowledge integration efficiency, and enhance user interaction quality and output logical coherence.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120297319A_ABST
    Figure CN120297319A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a multi-agent collaboration method, device and equipment based on a large model. The scheme comprises the following steps: acquiring input information of a user for the interactive plot generation system; performing intention recognition according to the input information; selecting a plurality of specified agents from a plurality of preset agents according to the identification result; aiming at each appointed intelligent agent, respectively executing a respective corresponding calculation task; and performing integration according to the task execution result to obtain output information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of computer technology, and particularly to a multi-agent collaboration method, apparatus, and device based on large models. Background Art

[0002] With the development of computer and Internet technologies, large model technologies have been continuously iterated, and applications related to large models have gradually come into people's view.

[0003] In the field of large models, an agent can be understood as an application that makes full use of the capabilities of large models and combines other technologies (such as knowledge bases) to enhance its performance and adaptability in specific tasks or environments.

[0004] Compared with a single agent, multi-agents can define agents with different responsibilities, enabling each agent to specialize in knowledge and capabilities in a specific field, and effectively simulating complex real-world environments through interactions between agents, providing more advanced capabilities. Multi-agent collaboration can often handle complex logical operations such as dialogue interactions, role-playing, and emotional expressions.

[0005] Based on this, a simpler and more efficient multi-agent collaboration solution is needed for large models. Summary of the Invention

[0006] One or more embodiments of this specification provide a multi-agent collaboration method, apparatus, device, and storage medium based on large models to solve the following technical problem: A simpler and more efficient multi-agent collaboration solution is needed for large models.

[0007] To solve the above technical problem, one or more embodiments of this specification are implemented as follows:

[0008] A multi-agent collaboration method based on large models provided by one or more embodiments of this specification includes:

[0009] Obtain input information of a user for an interactive plot generation system;

[0010] Perform intent recognition based on the input information;

[0011] According to the recognition result, select several specified agents from a preset plurality of agents;

[0012] For each specified agent, execute its respective corresponding computing task;

[0013] Integrate according to the task execution results to obtain output information.

[0014] A multi-agent collaboration apparatus based on large models provided by one or more embodiments of this specification includes:

[0015] An input information acquisition module that acquires input information of a user for an interactive plot generation system;

[0016] An intention recognition module that performs intention recognition based on the input information;

[0017] An agent selection module that selects a number of specified agents from a plurality of preset agents according to the recognition result;

[0018] A computing task execution module that respectively executes respective corresponding computing tasks for each specified agent;

[0019] An output information determination module that integrates according to the task execution result to obtain output information.

[0020] A multi-agent collaborative device based on a large model provided by one or more embodiments of this specification, including:

[0021] At least one processor; and,

[0022] A memory communicatively connected to the at least one processor; wherein,

[0023] The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to:

[0024] Acquire input information of a user for an interactive plot generation system;

[0025] Perform intention recognition based on the input information;

[0026] Select a number of specified agents from a plurality of preset agents according to the recognition result;

[0027] Respectively execute respective corresponding computing tasks for each specified agent;

[0028] Integrate according to the task execution result to obtain output information.

[0029] A non-volatile computer storage medium provided by one or more embodiments of this specification, storing computer-executable instructions, and the computer-executable instructions are set to:

[0030] Acquire input information of a user for an interactive plot generation system;

[0031] Perform intention recognition based on the input information;

[0032] Select a number of specified agents from a plurality of preset agents according to the recognition result;

[0033] For each specified agent, execute their respective corresponding computing tasks separately;

[0034] Integrate according to the task execution results to obtain output information.

[0035] The above at least one technical solution adopted by one or more embodiments of this specification can achieve the following beneficial effects:

[0036] By dynamically matching appropriate specified agents through intent recognition, reduce ineffective computing power consumption and implement a precise task allocation mechanism. Through a parallel computing architecture, support multiple agents to concurrently execute different tasks and shorten the processing delay of complex scenarios. Achieve plug-and-play through a preset professional agent pool, implement a modular collaboration paradigm, and reduce the system reconstruction cost. Selectively activate agents based on real-time recognition results to achieve dynamic resource optimization and avoid full resource occupancy. Through a distributed agent collaboration mechanism, improve the knowledge integration efficiency and effectively aggregate the processing efficiency of vertical domain knowledge bases. The multi-dimensional task result fusion algorithm ensures the logical coherence and emotional consistency of the output information and enhances the interaction quality between users and the system. Brief Description of the Drawings

[0037] In order to more clearly illustrate the technical solutions in the embodiments of this specification or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in this specification. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0038] Figure 1 A flowchart of a multi-agent collaboration method based on a large model provided for one or more embodiments of this specification;

[0039] Figure 2 An architecture diagram of a traditional multi-agent collaboration system in an application scenario provided for one or more embodiments of this specification;

[0040] Figure 3 An architecture diagram of an interactive plot generation system in an application scenario provided for one or more embodiments of this specification;

[0041] Figure 4 A conversion diagram of a dependency relationship in an application scenario provided for one or more embodiments of this specification;

[0042] Figure 5 A processing diagram of multi-modal information in an application scenario provided for one or more embodiments of this specification;

[0043] Figure 6A schematic diagram of the hidden score calculation process in an application scenario provided for one or more embodiments of this specification;

[0044] Figure 7 A schematic structural diagram of a multi-agent collaboration device based on a large model provided for one or more embodiments of this specification;

[0045] Figure 8 A schematic structural diagram of a multi-agent collaboration device based on a large model provided for one or more embodiments of this specification. Detailed implementation manners

[0046] Embodiments of this specification provide a multi-agent collaboration method, device, equipment, and storage medium based on a large model.

[0047] In order to enable those skilled in the art to better understand the technical solutions in this specification, the following will clearly and completely describe the technical solutions in the embodiments of this specification with reference to the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all of the embodiments. Based on the embodiments of this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.

[0048] Figure 1 A schematic diagram of the process of a multi-agent collaboration method based on a large model provided for one or more embodiments of this specification. This method can be applied to different business fields, such as the large model business field, the AI interactive plot business field, the Internet finance business field, the e-commerce business field, the instant messaging business field, the game business field, the official business field, etc. This process can be executed by a computing device in the corresponding field (such as an application program or server corresponding to a large model agent formed based on a large model), and some input parameters or intermediate results in the process allow manual intervention and adjustment to help improve accuracy.

[0049] Figure 2 A schematic diagram of the architecture of a traditional multi-agent collaboration system in an application scenario provided for one or more embodiments of this specification, which can include multiple modes: Agent Customization, Multi-Agent Conversations, Flexible Conversation Patterns. In an actual traditional solution, one or more of these modes are often selected to form the corresponding multi-agent collaboration system.

[0050] In the agent customization mode, there is a Conversable Agent which can participate in natural language conversations. It is an agent with the ability to understand, generate, and respond, and realizes conversations with users through built-in large models and other modules or interfaces.

[0051] In the multi-agent conversation mode, multiple agents collaborate to participate in the same conversation scenario and complete complex tasks through division of labor or cooperation. At this time, the multiple agents can have equal conversations, or a corresponding master-slave relationship can be set up.

[0052] In the flexible conversation mode, various conversation structures are supported, such as Joint chat, Hierarchical chat, etc. The chat levels between each agent and other agents can be set based on requirements to cooperate and complete corresponding work.

[0053] However, the above traditional multi-agent collaborative systems all have corresponding problems:

[0054] 1. The framework is often relatively complex and requires a large amount of settings and configurations according to different tasks, with a certain threshold for use; 2. The structure is loose, affecting the usability of application implementation; 3. The interface stability is often not very stable, and there is a certain timeout situation for interface calls; 4. The work focuses on processing and generating text, and it is difficult to support multiple types of outputs (such as images, audio, video, etc.); 5. The programming language is usually Python, which will affect the R & D coding efficiency to a certain extent.

[0055] Based on this, the multi-agent collaboration method based on a large model as Figure 1 shown is proposed. The process in Figure 1 can include the following steps:

[0056] S102: Obtain the input information of the user for the interactive plot generation system.

[0057] In the interactive plot generation system (for the convenience of description, hereinafter simply referred to as the plot system), the user can interact with the plot system through means such as text, pictures, and voice, and put forward their own role-playing requirements. The plot system generates a virtual environment for the user in the form of text, pictures, etc. according to the user's role-playing requirements.

[0058] In the virtual environment, the user is responsible for playing a certain virtual character (for the convenience of description, hereinafter the virtual character played by the user is simply referred to as the protagonist) or other virtual entities (such as virtual animals, virtual objects, etc.). The plot system understands the user's intention based on the information input by the user, and thus simulates the behaviors, conversations, emotions, etc. of each virtual entity, continuously advancing the plot, enabling the user to achieve immersive role-playing in this virtual plot.

[0059] The types of virtual plots can include various ones. For example, according to different user role-playing requirements, virtual plots of different types such as real-world scenarios and fantasy scenarios are generated for users.

[0060] The input information of users can include multiple modalities, such as in the forms of text, voice, pictures (including static pictures, animated pictures, etc.), videos, etc. In some plot systems, some special buttons can also be set for users. By clicking on this button, users can perform actions such as moving, attacking, and conversing of the role they play in the virtual plot, and the plot system can display the execution results of this action for users in forms such as text and pictures.

[0061] S104: Perform intent recognition according to the input information.

[0062] Figure 3 It is a schematic diagram of the architecture of an interactive plot generation system in an application scenario provided for one or more embodiments of this specification. In this specification, multiple agents in the interactive plot generation system work together, so it can also be called a multi-agent cooperation system. The plot system supports graph representation to manage the connections between agents, providing a clear and extensible way to handle the interactions between multiple agents.

[0063] As Figure 3 shown, based on user needs, a user query is performed. The plot system receives the information input (Input) by the user through the task manager. Among them, the task manager is mainly used to coordinate agent task allocation and scheduling to ensure the efficient execution of the process.

[0064] An agent set is preset, as Figure 3 shown, in which multiple agents such as Agent 1, Agent 2, and Agent 3 are set, and different agents are used to perform different tasks.

[0065] After obtaining the input information, the process and rules for task execution are defined through the workflow. During the execution of each task, first, the corresponding agent needs to perform intent recognition on the user's input information. Intent recognition is mainly used to analyze the goals or needs input by the user. For example, the goals of the user can include querying the weather, interacting with virtual characters in the plot, generating emoticons, etc. Among them, the corresponding agent can be built-in with a corresponding large model or a trained deep learning model. Based on this model and combined with other modules in the agent (such as a text preprocessing module, an image preprocessing module, a large model call interface, etc.), the intent recognition of the user is realized.

[0066] S106: Select several specified agents from a plurality of preset agents according to the recognition result.

[0067] After obtaining the user's intention, based on the recognition result, select several agents from the pre-stored agent set as the specified agents.

[0068] As Figure 3 shown, assume that the current virtual scenario is a chat scenario between the user and the generated virtual character. At this time, the recognized user's intention is to understand the mood of the virtual character in the current virtual plot. Based on the process in the workflow, corresponding agents such as the plot arrangement agent, the emoji agent, the game state machine agent, and the chat agent can be selected to perform corresponding tasks respectively.

[0069] In actual work, agents with other functions can also be pre-built based on requirements. For example, agents for image generation, user preference analysis, calculation of the favorability between the user and the virtual character, etc. are used to perform corresponding calculation tasks.

[0070] Each time, according to the user's input information and the intention recognition result, the number and type of the selected specified agents may be different.

[0071] In different business scenarios, the framework selects appropriate agents from the predefined agent set to implement custom workflow orchestration. Through the scheduling mechanism, each agent is coordinated to efficiently complete its own task execution.

[0072] S108: For each specified agent, execute its corresponding calculation task respectively.

[0073] As Figure 3 shown, the plot arrangement agent mainly dynamically generates or adjusts the branch logic of the plot. For example, according to the user's intention recognition result and relevant information such as the recognized user's mood information, the interaction frequency and interaction scenarios between the protagonist and the current character in the follow-up can be adjusted.

[0074] The emoji agent is responsible for generating, recommending, or managing emoji content, and can increase the user's immersion by adding corresponding emojis when the virtual character interacts with the user.

[0075] The game state machine agent manages the game state (such as character attributes, plot progress) and drives state migration. Character attributes include the attributes of the protagonist or other virtual characters. For example, in a martial arts scenario, character attributes can include the user's health value, magic value, force value, etc. In a real-life scenario, character attributes can include the favorability and asset value between the protagonist and other virtual characters. These character attributes can be displayed or hidden.

[0076] The chat agent processes the natural language dialogue interaction between the user and the agent, making the conversation between the virtual character and the user smoother and more in line with the way humans speak, increasing the user's sense of immersion.

[0077] In actual work, each designated agent can communicate with each other based on needs or work completely independently during the execution of computing tasks. For example, the emoticon agent and the chat agent communicate with each other, and the emoticons generated by the emoticon agent are added to the conversation content generated by the chat agent, thereby increasing the fun of the conversation.

[0078] S110: Integrate the task execution results to obtain output information.

[0079] Different agents obtain corresponding task execution results for their respective computing tasks. The execution results of each task are integrated to obtain output information, which is then output to the user through the plot system. The form of output is also based on demand and can be set as a message or action. Its modality can include text, pictures, voice, etc., or multiple modalities can be selected and combined.

[0080] In addition, the plot system can be developed in JAVA, which is usually a programming language that business developers are more familiar with than Python, thereby reducing coding efficiency issues caused by the language.

[0081] Through intention recognition, the appropriate designated agent is dynamically matched to reduce the consumption of ineffective computing power and realize the precise task allocation mechanism. Through the parallel computing architecture, multiple agents are supported to execute differentiated tasks concurrently, shortening the processing delay of complex scenarios. Through the preset specialized agent pool, plug-and-play is realized, a modular collaboration paradigm is realized, and the cost of system reconstruction is reduced. Based on the real-time recognition results, the agent is selectively activated to realize dynamic resource optimization and avoid the occupation of all resources. Through the distributed agent collaboration mechanism, the efficiency of knowledge integration is improved, and the processing efficiency of the vertical field knowledge base is effectively aggregated. The multi-dimensional task result fusion algorithm ensures the logical coherence and emotional consistency of the output information, and enhances the quality of interaction between users and the system.

[0082] based on Figure 1 This specification also provides some specific implementation plans and extension plans of the method, which will be described below.

[0083] In one or more embodiments of the present specification, since the selection of the designated agents is based on different user input information, and different agents are selected, there is no strict hierarchical relationship or sequential processing relationship among the designated agents. Moreover, each designated agent, as a complete agent, can independently execute corresponding computing tasks. Therefore, when the designated agents execute computing tasks, multiple designated agents are supported to execute their respective corresponding computing tasks in parallel, avoiding the performance bottleneck caused by serial execution, thereby achieving fast response.

[0084] As Figure 3 shown, the short-term or long-term interaction data of the user (such as user behavior logs, role memories, etc.) is pre-stored in the memory module. And a Context Manager is provided to manage the context information of the conversation or task (such as historical records, user preferences, etc.).

[0085] During the execution of the computing tasks of the designated agents, for each designated agent, through context management, the corresponding context information is retrieved from the memory module, and according to the context information, the respective corresponding computing tasks are executed.

[0086] The Context Manager can directly interact with each designated agent, so that role memory and the sharing and exchange of other context information can be realized through context management. For different designated agents, seamless communication can be carried out with other designated agents based on their respective needs, enhancing the flexibility and consistency of collaboration.

[0087] In addition, as Figure 3 shown, after the execution of the computing tasks of the designated agents is completed, the task execution results can be used to update the context information recorded in the Context Manager.

[0088] In one or more embodiments of the present specification, although there is no strict sequence among the designated agents each time a designated agent is selected according to the user input information, there are still some dependencies among them.

[0089] Figure 4 This is a schematic diagram of the conversion of the dependency relationship in an application scenario provided for one or more embodiments of the present specification.

[0090] Among them, nodes A to G respectively represent different designated agents. There is a dotted-line connection between them, indicating a dependency relationship. For example, if node C points to node A with a dotted line, it means that when node C executes a computing task, it depends on the task execution result of node A. For example, if node C is a chat agent and node A is an emoji agent, the chat content finally output by node C depends on the emojis output by node A, and the chat content is finally combined.

[0091] To determine the dependency relationship, an agent specifically used for dependency analysis can be set. This agent combines static configuration and dynamic analysis to determine the dependency relationship between the selected designated agents this time. Among them, static configuration means that a corresponding dependency relationship table is preset, and this dependency relationship table stipulates the dependency relationships between all agents. Dynamic analysis means that this agent used for dependency analysis combines the context information input this time to dynamically analyze and adjust the dependency relationship between the obtained designated agents.

[0092] Since the selection of each designated agent is adaptively selected according to the user's input information, the dependency relationships between the designated agents are not fixed. If only relying on the dependency relationship, once an agent fails or malfunctions, it is likely to cause the failure of generating the output information this time.

[0093] Based on this, as Figure 4 shown, according to the dependency relationships between the designated agents, a level relationship between the designated agents is generated. The level relationship can also be called a hierarchical call relationship. For example, if node C depends on node A, then the level of node A is higher than that of node C, or it can be said that the hierarchical call level of node A is higher than that of node C, so as to finally generate the level relationship between each designated agent.

[0094] At this time, for each designated agent, its corresponding computing task is executed separately. For the designated agents with a hierarchical relationship, the computing tasks need to be executed in order. For some designated agents whose hierarchical relationships are not directly related (such as Figure 4 nodes B and C, nodes D and C in), the computing tasks can be executed in parallel.

[0095] During the execution of the computing tasks, each designated agent communicates with other designated agents based on the business requirements and according to the level relationship. This communication is not a mandatory requirement, but is carried out based on their respective business requirements. For example, the emoji agent and the chat agent need to conduct business communication to make the finally generated emojis and chat content more in line with the user's needs.

[0096] During the execution of a computing task, based on environmental requirements and according to the level relationship, a degradation strategy is executed. The degradation strategy refers to performing a degradation action on a specified agent when a failure or anomaly occurs in that agent. The degradation action can include: replacing the content output by it with a default value, skipping its corresponding dependency link, etc. By managing the dependencies between agents through the degradation strategy, the basic functions of the link can be maintained under specific conditions, ensuring the stability and reliability of the service, achieving decoupling between agents, and ensuring the high availability of agents.

[0097] Regard the workflow as a graph, where the nodes represent independent agents and the edges represent links. Managing agent connections through the graph representation provides a clear, intuitive, and flexible choreography collaboration ability, which has high availability and scalability in complex task processing.

[0098] Support the parallel execution of multiple agents, avoid the performance bottleneck caused by serial execution, and thus achieve fast response. At the same time, by managing the dependencies between agents through the degradation strategy, the basic functions of the link can be maintained under specific conditions, ensuring the stability and reliability of the service.

[0099] Furthermore, when each specified agent in the scenario system executes a computing task, the modalities of the data processed by different specified agents may be different, and the modal information of the final task execution result may also be different. For example, the emoticons output by the emoticon agent are usually in the picture modality, while the scenario choreography output by the scenario choreography agent is usually in the text modality. At this time, the task execution result corresponds to multiple modal information.

[0100] To further enhance the user's immersion during the scenario play, based on the communication capabilities between the specified agents, the task execution results of at least some of the specified agents can be personalized updated for the user, so that the final output information can be closer to the user, enhancing the user's immersion in the scenario system.

[0101] Based on this, determine whether communication occurs between the specified agent for processing the first preset modality and the specified agent for processing the second preset modality. The first preset modality and the second preset modality are preset modalities. The first preset modality is usually the text modality, which is convenient for updating information in other modalities. The second preset modality can be the picture modality, video modality, voice modality, etc.

[0102] If communication occurs, it means there may be a dependency relationship between the two, and then the task execution result corresponding to the second preset modality can be updated according to the task execution result corresponding to the first preset modality. At this time, the output information is obtained by integrating the updated task execution results.

[0103] For example, when the specified agent is an emoji agent, the corresponding task processing result is usually an emoji, and emojis are usually in the image modality, assuming it meets the limitations of the second preset modality. In the traditional solution, the emojis generated by the emoji agent are usually searched in a preset emoji library or on the Internet to obtain corresponding emojis, or simply by stacking several emojis to generate new emojis. Although it can reflect the current state of the virtual character, the degree of interaction with the user is still insufficient.

[0104] At this time, the generated emoji is updated through the text content in the first preset modality. The update methods can include: adding the text content to the emoji, adjusting the state of the target in the emoji according to the text content, etc. For example, the originally generated emoji is a cute expression, indicating that the current mood value of the virtual character the user is chatting with is happy. According to the text description of the plot in the plot arrangement agent, it can be known that the virtual character is happy because he is currently playing in the amusement park with the protagonist played by the user. At this time, through the text description in the first preset modality, the emoji in the second preset modality is updated, adding the content corresponding to the amusement park as the background to the cute expression, or adding corresponding text to indicate that it is very happy to play in the amusement park.

[0105] For another example, when the specified agent is a voice agent, which is used to convert the text content output by the chat agent into voice content, it can refer to the plot in the plot arrangement agent to adjust the tone and emotion of the output voice. It can also, according to the subsequent plot arrangement, after adjusting the emotion of the current voice, have a certain hint effect on the user. For example, in the subsequent plot arrangement, the protagonist and the virtual character need to be separated for a while, and the mood value of the virtual character will change from the current happiness to sadness. Assuming that the emotion corresponding to the voice at this time is happiness, then on the basis of this happiness, the tone can be gradually updated to simulate the state of the emotion changing from happiness to sadness step by step.

[0106] In one or more embodiments of this specification, when performing intent recognition based on input information, there may also be multiple modality information in the input information. For example, after the user inputs a piece of text and an emoji, the modality information at this time includes the text modality and the image modality.

[0107] At this time, if it is determined that there is multiple modality information in the input information, when directly performing intent recognition, it may be difficult to perform intent recognition due to the lack of unity between the multiple modality information.

[0108] Based on this, multiple modal information is mapped to a shared semantic space, and based on the shared semantic space, intent recognition is performed on the input information. For example, in an intent recognition agent, a pre-trained multi-modal model is preset, which can encode inputs such as text, images, and speech into a unified high-dimensional vector. At this time, the encoder can be optimized through contrastive learning to ensure that similar semantics of different modalities are close in distance in the shared semantic space.

[0109] Of course, since the confidence levels of different modalities may be different, weights for each modal information can be set based on requirements, and certain modal information can be trusted preferentially. In addition, quality inputs (such as blurred pictures, background noise, etc.) can be filtered or de-weighted.

[0110] At this time, in the shared semantic space, different modal information is unified, and similar modal information is closer in distance in the space, making it more convenient for user intent recognition.

[0111] Figure 5 This is a schematic diagram of the processing of multi-modal information in an application scenario provided for one or more embodiments of this specification. In real-life chats, there are various communication methods, not just single text conversations. In addition to text, we can also use emoticons, pictures, and voice to enrich our communication and make the communication more vivid and emotional.

[0112] Based on this, a multi-modal component is introduced to overcome the limitations of traditional single text models and process cross-modal tasks in a more efficient manner through the collaborative capabilities of agents.

[0113] The ToolAgent, as a part of the plot system, receives incoming custom tool parameters during the workflow orchestration phase, constructs corresponding tool call prompts, and hands them over to the large model for execution, and finally returns the tool decision results. The FunctionCallback Manager will execute corresponding function calls based on these decision results, such as sending emoticons or pictures.

[0114] Specifically, as Figure 5 shown, the Task Manager triggers the corresponding ToolAgent and corresponding agents (such as Figure 5 Agent A shown in

[0115] The tool agent receives the input, parses the custom tool parameters (such as emoji type, image resolution), constructs a tool call prompt, and submits the prompt to the large language model (LLM). The large language model analyzes the user's intention, generates corresponding LLM suggestions, and returns the tool decision result.

[0116] The function invoker (functionInvoke) receives the decision result, invokes the tool (execute toolinvoke), routes the tool through the function callback manager (FunctionCallback Manager) to the tool module (Tool), and the tool route selects the corresponding tool (such as JSON function, string function) and binds the execution logic.

[0117] In the tool module, the specific function is executed through function callback. For example, if generating an emoji, the JSON function is called. The JSON function has a pre-set implementation interface (implements), and the processing result is returned through the JSON function callback interface (JSONFunction Callback). Or, if processing an image, the string function is called to parse the file path and trigger the image processing service, and the processing result is returned through the string function callback interface (StringFunction Callback).

[0118] In one or more embodiments of this specification, when performing intention recognition, it is possible to perform intention recognition only based on the user's current input information and recent context information. For example, using the corresponding agent, keyword extraction is performed on the current input information and recent context information, marked through entity recognition, and sentiment analysis is performed, and then intention classification and mapping are performed, etc., to achieve the user's intention recognition.

[0119] However, it is difficult to classify the expressions of different users in this way. For example, some users are more introverted, and their expressions are relatively simple, which does not mean that they are not interested. And some extroverted users have more content in their expressions, which does not necessarily mean that they are interested.

[0120] Based on this, as Figure 2As shown, a long-term memory module is preset. In the components of the long-term memory module, corresponding agents are set. They are called through corresponding function calls and perform corresponding actions through prompt handles to obtain the user portrait corresponding to the user. The long-term memory module can record the user's historical interaction data (such as, common meme types, high-frequency conversation topics, plot selection preferences, etc.), and generate dynamic labels (such as, introverted users, mystery lovers, etc.) through large models or time series models. At the same time, combined with the results of multi-modal sentiment analysis, the user's emotion labels are determined, such as, optimistic, introverted, prone to anxiety, etc.

[0121] Taking these labels as part of the user portrait, a time decay factor can also be introduced to reduce the impact of past behaviors on the current portrait and avoid portrait rigidity. At the same time, the user portrait data needs to be stored anonymously, and users are provided with viewing and deletion permissions for privacy protection.

[0122] Synchronously with obtaining the user portrait, multiple modal information can also be cross-modally aligned based on the shared semantic space.

[0123] At this time, based on the obtained multiple modal information and the user labels corresponding to the user portrait, the input information is intention-recognized to obtain intention recognition results at multiple levels. When performing intention recognition in this way, the weights of modal information and others can be adjusted based on the user labels to obtain intention recognition results that are more in line with the user's situation.

[0124] Among them, the intention recognition results include multiple levels. For example, the intentions are divided into core intentions (including unlocking new plots) and secondary intentions (including expressing emotions, etc.), and the core intentions are processed with higher priority.

[0125] Figure 6 This is a schematic diagram of the hidden score calculation process in an application scenario provided by one or more embodiments of this specification. The agent router routes the calculation tasks to the specified agents, and the specified agents execute their respective corresponding calculation tasks. In order to smoothly promote the development of the plot and create a goal for the user in the plot role-playing, specific specified agents (referred to as the first specified agents here) can be set, which can perform hidden calculations based on the intention recognition results and in combination with the context information to obtain the hidden score between the user and the virtual entities that have been generated in the interactive plot generation system.

[0126] Among them, virtual entities include virtual characters, virtual events, virtual objects, etc. The hidden score is used to promote the plot development between the user and the corresponding virtual entity. For example, unlocking different virtual characters, generating different plots with virtual characters, etc. It can be directly presented to the user based on requirements, or hidden from the user. For example, when the hidden score between the user and a certain virtual entity reaches a threshold, the corresponding plot can be triggered.

[0127] The calculation of the hidden score can be obtained through the user's intention. For example, in the user's intention, if a corresponding positive action is performed on a certain virtual entity (such as sending a positive emoji, selecting a more positive option for the virtual character in some choices, etc.), the corresponding hidden score can be increased. Conversely, it may reduce the hidden score of the user for the virtual entity.

[0128] As Figure 6 shown, when calculating the hidden score, after the user answers the question, the plot is arranged by the specified intelligent agent through other channels, and the hidden score is calculated asynchronously for the recommendation of the next round of user questions.

[0129] Furthermore, in an application scenario, a plot game intelligent agent that focuses on "unlocking different characters through chatting" can be implemented. The user can interact with virtual characters through conversations, accumulate favorability as the hidden score, unlock new plots such as photo rewards, and start new character challenges after obtaining the rewards. The rewards that the user can unlock are mainly promoted by the hidden score.

[0130] When calculating the hidden score, the conversation content of the user can be scored in the background to quantify the intimacy and favorability between the user and the virtual character as the hidden score. The increase or decrease of the score directly affects the development trend of the subsequent plot, further enhancing the level and complexity of the interaction. In order to reduce the interference to the main link of the conversation, an asynchronous link is used to ensure the independence between the score calculation and the real-time user conversation, thus avoiding the impact on the smooth experience of the user.

[0131] However, the use of an asynchronous link for hidden score calculation and reward distribution may lead to a lag in plot unlocking or reward feedback, interrupting the user experience.

[0132] Based on this, when calculating the hidden score, determine the core intention and secondary intention in the intention recognition result; and determine the virtual entities that have been generated in the interactive plot generation system. As mentioned above, multiple levels are divided in the intention recognition result, and the core intention and secondary intention are determined here. Among them, there is usually only one core intention, while the secondary intention can include multiple ones.

[0133] At this time, according to the core intention, combined with the context information of the current input information, a hidden calculation is performed to obtain the first hidden score between the user and the virtual entity. The first hidden score is used to promote the plot development between the user and the corresponding virtual entity. Since the first hidden score only considers the core intention and the context information of the current input information, its calculation speed is relatively fast, and it can be called real-time lightweight scoring here. For real-time lightweight scoring, even for asynchronous link calculations, it can ensure that the time consumed is short, approximately producing real-time effects and reducing the lag in plot development.

[0134] Meanwhile, a hidden calculation can also be performed according to the core intention and secondary intention, combined with the context information of the long-term memory module, to obtain the second hidden score between the user and the virtual entity; the second hidden score is used to correct the plot development that has been promoted.

[0135] The first hidden score can basically reflect the user's intention and obtain a relatively accurate hidden score for promoting plot development. On the basis of the core intention, the second hidden score combines the secondary intention and also considers the context information of the long-term memory module to achieve a more accurate determination of the hidden score, which can be called asynchronous depth analysis scoring here.

[0136] Through asynchronous depth analysis scoring, the plot development that has been obtained is corrected, so that the subsequent plot can be more in line with the user's intention at the detail level. For example, the correction process can include fine-tuning the development speed of the plot development determined by the first hidden score, the characters appearing during the development process, etc. according to the second hidden score.

[0137] Furthermore, after the hidden score is calculated by the first designated agent, the second designated agent can be used to set the plot development according to the hidden score.

[0138] Through the second designated agent, according to the first hidden score, multiple virtual plot development frameworks are generated and pre-stored. The virtual plot development framework is not the specific plot development content, but mainly represents the general framework of the plot development. For example, the framework can include whether the user can interact with a certain virtual character again in the subsequent plot development, whether a certain virtual event may be triggered, etc. Each virtual plot development framework can be used as a plot development story line and pre-stored in advance, so that when the user hits a certain story line later, on the basis of this virtual plot development framework, further arrangement of plot details can be carried out.

[0139] At this time, the weights corresponding to the first hidden score and the second hidden score can be obtained according to the total usage duration of the user in the interactive plot generation system. Generally speaking, it is considered that the longer the total usage duration, the deeper the understanding of the user, and the higher the weight of the second hidden score. Corresponding upper limits can also be set. When the total usage duration reaches a certain upper limit value, it is considered that the user has been understood sufficiently, and the weight value of the second hidden score also reaches the corresponding upper limit value.

[0140] According to the first hidden score, the second hidden score, and the corresponding weights respectively, multiple virtual plot development frameworks are corrected. The higher the weight of the second hidden score, the higher the degree of considering the second hidden score when the first hidden score and the second hidden score conflict during correction. For example, the user profile of a certain user includes "likes to be sarcastic". Through the first hidden score, it is judged that the user's intention is to praise a certain virtual character, while through the second hidden score, it is judged that the user's intention is a superficial praise but actually sarcasm of the virtual character. When the weight of the second hidden score is relatively high, the judgment result of the second hidden score is more adopted, and the content related to the virtual character in the subsequent virtual plot development framework is corrected.

[0141] In one or more embodiments of this specification, during the process of the user's role-playing, as the usage duration increases and the plot progresses, the user is prone to fatigue.

[0142] Based on this, when specifying an intelligent agent to execute a computing task, the first plot progress of the user in the current first interactive plot is obtained, and the second plot progress of the user in other second interactive plots of the same type is obtained. The first interactive plot is the plot that the user is currently in, and the second interactive plot is the plot that the user has executed in the history. Each interactive plot can be set with a corresponding type label, such as, urban emotion type, martial arts type, etc.

[0143] If it is determined that the first plot progress and / or the second plot progress meet the preset progress, then through the third specified intelligent agent, according to the first plot progress, data reconstruction is performed and converted into a third plot progress that matches the second interactive plot.

[0144] The plot progress can be determined by the user's play duration, the number of conversations, the completion of goals, etc. When the first plot progress meets the preset progress, it means that the user has invested a lot of energy in the current first interactive plot and may feel fatigued. And when the second plot progress meets the preset progress, it means that the user has invested a lot of energy in the same type of plot before. At this time, even if the first plot progress does not meet the preset progress, the user may still feel fatigued.

[0145] The preset progress corresponding to the first plot progress and the second plot progress can be set to different progresses based on demand. When there are multiple second interactive plots, it can be determined whether the second plot progress meets the preset progress by calculating the average, selecting the highest plot progress, etc.

[0146] When it is determined that fatigue may occur, the data of the first plot progress is reconstructed through a third designated intelligent agent, and the first plot progress is mapped to the second interactive plot, allowing the user to travel from the first interactive plot to the second interactive plot, using a method similar to time travel to increase the freshness of the user's plot experience.

[0147] The second interactive plot can be a plot that the user has already experienced, thereby increasing the user's interest in the overall relationship between the plots, playing a similar effect to the "plot universe", increasing the user's sense of identity with the plot itself and the plot system. Of course, it can also be a currently popular plot that the user has not yet experienced, thereby further increasing the user's sense of freshness, and can also promote other plots to the user.

[0148] Since the first interactive plot and the second interactive plot are of the same type, the progress of the first plot can be converted into the progress of the third plot matching the second interactive plot by means of data reconstruction.

[0149] The scope of data reconstruction can be set based on demand. For example, it can be the attribute value, hidden points, virtual equipment and other contents of the protagonist. Since the two belong to the same type of plot, similar contents are usually set in the second interactive plot. For example, the "magic value" in the first interactive plot corresponds to the "energy value" in the second interactive plot after data reconstruction. The corresponding relationship during data reconstruction can be determined by comparing the similarity between the contents in the scope of data reconstruction and the contents in the second interactive plot.

[0150] For some content that is difficult to reconstruct data, it can be directly transplanted. For example, if the favorability between the protagonist and a virtual character in the first interactive plot is difficult to reconstruct data in the second interactive plot, the virtual character and the corresponding favorability can be directly transplanted to the second interactive plot, so that the protagonist and the virtual character in the first interactive plot perform corresponding tasks in the second interactive plot, which increases the user's sense of freshness.

[0151] Furthermore, in addition to the third designated agent, a fourth designated agent and a fifth designated agent may be set up to perform corresponding actions in the cross-plot arrangement of the first interactive plot and the second interactive plot, respectively.

[0152] Through the fourth designated intelligent agent, a first guiding plot is generated according to the progress of the first plot, and the user is guided from the first interactive plot to the second interactive plot through the first guiding plot.

[0153] Through the fifth designated intelligent agent, the user's actions in the first interactive plot and the second interactive plot are extracted and mapped, and a second guidance plot is generated according to the mapping relationship. Through the second guidance plot, the user is allowed to perform corresponding actions in the second interactive plot.

[0154] The fourth designated agent and the fifth designated agent do not have to be set, but can be set based on demand. Their main purpose is to guide the user so that the user can perform actions in the cross-plot arrangement more smoothly.

[0155] The first guiding plot is mainly used for text guidance, which can add a background story to the cross-plot arrangement, so as to make the user's crossing more reasonable. For example, the background story in the fantasy scene may include some secret places, time and space channels, etc. In the real scene, the protagonist may come to the area where the second interactive plot is located for some reasons such as tourism.

[0156] The second guidance plot is mainly to guide the user's operation. In the first interactive plot and the second interactive plot, when the user wants to execute the same command, the actions that need to be performed may be different. For example, when the user wants to perform the action of "flying", in the first interactive plot, it is only necessary to issue the command "use props to fly", while in the second interactive plot, it is necessary to issue the command "use mount to fly".

[0157] At this time, the fifth designated intelligent agent extracts the user's actions in the two interactive plots and maps them, establishing a mapping relationship between the instructions of the same target, thereby generating a second guided plot, which enables the user to perform corresponding actions in the second interactive plot through new instructions.

[0158] Based on the same idea, one or more embodiments of this specification also provide devices and apparatuses corresponding to the above methods, such as Figure 7 , Figure 8 shown.

[0159] Figure 7 A schematic diagram of a multi-agent collaboration device based on a large model is provided for one or more embodiments of this specification, and the device includes:

[0160] An input information acquisition module 702 is used to acquire user input information for the interactive scenario generation system;

[0161] An intention recognition module 704 performs intention recognition according to the input information;

[0162] The agent selection module 706 selects several specified agents from a preset plurality of agents according to the recognition result;

[0163] The computing task execution module 708 executes the respective corresponding computing tasks for each of the specified agents;

[0164] The output information determination module 710 integrates according to the task execution results to obtain the output information.

[0165] Optionally, the computing task execution module 708 executes the respective corresponding computing tasks for the several specified agents in parallel;

[0166] Wherein, during the execution of the computing task, for each specified agent, the corresponding context information is obtained, and the respective corresponding computing tasks are executed according to the context information.

[0167] Optionally, the computing task execution module 708 generates a level relationship between the specified agents according to the dependency relationship between the specified agents;

[0168] For each specified agent, the respective corresponding computing tasks are executed;

[0169] During the execution of the computing task, each specified agent communicates with other specified agents based on business requirements according to the level relationship;

[0170] And during the execution of the computing task, a degradation strategy is executed according to the level relationship based on environmental requirements.

[0171] Optionally, the output information determination module 710 determines whether to communicate between the specified agent for processing the first preset modality and the specified agent for processing the second preset modality if the task execution result corresponds to multiple modality information;

[0172] If communication is performed, the task execution result corresponding to the second preset modality is updated according to the task execution result corresponding to the first preset modality;

[0173] The output information is obtained by integrating according to the updated task execution results.

[0174] Optionally, the intention recognition module 704 determines that there are multiple modality information in the input information;

[0175] Map the multiple modality information to a shared semantic space;

[0176] Based on the shared semantic space, intention recognition is performed on the input information.

[0177] Optionally, based on the long-term memory module, the intention recognition module 704 obtains a user profile corresponding to the user;

[0178] Based on the shared semantic space, cross-modal alignment of the multiple modal information is performed;

[0179] According to the multiple modal information and the user tags corresponding to the user profile, intention recognition is performed on the input information to obtain intention recognition results at multiple levels;

[0180] For the first specified intelligent agent, the computing task execution module 708 performs hidden computing according to the intention recognition result in combination with the context information to obtain a hidden score between the user and the virtual entity that has been generated in the interactive plot generation system; the hidden score is used to promote the plot development between the user and the corresponding virtual entity.

[0181] Optionally, the computing task execution module 708 determines the core intention and secondary intention in the intention recognition result; and determines the virtual entity that has been generated in the interactive plot generation system;

[0182] According to the core intention, hidden computing is performed in combination with the context information of the current input information to obtain a first hidden score between the user and the virtual entity; the first hidden score is used to promote the plot development between the user and the corresponding virtual entity;

[0183] According to the core intention and the secondary intention, hidden computing is performed in combination with the context information of the long-term memory module to obtain a second hidden score between the user and the virtual entity; the second hidden score is used to correct the plot development that has been promoted.

[0184] Optionally, the computing task execution module 708 generates and pre-stores multiple virtual plot development frameworks according to the first hidden score through a second specified intelligent agent;

[0185] According to the total usage duration of the user in the interactive plot generation system, the weights corresponding to the first hidden score and the second hidden score are obtained;

[0186] According to the first hidden score, the second hidden score, and the corresponding weights respectively, the multiple virtual plot development frameworks are corrected.

[0187] Optionally, the computing task execution module 708 obtains the first plot progress of the user in the current first interactive plot, and obtains the second plot progress of the user in other second interactive plots of the same type;

[0188] If it is determined that the first plot progress and / or the second plot progress meet the preset progress, then through a third designated intelligent agent, according to the first plot progress, data reconstruction is performed and converted into a third plot progress that matches the second interactive plot.

[0189] Optionally, the computing task execution module 708 generates a first guiding plot through a fourth designated intelligent agent according to the first plot progress, and guides the user from the first interactive plot to the second interactive plot through the first guiding plot;

[0190] and / or,

[0191] Through a fifth designated intelligent agent, the actions performed by the user in the first interactive plot and the second interactive plot are extracted and mapped, and a second guiding plot is generated according to the mapping relationship. Through the second guiding plot, the user performs corresponding actions in the second interactive plot.

[0192] Figure 8 A schematic structural diagram of a multi-intelligent agent collaborative device based on a large model provided by one or more embodiments of this specification, the device includes:

[0193] At least one processor; and,

[0194] A memory communicatively connected to the at least one processor; wherein,

[0195] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can:

[0196] Obtain input information of the user for the interactive plot generation system;

[0197] Perform intent recognition according to the input information;

[0198] According to the recognition result, select several designated intelligent agents from a preset plurality of intelligent agents;

[0199] For each designated intelligent agent, execute its respective corresponding computing task;

[0200] Integrate according to the task execution results to obtain output information.

[0201] Based on the same idea, one or more embodiments of this specification also provide a non-volatile computer storage medium corresponding to the above method, storing computer-executable instructions, and the computer-executable instructions are set as:

[0202] Obtain input information of the user for the interactive plot generation system;

[0203] Perform intent recognition based on the input information;

[0204] Based on the recognition result, select several specified agents from multiple preset agents;

[0205] For each specified agent, execute its respective corresponding computing task;

[0206] Integrate according to the task execution result to obtain the output information.

[0207] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to circuit structures such as diodes, transistors, switches, etc.) or software improvements (improvements to method flows). However, with the development of technology, many method flow improvements today can be regarded as direct improvements to hardware circuit structures. Designers almost always obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that an improvement to a method flow cannot be implemented using a hardware entity module. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is an integrated circuit whose logical function is determined by the user programming the device. Designers can program themselves to "integrate" a digital system onto a single PLD, without having to ask a chip manufacturer to design and fabricate a dedicated integrated circuit chip. Moreover, today, instead of manually fabricating integrated circuit chips, this programming is mostly implemented using "logic compiler" software, which is similar to the software compilers used in program development and writing. The original code before compilation also has to be written in a specific programming language, which is called a Hardware Description Language (HDL), and there is not just one type of HDL, but many, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones currently are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also be aware that by simply performing a little logical programming on the method flow using the above-mentioned several hardware description languages and programming it into an integrated circuit, it is easy to obtain the hardware circuit that implements the logical method flow.

[0208] The controller can be implemented in any suitable manner. For example, the controller can take the form of, for example, a microprocessor or a processor and a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of the controller include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that, in addition to implementing the controller in the form of pure computer-readable program code, it is entirely possible to logically program the method steps to enable the controller to be implemented in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers, and embedded microcontrollers to achieve the same function. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be regarded as the structures within the hardware component. Or even, the devices for implementing various functions can be regarded as either software modules for implementing the method or the structures within the hardware component.

[0209] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.

[0210] For the convenience of description, when describing the above devices, they are described separately as various units according to their functions. Of course, when implementing this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0211] Those skilled in the art should understand that the embodiments of this specification can be provided as a method, a system, or a computer program product. Therefore, the embodiments of this specification can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of this specification can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program code.

[0212] This specification is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the specification. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and combinations of flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to produce a machine, such that the instructions executed by the processor of the computer or other programmable data processing device generate means for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or blocks.

[0213] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including instruction means that implement the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or blocks.

[0214] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or blocks.

[0215] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0216] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). Memory is an example of computer-readable media.

[0217] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0218] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0219] This specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.

[0220] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device, equipment, and non-volatile computer storage medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0221] The above description has been made of specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the drawings do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0222] The above is only one or more embodiments of this specification and is not intended to limit this specification. For those skilled in the art, various changes and modifications can be made to one or more embodiments of this specification. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of one or more embodiments of this specification shall be included within the scope of the claims of this specification.

Claims

1. A multi-agent collaboration method based on a large model, comprising: Obtaining input information of a user for an interactive plot generation system; Performing intent recognition according to the input information; According to the recognition result, selecting several specified agents from a plurality of preset agents; For each specified agent, respectively performing its corresponding computing task; Integrating according to the task execution results to obtain output information.

2. The method according to claim 1, for each specified agent, respectively performing its corresponding computing task, specifically comprising: For the several specified agents, parallelly performing their corresponding computing tasks; Wherein, during the execution of the computing task, for each specified agent, obtaining corresponding context information, and according to the context information, performing its corresponding computing task.

3. The method according to claim 1, for each specified agent, respectively performing its corresponding computing task, specifically comprising: Generating a level relationship between the specified agents according to the dependency relationship between the specified agents; For each specified agent, respectively performing its corresponding computing task; During the execution of the computing task, each specified agent communicates with other specified agents according to the level relationship based on business requirements; And during the execution of the computing task, based on environmental requirements, performing a degradation strategy according to the level relationship.

4. The method according to claim 3, integrating according to the task execution results to obtain output information, specifically comprising: If the task execution results correspond to multiple modal information, determining whether to communicate between the specified agent for processing the first preset modality and the specified agent for processing the second preset modality; If communication is performed, updating the task execution result corresponding to the second preset modality according to the task execution result corresponding to the first preset modality; Integrating according to the updated task execution results to obtain output information.

5. The method according to claim 1, performing intent recognition according to the input information, specifically comprising: Determining that there are multiple modal information in the input information; Mapping the multiple modal information to a shared semantic space; Performing intent recognition on the input information based on the shared semantic space.

6. The method according to claim 5, performing intent recognition on the input information based on the shared semantic space, specifically comprising: Obtaining a user portrait corresponding to the user based on a long-term memory module; Performing cross-modal alignment on the multiple modal information based on the shared semantic space; Performing intent recognition on the input information according to the multiple modal information and the user labels corresponding to the user portrait to obtain intent recognition results at multiple levels; For each specified agent, respectively performing its corresponding computing task, specifically comprising: For the first specified agent, performing hidden calculation according to the intent recognition result in combination with context information to obtain a hidden score between the user and the virtual entity already generated in the interactive plot generation system; The hidden score is used to promote the plot development between the user and the corresponding virtual entity.

7. The method according to claim 6, which performs hidden calculation by combining context information based on the intention recognition result to obtain a hidden score between the user and the virtual entity already generated in the interactive plot generation system, specifically including: Determine the core intention and secondary intention in the intention recognition result; And determine the virtual entities already generated in the interactive plot generation system; Perform hidden calculation according to the core intention by combining the context information of the current input information to obtain a first hidden score between the user and the virtual entity; The first hidden score is used to promote the plot development between the user and the corresponding virtual entity; Perform hidden calculation according to the core intention and the secondary intention by combining the context information of the long-term memory module to obtain a second hidden score between the user and the virtual entity; The second hidden score is used to correct the plot development that has been promoted.

8. The method according to claim 7, which respectively executes respective corresponding calculation tasks for each designated intelligent agent, specifically including: Through the second designated intelligent agent, generate multiple virtual plot development frameworks according to the first hidden score and pre-store them; Obtain the weights corresponding to the first hidden score and the second hidden score according to the total usage duration of the user in the interactive plot generation system; Correct the multiple virtual plot development frameworks according to the first hidden score, the second hidden score, and their respective corresponding weights.

9. The method according to claim 1, which respectively executes respective corresponding calculation tasks for each designated intelligent agent, specifically including: Obtain the first plot progress of the user in the current first interactive plot, and obtain the second plot progress of the user in other second interactive plots of the same type; If it is determined that the first plot progress and / or the second plot progress meet the preset progress, then through the third designated intelligent agent, perform data reconstruction according to the first plot progress and convert it into a third plot progress matching the second interactive plot.

10. The method according to claim 9, the method further includes: Through the fourth designated intelligent agent, generate a first guiding plot according to the first plot progress, and guide the user from the first interactive plot to the second interactive plot through the first guiding plot; And / or, Through the fifth designated intelligent agent, extract and map the actions performed by the user in the first interactive plot and the second interactive plot, generate a second guiding plot according to the mapping relationship, and make the user perform corresponding actions in the second interactive plot through the second guiding plot.

11. A multi-agent collaborative device based on a large model, including: An input information acquisition module, which acquires the input information of the user for the interactive plot generation system; An intention recognition module, which performs intention recognition according to the input information; An intelligent agent selection module, which selects several designated intelligent agents from a preset plurality of intelligent agents according to the recognition result; A calculation task execution module, which respectively executes respective corresponding calculation tasks for each designated intelligent agent; The output information determination module integrates based on the task execution result to obtain the output information.

12. The device according to claim 11, wherein the computing task execution module executes the respective corresponding computing tasks in parallel for the several specified agents; Among them, During the execution of the computing task, for each specified agent, the corresponding context information is obtained, and the respective corresponding computing tasks are executed according to the context information.

13. The device according to claim 11, wherein the computing task execution module generates a level relationship between the specified agents according to the dependency relationship between the specified agents; For each specified agent, the respective corresponding computing tasks are executed separately; During the execution of the computing task, each specified agent communicates with other specified agents according to the level relationship based on the business requirements. And during the execution of the computing task, a downgrading strategy is executed according to the level relationship based on the environmental requirements.

14. The device according to claim 13, wherein the output information determination module determines whether communication is to be performed between the specified agent for processing the first preset modality and the specified agent for processing the second preset modality if the task execution result corresponds to multiple modality information; If communication is to be performed, the task execution result corresponding to the second preset modality is updated according to the task execution result corresponding to the first preset modality; The output information is obtained by integrating according to the updated task execution result.

15. The device according to claim 11, wherein the intention recognition module determines that there are multiple modality information in the input information; The multiple modality information is mapped to a shared semantic space; Based on the shared semantic space, intention recognition is performed on the input information.

16. The device according to claim 15, wherein the intention recognition module obtains a user profile corresponding to the user based on a long-term memory module; Based on the shared semantic space, the multiple modality information is aligned across modalities; According to the multiple modality information and the user tags corresponding to the user profile, intention recognition is performed on the input information to obtain intention recognition results at multiple levels; The computing task execution module performs hidden calculation for a first specified agent according to the intention recognition result in combination with the context information to obtain a hidden score between the user and the virtual entity already generated in the interactive plot generation system; The hidden score is used to promote the plot development between the user and the corresponding virtual entity.

17. The device according to claim 16, wherein the computing task execution module determines the core intention and the secondary intention in the intention recognition result; and determines the virtual entity already generated in the interactive plot generation system; According to the core intention, hidden calculation is performed in combination with the context information of the current input information to obtain a first hidden score between the user and the virtual entity; The first hidden score is used to promote the plot development between the user and the corresponding virtual entity; Perform hidden calculations based on the core intention, the secondary intention, and the context information of the long-term memory module to obtain a second hidden score between the user and the virtual entity; The second hidden score is used to correct the plot development that has been promoted.

18. The device according to claim 17, wherein the calculation task execution module generates and pre-stores a plurality of virtual plot development frameworks through a second designated intelligent agent according to the first hidden score; Obtain the weights corresponding to the first hidden score and the second hidden score according to the total usage duration of the user in the interactive plot generation system; Correct the plurality of virtual plot development frameworks according to the first hidden score, the second hidden score, and the respective corresponding weights.

19. The device according to claim 11, wherein the calculation task execution module obtains the first plot progress of the user in the current first interactive plot and obtains the second plot progress of the user in other second interactive plots of the same type; If it is determined that the first plot progress and / or the second plot progress meet the preset progress, then through a third designated intelligent agent, perform data reconstruction according to the first plot progress and convert it into a third plot progress matching the second interactive plot.

20. The device according to claim 19, wherein the calculation task execution module generates a first guiding plot according to the first plot progress through a fourth designated intelligent agent, and guides the user from the first interactive plot to the second interactive plot through the first guiding plot; and / or Extract and map the user's executed actions in the first interactive plot and the second interactive plot through a fifth designated intelligent agent, generate a second guiding plot according to the mapping relationship, and make the user perform corresponding actions in the second interactive plot through the second guiding plot.

21. A multi-agent collaborative device based on a large model, comprising: At least one processor; And, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can: Obtain the input information of the user for the interactive plot generation system; Perform intention recognition according to the input information; Select several designated intelligent agents from a plurality of preset intelligent agents according to the recognition result; For each designated intelligent agent, execute its respective corresponding calculation task; Integrate according to the task execution results to obtain output information.

Citation Information

Cited By

  • Greenhouse gas emission analysis method and system based on multi-mode intelligent agent

    CN120892536A

  • Content generation method and device, equipment and storage medium

    CN120956985A

  • Intelligent airport data processing method and device and electronic equipment

    CN121031646A

  • Data transmission method in intelligent agent system, intelligent agent system, equipment and medium

    CN121334271A

  • Multi-agent and multi-tool cooperation method and storage medium

    CN121636077A