Autonomous agent construction method and device based on multi-modal large model

By fusing cross-modal projection, dynamic attention, and adaptive networks, and combining meta-learning and distributed inference architecture, the efficiency and accuracy issues of multimodal autonomous agents are solved, achieving efficient adaptation and real-time decision-making, and promoting the application of multimodal agents in complex scenarios.

CN120952042APending Publication Date: 2025-11-14XINZHIDI (BEIJING) TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511043202.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-28
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Existing multimodal autonomous intelligent agents suffer from insufficient modality fusion efficiency and accuracy, weak adaptive capabilities, and a prominent contradiction between computing resources and real-time performance, making it difficult to make stable decisions in complex scenarios.

Method used

By mapping cross-modal projections to a unified semantic space, using dynamic attention mechanisms and adaptive network models for weighted fusion, and combining meta-learning and continuous learning frameworks, along with model compression and distributed inference architecture, we can achieve efficient training and deployment of autonomous intelligent agents.

Benefits of technology

It improves the efficiency and accuracy of multimodal fusion, enhances the generalization and adaptability of intelligent agents, optimizes the utilization of computing resources, improves the naturalness and real-time performance of human-computer interaction, and is suitable for applications in multiple fields.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120952042A_ABST
    Figure CN120952042A_ABST
Patent Text Reader

Abstract

The invention discloses an autonomous agent construction method and device based on a multi-modal large model, and the method comprises the steps: obtaining multi-modal data, and carrying out the preprocessing of the multi-modal data; extracting basic features of each modal from the preprocessed multi-modal data, and mapping the basic features to a unified semantic space to generate semantic representation of unified dimensions; on the basis of semantic representation of a unified dimension, utilizing a dynamic attention mechanism to dynamically calculate an association weight of each modal, and performing weighted fusion through an adaptive network model to obtain an effective feature; training and optimizing the initial autonomous agent by using the effective features so as to obtain a trained autonomous agent; wherein when the initial autonomous agent is trained, a meta-learning stage and a continuous learning stage are included. According to the method, the multi-modal fusion efficiency and precision are remarkably improved through the dynamic attention and adaptive gating network, and the generalization and adaptive capabilities of the intelligent agent are improved based on a mixed training framework of meta learning and continuous learning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, specifically to a method and apparatus for constructing autonomous intelligent agents based on a multimodal large model. Background Technology

[0002] In recent years, multimodal large models have made significant progress in the field of artificial intelligence. By integrating data from multiple modalities such as text, images, speech, and video, multimodal large models can understand complex information from different dimensions and achieve intelligent interactions that are closer to human cognition. During training, multimodal large models employ pre-training strategies such as contrastive learning and masked language modeling, effectively improving the model's ability to represent multi-source data.

[0003] Autonomous agents are intelligent systems capable of perceiving their environment, making autonomous decisions, and executing actions. They are widely used in fields such as intelligent robots, autonomous driving, and intelligent customer service. Traditional autonomous agents typically build perception modules based on single-modal data (such as vision or text) and make decisions through preset rules or reinforcement learning algorithms.

[0004] With the rapid development of science and technology, some intelligent agents have begun to attempt to combine multimodal information to enhance their environmental understanding. Currently, there are two main methods for constructing multimodal autonomous intelligent agents: one is to input multimodal data into independent single-modal models and then integrate the results through a later fusion strategy; the other is to use a unified large multimodal model to process all data, but this method has high model training complexity and its inference efficiency is difficult to meet the requirements in scenarios with high real-time requirements. In addition, existing intelligent agents have shortcomings in autonomous learning and adaptive capabilities, and they are unable to flexibly adjust their strategies when faced with unseen tasks or environmental changes.

[0005] In summary, existing technologies suffer from insufficient efficiency and accuracy in modal fusion. Specifically, multimodal fusion methods often rely on fixed feature concatenation or weighted summation strategies, failing to dynamically capture deep semantic relationships between different modalities. For example, when processing image-text information, simple feature fusion struggles to understand the logical correspondence between complex scenes in images and text descriptions, leading to semantic biases in agent decision-making. Furthermore, the heterogeneity of multimodal data (such as differences in data format and time scale) further exacerbates the fusion difficulty, impacting overall model performance. Weak generalization and adaptive capabilities also exist. Traditional multimodal autonomous agents are typically trained on specific datasets, resulting in insufficient generalization when facing long-tail problems in the real world (such as rare scenes and noisy data).

[0006] Furthermore, intelligent agents lack mechanisms for dynamically updating knowledge and optimizing strategies, making it difficult for them to autonomously adjust their decision-making logic according to environmental changes, resulting in unstable performance in complex and open scenarios. The conflict between computational resources and real-time performance is prominent; multimodal large-scale models have a massive number of parameters, and the training and inference processes place extremely high demands on computational resources. Existing systems struggle to balance model performance and computational efficiency in edge devices or real-time interactive scenarios. Summary of the Invention

[0007] To address this, this application provides a method and apparatus for constructing autonomous intelligent agents based on multimodal large models, in order to solve the problems of low efficiency and accuracy of multimodal data fusion and poor adaptive ability of intelligent agents in the prior art.

[0008] To achieve the above objectives, this application provides the following technical solution:

[0009] Firstly, a method for constructing autonomous intelligent agents based on a multimodal large model includes:

[0010] Step 1: Acquire multimodal data and perform preprocessing; the multimodal data includes text data, image data, and voice data; the preprocessing includes data cleaning and standardization;

[0011] Step 2: Extract the basic features of each modality from the preprocessed multimodal data, and map the basic features of each modality to a unified semantic space through a cross-modal projection matrix to generate a semantic representation of a unified dimension.

[0012] Step 3: Based on the unified dimensional semantic representation, the association weights of each modality are dynamically calculated using a dynamic attention mechanism, and then weighted and fused through an adaptive network model to obtain effective features;

[0013] Step 4: Use the effective features to train and optimize the initial autonomous agent to obtain a trained autonomous agent; the training of the initial autonomous agent includes a meta-learning stage and a continuous learning stage; the optimization of the initial autonomous agent includes reinforcement learning environment exploration and strategy optimization.

[0014] Step 5: Compress the trained autonomous agent model and deploy it to various nodes through a distributed inference architecture.

[0015] Preferably, step 3, which involves weighted fusion using an adaptive network model, specifically includes generating a gating signal using a multilayer perceptron and performing weighted fusion processing in conjunction with task labels.

[0016] Preferably, when generating gating signals through a multilayer perceptron, for time-series data, a gating loop unit is used to adjust the fusion path step by step, and the state is updated by updating the gate and resetting the gate; for non-time-series data, an attention gating mechanism is used to locate the corresponding regions across modalities, and these regions are weighted and fused.

[0017] As a preferred option, in step 4, the initial autonomous agent is trained using a massive amount of task samples covering multiple domains and scenarios during the meta-learning stage, enabling the initial autonomous agent to master the ability to quickly learn new tasks, and to achieve efficient understanding and preliminary decision-making of unknown tasks through fine-tuning with small samples.

[0018] Preferably, in step 4, the initial autonomous agent adopts dynamic knowledge distillation and incremental learning algorithms during the continuous learning phase to prevent the initial autonomous agent from forgetting old knowledge when learning new tasks.

[0019] Preferably, the autonomous intelligent agent is also capable of performing structured analysis and emotion recognition on multimodal inputs using knowledge graphs and sentiment computing.

[0020] Secondly, an autonomous intelligent agent construction device based on a multimodal large model includes:

[0021] A multimodal data processing module is used to acquire multimodal data and perform preprocessing; the multimodal data includes text data, image data, and voice data; the preprocessing includes data cleaning and standardization;

[0022] The feature extraction module is used to extract the basic features of each modality from the preprocessed multimodal data, and to map the basic features of each modality to a unified semantic space through a cross-modal projection matrix to generate a semantic representation of a unified dimension.

[0023] The feature fusion module is used to dynamically calculate the association weights of each modality based on a unified dimension semantic representation, and then perform weighted fusion through an adaptive network model to obtain effective features.

[0024] The model training module is used to train and optimize the initial autonomous agent using the effective features, thereby obtaining a trained autonomous agent; the training of the initial autonomous agent includes a meta-learning stage and a continuous learning stage; the optimization of the initial autonomous agent includes reinforcement learning environment exploration and policy optimization.

[0025] The model deployment module is used to compress the trained autonomous agent model and deploy it to various nodes through a distributed inference architecture.

[0026] Thirdly, a computer device includes a memory and a processor, the memory storing a computer program, the processor executing the computer program to implement the steps of a method for constructing an autonomous intelligent agent based on a multimodal large model.

[0027] Fourthly, a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of a method for constructing an autonomous intelligent agent based on a multimodal large model.

[0028] Fifthly, a computer program product includes a computer program or instructions that, when executed by a processor, implement steps of a method for constructing an autonomous intelligent agent based on a multimodal large model.

[0029] Compared with the prior art, this application has at least the following beneficial effects:

[0030] This application provides a method for constructing an autonomous agent based on a multimodal large model. The method involves acquiring and preprocessing multimodal data; extracting basic features of each modality from the preprocessed multimodal data and mapping them to a unified semantic space to generate a unified-dimensional semantic representation; dynamically calculating the association weights of each modality using a dynamic attention mechanism based on the unified-dimensional semantic representation, and then weighted fusion through an adaptive network model to obtain effective features; training and optimizing an initial autonomous agent using these effective features to obtain a trained autonomous agent; the initial autonomous agent training includes a meta-learning stage and a continuous learning stage; the initial autonomous agent optimization includes reinforcement learning environment exploration and policy optimization; and the trained autonomous agent is model compressed and deployed to various nodes through a distributed inference architecture. This application significantly improves the efficiency and accuracy of multimodal fusion through dynamic attention and adaptive gating networks, and, based on a hybrid training framework of meta-learning and continuous learning, enables the agent to acquire rapid learning capabilities, thereby improving the agent's generalization and adaptive capabilities. Attached Figure Description

[0031] To more intuitively illustrate the prior art and this application, exemplary drawings are provided below. It should be understood that the specific shapes and structures shown in the drawings should not generally be regarded as limiting conditions for implementing this application; for example, based on the technical concept disclosed in this application and the exemplary drawings, those skilled in the art are able to easily make conventional adjustments or further optimizations to the addition / reduction / classification, specific shapes, positional relationships, connection methods, size ratios, etc. of certain units (components).

[0032] Figure 1 A flowchart illustrating a method for constructing an autonomous intelligent agent based on a multimodal large model, as provided in Embodiment 1 of this application;

[0033] Figure 2 A schematic diagram of the structure of an autonomous intelligent agent construction method based on a multimodal large model provided in Embodiment 1 of this application;

[0034] Figure 3 This is a schematic diagram of the interaction process between the autonomous intelligent agent and the user in the intelligent customer service product recommendation scenario provided in Embodiment 1 of this application. Detailed Implementation

[0035] The present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0036] In the description of this application: unless otherwise stated, "a plurality of" means two or more. The terms "first," "second," "third," etc., in this application are intended to distinguish the objects referred to and do not have any special meaning in terms of technical connotation (e.g., they should not be construed as an emphasis on importance or order). Expressions such as "comprising," "including," and "having" also mean "not limited to" (certain units, components, materials, steps, etc.).

[0037] The terms used in this application, such as "upper," "lower," "left," "right," and "middle," are generally used to indicate the general relative positional relationship for the purpose of intuitive understanding by referring to the accompanying drawings, and are not absolute limitations on the positional relationship in the actual product.

[0038] Example 1

[0039] Please see Figure 1 and Figure 2 This embodiment provides a method for constructing autonomous intelligent agents based on a multimodal large model, including:

[0040] S1: Acquire multimodal data and perform preprocessing; multimodal data includes text data, image data, and voice data; preprocessing includes data cleaning and standardization;

[0041] Specifically, this step involves collecting multimodal data such as text, images, and voice, and then cleaning and standardizing it to provide high-quality data for subsequent processing.

[0042] S2: Extract the basic features of each modality from the preprocessed multimodal data, and map the basic features of each modality to a unified semantic space through a cross-modal projection matrix to generate a unified semantic representation.

[0043] Specifically, after extracting the basic features of each modality in this step, they are mapped to a unified semantic space through a cross-modal projection matrix to generate a semantic representation of a unified dimension. This semantic representation contains semantic information shared across modalities, providing basic data for subsequent dynamic fusion.

[0044] S3: Based on a unified dimension of semantic representation, the association weights of each modality are dynamically calculated using a dynamic attention mechanism, and weighted fusion is performed through an adaptive network model to obtain effective features;

[0045] Specifically, this step first calculates the dynamic weights of each modality and the current task based on the generated unified semantic representation through a dynamic attention mechanism. Then, based on the output dynamic weights, an adaptive gating network is introduced to adjust the fusion path, solve the multimodal heterogeneity problem, and improve the semantic alignment accuracy.

[0046] Weighted fusion using an adaptive network model includes: generating gating signals using a multilayer perceptron and performing weighted fusion processing in conjunction with task labels; wherein, when generating gating signals using a multilayer perceptron, for time-series data, a gating loop unit is used to adjust the fusion path step by step, and the state is updated by updating the gate and resetting the gate; for non-time-series data, an attention gating mechanism is used to locate the corresponding regions across modalities, and these regions are then weighted and fused.

[0047] This step introduces a contrastive learning loss function during weighted fusion to force the feature similarity of matched modal pairs to be higher than that of unmatched pairs. Then, the gating network parameters are dynamically updated based on backpropagation to optimize semantic alignment accuracy.

[0048] S4: Train and optimize the initial autonomous agent using effective features to obtain a well-trained autonomous agent; the training of the initial autonomous agent includes a meta-learning stage and a continuous learning stage; the optimization of the initial autonomous agent includes reinforcement learning environment exploration and policy optimization.

[0049] Specifically, in the meta-learning phase, the initial autonomous agent is trained using massive task samples covering multiple domains and scenarios, enabling it to quickly learn new tasks. Fine-tuning with small samples further enhances its ability to efficiently understand and make initial decisions about unknown tasks. In the continuous learning phase, dynamic knowledge distillation and incremental learning algorithms are employed to prevent the agent from forgetting old knowledge when learning new tasks. The meta-learning phase allows the model to quickly master basic learning capabilities, while the continuous learning phase consolidates and expands its knowledge on more diverse data. This hybrid training framework of meta-learning and continuous learning improves the agent's generalization and adaptive capabilities.

[0050] This embodiment also incorporates the environment exploration mechanism in reinforcement learning, enabling the agent to autonomously discover and adapt to new scenarios during interaction with the environment, optimize decision-making strategies, and improve the model's decision-making and adaptability.

[0051] It should be noted that this embodiment pre-builds a structured knowledge base (RDS): an unstructured knowledge base is constructed using Alibaba Cloud's native structured RDS knowledge base capabilities. The corresponding management platform provides unstructured knowledge base capabilities, allowing administrators to upload manually provided feedback or reinforcement learning documents to the unstructured knowledge base for agents to learn from. A custom agent creation function is specifically added to the management backend. Administrators can flexibly create agents and configure independent knowledge base data for each newly created agent, thereby meeting diverse business needs and personalized management requirements for calling pre-trained multimodal large models.

[0052] S5: Compress the trained autonomous agent model and deploy it to various nodes through a distributed inference architecture.

[0053] Specifically, this embodiment addresses the conflict between computing resources and real-time performance by employing a lightweight model architecture and inference acceleration technology, proposing an optimization scheme that combines model compression and distributed inference. By designing a distributed inference architecture, model inference tasks are distributed to edge devices and the cloud for collaborative processing, achieving efficient utilization of computing resources. Furthermore, based on the model's task requirements, inference accuracy is dynamically adjusted, improving inference speed while ensuring accuracy for critical tasks.

[0054] After training and optimization, the model undergoes compression, distributed deployment, and dynamic accuracy adjustment in "model deployment and inference," followed by multimodal inference, policy optimization interaction, and human-computer feedback in "agent decision-making and interaction," thus achieving a closed loop from training to practical application. The entire training process relies on the quality of the initial data and feature processing, and subsequent deployment and interaction also feed back into model iteration, jointly supporting the development of multimodal intelligent models.

[0055] The autonomous intelligent agent construction method based on a multimodal large model provided in this embodiment also includes: the autonomous intelligent agent can use knowledge graphs and sentiment computing to perform structured analysis and sentiment recognition on multimodal inputs.

[0056] Specifically, this step involves constructing a deep reasoning engine based on a multimodal large model, integrating knowledge graphs and logical reasoning modules to enable the agent to perform structured analysis and complex reasoning on multimodal information. At the human-computer interaction level, affective computing and intent recognition technologies are introduced to accurately understand user intent in conjunction with multimodal input and generate natural, context-appropriate feedback. Reinforcement learning is used to optimize interaction strategies, improving the fluency of human-computer interaction and user experience.

[0057] The autonomous intelligent agent construction method based on a multimodal large model provided in this embodiment also features business database data isolation. Specifically, to ensure the security and independence of business data, this embodiment implements strict isolation of knowledge base data in the business database. Through data source information annotation, data uploaded by different entities can be clearly distinguished, thereby achieving effective isolation between data, preventing data confusion and leakage, and ensuring the integrity and confidentiality of data from each entity. Efficient isolation of intelligent agents in the cloud environment is achieved by creating multiple independent intelligent agents. Each intelligent agent is precisely associated with a specific knowledge base. This targeted association mechanism establishes clear boundaries for data access and interaction. In this way, data flow between different intelligent agents strictly follows predetermined rules, fundamentally eliminating mutual data interference and ensuring that the data processed and output by each intelligent agent has high accuracy and reliability.

[0058] This embodiment utilizes a method for constructing an autonomous intelligent agent based on a multimodal large model to build an autonomous intelligent agent applicable to intelligent customer service product recommendation scenarios. The complete processing flow of the autonomous intelligent agent after the user inputs their intent is as follows: Figure 3 As shown:

[0059] Users initiate requests via text ("I want a short-sleeved shirt") or voice; Multimodal preprocessing: The fusion layer module is invoked to convert the input into a unified semantic representation; Intent determination: A dynamic attention mechanism calculates weights based on the unified representation and combines them with the semantics related to "short-sleeved shirt" in the knowledge base to determine that the user's intent is "product recommendation"; Agent matching: The decision layer invokes a "product recommendation agent" trained by meta-learning and loads user historical data through a continuous learning module; Multimodal fusion reasoning: An adaptive gating network fuses user input and historical data (user browsing history) to generate recommendation features; Feedback generation: Combining sentiment analysis, a natural language response is generated ("I recommend 3 cotton short-sleeved shirts for you, suitable for summer wear. Would you like to see the details?") along with a product image to complete the interaction.

[0060] Therefore, this intelligent agent can analyze user-input dialogue, determine the user's intent, input task labels, and generate task feature vectors using a multilayer perceptron (MLP). It calculates the cosine similarity between the semantic representations of each modality and the task feature vectors, using this as initial attention weights. It then performs fine-grained analysis of the user-input dialogue (text / speech) or scene (image / video), focusing on key information to determine the user's intent. This step is a crucial decision point. Based on different dialogue intents, different working intelligent agents are matched to perform dialogue recommendations.

[0061] The autonomous intelligent agent construction method based on a multimodal large model provided in this embodiment has the following advantages:

[0062] I. Significantly Improved Efficiency and Accuracy of Multimodal Fusion

[0063] Existing technologies employ fixed fusion strategies, which struggle to handle the heterogeneity and semantic differences of modal data. In contrast, the hierarchical dynamic fusion mechanism constructed in this embodiment, through dynamic attention and adaptive gating networks, can adjust the weights and fusion paths of each modality in real time according to task requirements. This improves the accuracy of cross-modal semantic alignment and effectively avoids information redundancy and bias. In complex tasks such as text-based question answering and multimodal instruction parsing, the accuracy and response speed of the agent's decision-making are significantly better than traditional methods, providing a more reliable data foundation for subsequent intelligent decision-making.

[0064] II. Breakthroughs in Model Generalization and Adaptability

[0065] Traditional autonomous agents exhibit weak generalization ability when facing new scenarios or tasks, and are prone to "catastrophic forgetting." This embodiment is based on a hybrid training framework of meta-learning and continuous learning, enabling the agent to acquire rapid learning capabilities through meta-learning, and combining dynamic knowledge distillation and incremental learning algorithms to achieve continuous accumulation and updating of knowledge.

[0066] III. Efficient utilization of computing resources and real-time performance assurance

[0067] To address the conflict between the computational demands of multimodal large-scale models and real-time applications, this embodiment employs model compression and a distributed inference architecture. This distributed inference model, which integrates edge devices and the cloud, not only alleviates the computational burden on individual devices but also dynamically adjusts inference accuracy based on task urgency. In real-time intelligent customer service scenarios, it ensures accuracy for critical tasks while meeting millisecond-level response requirements, significantly improving system usability and deployment flexibility.

[0068] IV. Upgrading the Naturalness and Intelligence of Human-Computer Interaction

[0069] Existing intelligent agents struggle to understand complex user intentions in multimodal interactions, resulting in a stiff and unnatural user experience. This embodiment integrates knowledge graph and affective computing technologies, enabling the intelligent agent to perform structured analysis and emotion recognition on multimodal inputs, effectively improving the accuracy of intent understanding. Simultaneously, by optimizing interaction strategies through reinforcement learning, the intelligent agent can generate human-like feedback based on user emotions and context, significantly enhancing user trust and user experience, and driving the transformation of human-computer interaction from "functional interaction" to "natural and intelligent interaction."

[0070] V. Application Scenario Expansion and Technological Ecosystem Value

[0071] The autonomous intelligent agent technology system constructed in this embodiment can be quickly adapted to the needs of multiple fields such as smart healthcare, intelligent transportation, and industrial robots, reducing the technology migration costs for cross-scenario applications. Its modular design and open interfaces facilitate integration with existing systems, promote the ecological integration of multimodal artificial intelligence technologies, and provide a universal and efficient solution for the intelligent upgrading of industries, with significant economic benefits and social value.

[0072] In summary, the autonomous intelligent agent construction method based on a multimodal large model provided in this embodiment breaks through the limitations of traditional multimodal data fusion efficiency and accuracy, realizing deep dynamic association and efficient integration of cross-modal semantics; it enhances the generalization and adaptive capabilities of the intelligent agent in complex open scenarios, endowing it with a mechanism for autonomous learning and dynamic optimization of decision-making strategies; it balances the contradiction between the high computing power requirements of the multimodal large model and real-time application scenarios, optimizing the model architecture and inference algorithm to adapt to edge devices; it enhances the deep reasoning capability of the intelligent agent based on multimodal information, realizing natural, smooth and accurate human-computer interaction, and ultimately constructing an autonomous intelligent agent with strong perception, high intelligence and high adaptability, promoting the application of multimodal artificial intelligence technology in multiple fields.

[0073] Example 2

[0074] This embodiment provides an autonomous intelligent agent construction device based on a multimodal large model, including:

[0075] A multimodal data processing module is used to acquire multimodal data and perform preprocessing; the multimodal data includes text data, image data, and voice data; the preprocessing includes data cleaning and standardization;

[0076] The feature extraction module is used to extract the basic features of each modality from the preprocessed multimodal data, and to map the basic features of each modality to a unified semantic space through a cross-modal projection matrix to generate a semantic representation of a unified dimension.

[0077] The feature fusion module is used to dynamically calculate the association weights of each modality based on a unified dimension semantic representation, and then perform weighted fusion through an adaptive network model to obtain effective features.

[0078] The model training module is used to train and optimize the initial autonomous agent using the effective features, thereby obtaining a trained autonomous agent; the training of the initial autonomous agent includes a meta-learning stage and a continuous learning stage; the optimization of the initial autonomous agent includes reinforcement learning environment exploration and policy optimization.

[0079] The model deployment module is used to compress the trained autonomous agent model and deploy it to various nodes through a distributed inference architecture.

[0080] For details on the specific implementation of each module in the autonomous intelligent agent construction device based on a multimodal large model, please refer to the above description of the limitations of the autonomous intelligent agent construction method based on a multimodal large model, which will not be repeated here.

[0081] Example 3

[0082] This embodiment provides a computer device, including a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the steps of a method for constructing an autonomous intelligent agent based on a multimodal large model.

[0083] Example 4

[0084] This embodiment provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of a method for constructing an autonomous intelligent agent based on a multimodal large model.

[0085] Example 5

[0086] This embodiment provides a computer program product, including a computer program or instructions, which, when executed by a processor, implements the steps of a method for constructing an autonomous intelligent agent based on a multimodal large model.

[0087] The technical features of the above embodiments can be combined in any way (as long as there is no contradiction in the combination of these technical features). For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described; these embodiments not explicitly written should also be considered to be within the scope of this specification.

Claims

1. A method for constructing autonomous intelligent agents based on a multimodal large model, characterized in that, include: Step 1: Acquire multimodal data and preprocess it; the multimodal data includes text data, image data, and voice data; The preprocessing includes data cleaning and standardization; Step 2: Extract the basic features of each modality from the preprocessed multimodal data, and map the basic features of each modality to a unified semantic space through a cross-modal projection matrix to generate a semantic representation of a unified dimension. Step 3: Based on the unified dimensional semantic representation, the association weights of each modality are dynamically calculated using a dynamic attention mechanism, and then weighted and fused through an adaptive network model to obtain effective features; Step 4: Use the effective features to train and optimize the initial autonomous agent to obtain a trained autonomous agent; the training of the initial autonomous agent includes a meta-learning stage and a continuous learning stage; the optimization of the initial autonomous agent includes reinforcement learning environment exploration and strategy optimization. Step 5: Compress the trained autonomous agent model and deploy it to various nodes through a distributed inference architecture.

2. The method for constructing an autonomous intelligent agent based on a multimodal large model according to claim 1, characterized in that, Step 3, specifically the weighted fusion using an adaptive network model, includes: generating a gating signal using a multilayer perceptron and performing weighted fusion processing in conjunction with task labels.

3. The method for constructing an autonomous intelligent agent based on a multimodal large model according to claim 2, characterized in that, When generating gated signals using a multilayer perceptron, for time-series data, a gated loop unit is used to adjust the fusion path step by step, and the state is updated by updating the gate and resetting the gate; for non-time-series data, an attention gating mechanism is used to locate the corresponding regions across modalities, and these regions are weighted and fused.

4. The method for constructing an autonomous intelligent agent based on a multimodal large model according to claim 1, characterized in that, In step 4, the initial autonomous agent is trained using a massive amount of task samples covering multiple domains and scenarios during the meta-learning stage, enabling the initial autonomous agent to master the ability to quickly learn new tasks, and to achieve efficient understanding and preliminary decision-making for unknown tasks through fine-tuning with small samples.

5. The method for constructing an autonomous intelligent agent based on a multimodal large model according to claim 1, characterized in that, In step 4, the initial autonomous agent adopts dynamic knowledge distillation and incremental learning algorithms during the continuous learning phase to prevent the initial autonomous agent from forgetting old knowledge when learning new tasks.

6. The method for constructing an autonomous intelligent agent based on a multimodal large model according to claim 1, characterized in that, Also includes: The autonomous intelligent agent can use knowledge graphs and sentiment computing to perform structured analysis and sentiment recognition on multimodal inputs.

7. A device for constructing autonomous intelligent agents based on a multimodal large model, characterized in that, include: A multimodal data processing module is used to acquire multimodal data and perform preprocessing; the multimodal data includes text data, image data, and voice data. The preprocessing includes data cleaning and standardization; The feature extraction module is used to extract the basic features of each modality from the preprocessed multimodal data, and to map the basic features of each modality to a unified semantic space through a cross-modal projection matrix to generate a semantic representation of a unified dimension. The feature fusion module is used to dynamically calculate the association weights of each modality based on a unified dimension semantic representation, and then perform weighted fusion through an adaptive network model to obtain effective features. The model training module is used to train and optimize the initial autonomous agent using the effective features, thereby obtaining a trained autonomous agent; the training of the initial autonomous agent includes a meta-learning stage and a continuous learning stage; the optimization of the initial autonomous agent includes reinforcement learning environment exploration and policy optimization. The model deployment module is used to compress the trained autonomous agent model and deploy it to various nodes through a distributed inference architecture.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.

10. A computer program product, comprising a computer program or instructions, characterized in that, When the computer program or instructions are executed by a processor, they implement the steps of the method according to any one of claims 1 to 6.