Intelligent agent architecture based on multi-modal large model
Through modular design and multimodal large model agent architecture, the shortcomings of existing agents in multimodal data processing, dynamic information management, personalized functions and task planning are solved, efficient and flexible task execution and personalized feedback are achieved, and the adaptability and execution efficiency of agents in complex scenarios are improved.
Patent Information
- Application Number
- CN202411928184.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-25
- Publication Date
- 2025-05-27
AI Technical Summary
The existing agent architecture has shortcomings in multimodal data processing, dynamic information management, personalized functions, task planning and tool matching, and it is difficult to meet the needs of complex application scenarios, resulting in insufficient perception capabilities, low task execution efficiency, poor adaptability and inaccurate feedback.
The modularly designed agent architecture adopts a modular design, including perception module, memory module, portrait configuration module, planning module, tool usage module and action module, combined with multimodal large models and reinforcement learning, realize multimodal data processing, dynamic information management, personalized configuration and precise tool matching. Through technical means such as Decode-Only Transformer architecture, attention mechanism, pre-trained large language model and improved topological sorting algorithm, the cross-modal understanding, task planning and execution capabilities of agents are improved.
It realizes unified processing and integration of multimodal data, dynamic information management and task context maintenance, personalized task configuration and dynamic adjustment, task decomposition and path optimization, accurate tool matching and resource optimization, improves the efficiency and adaptability of task execution, and provides personalized feedback and efficient task execution.
Smart Images

Figure CN120046645A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence, and more specifically to an intelligent agent architecture based on a multimodal large model. Background Art
[0002] Intelligent Agent is a term in the field of artificial intelligence (AI), which refers to a computer system or software entity that can independently perform tasks, make decisions and adapt to environmental changes. Therefore, intelligent agents play a very important role in the realization of AI functions. However, most existing intelligent agents are single-function artificial intelligence systems, which are difficult to meet the requirements of multi-capability combinations in complex application scenarios. With the rapid development of multimodal large models, intelligent agents can use the core capabilities of large models such as perception, reasoning, and memory to cope with diverse tasks. However, there is currently a lack of a general intelligent agent architecture that can efficiently integrate and coordinate these capabilities, especially in terms of multimodal data understanding, tool calling, and self-adaptation in complex environments. The current existing intelligent agent architecture faces the following problems in practical applications:
[0003] (1) Inadequate multimodal data processing: The current architecture has limited ability to integrate multimodal inputs (such as text, images, and voice), making it difficult to achieve efficient and unified internal representation, resulting in insufficient perception capabilities and affecting task understanding and execution.
[0004] (2) Imperfect dynamic information management: The memory module lacks flexibility in the storage and call of contextual information, making it difficult to balance the immediate needs of short-term tasks and the accumulation and review of long-term knowledge, resulting in insufficient task continuity and adaptability;
[0005] (3) Lack of personalized functions: Existing intelligent agents have weak adaptability to user needs and preference characteristics, and it is difficult to dynamically adjust the characteristics of intelligent agents, resulting in low accuracy of task output and low user satisfaction;
[0006] (4) Low efficiency of task planning: Task planning often relies on static algorithms or predefined rules, which cannot dynamically optimize task decomposition and path selection, and is difficult to adapt to complex and changing task requirements, which easily leads to low execution efficiency;
[0007] (5) Inaccurate tool matching and calling: The existing architecture lacks intelligent mechanisms for tool selection and resource calling, which easily leads to problems such as improper tool matching or resource waste, and the task cannot be completed efficiently;
[0008] (6) Difficulty in integrating technical modules: Many architectures are highly dependent on specific models or tools in module design, lack modular design, and are difficult to achieve flexible expansion and cross-domain application;
[0009] These problems limit the practical application of existing intelligent agents in complex scenarios and make it difficult to meet the demands for efficiency, intelligence and personalization. Summary of the invention
[0010] In view of the shortcomings of the prior art, the purpose of the present invention is to provide an intelligent agent architecture based on a multimodal large model, constructing a modular, intelligent, and dynamically adaptable intelligent agent architecture that can realize multimodal data processing, task decomposition and optimized execution, information management and personalized configuration and other functions in complex tasks, providing comprehensive technical support for efficient and automated processing of tasks, and being used to solve many key technical problems of intelligent agents in the process of complex task processing.
[0011] To achieve the above object, the present invention provides the following technical solution: an intelligent agent architecture based on a multimodal large model, comprising an external support module and an internal core module, characterized in that: the external support module comprises:
[0012] Input module, used to support multi-modal information reception, including text, image, video, audio and sensor modal data. Data sources include cameras, sensors, microphones and text input;
[0013] The output module is responsible for generating feedback information that users can understand, including text, voice, image and chart display;
[0014] The large model capability module is used to provide support for the internal core modules through the multimodal large model capability; the internal core modules include:
[0015] The perception module is based on the Decode-Only Transformer architecture and is combined with the subsequent portrait configuration module to transform multimodal input into a unified internal representation, thereby achieving cross-modal understanding and decision-making capabilities, laying the foundation for subsequent task processing;
[0016] Memory module, used to store and call dynamic information through attention mechanism;
[0017] The portrait configuration module is used to provide preference information to the perception module, subsequent planning module and action module by setting the personalized characteristics of the intelligent agent, and dynamically update the portrait configuration using reinforcement learning to ensure that the decision and behavior are consistent with the role characteristics of the intelligent agent;
[0018] Planning module: Based on the data of perception module, memory module and portrait configuration module, as well as the support of subsequent tool usage module, the planning module adopts pre-trained large language model and improved topological sorting algorithm to realize task decomposition and execution path formulation;
[0019] Tool usage module: Based on the subtask execution path provided by the planning module, the tool usage module combines semantic matching and reinforcement learning to achieve accurate tool matching and optimization, ensuring that the external tools or APIs required for task execution accurately correspond to the planned task path;
[0020] The action module combines the recommended tool information of the tool call module and the personalized data of the portrait configuration module, uses the large language model to generate accurate tool call instructions, and executes the task. At the same time, it adjusts the task execution process according to the execution feedback to ensure that the task is completed as expected and provide personalized feedback. As a further improvement of the present invention, the perception module is based on the Decode-Only Transformer architecture, combined with the subsequent portrait configuration module, to convert multimodal input into a unified internal representation, thereby realizing cross-modal understanding and decision-making capabilities, and laying the foundation for subsequent task processing. The specific steps are as follows:
[0021] Step 11: Input data. The data input in this step includes two parts: ① original data including text, images, videos, audio, and sensors; ② personalized data of the portrait configuration module, including user preferences, historical behaviors, and demand characteristics;
[0022] Step 12, perform data preprocessing. The specific preprocessing steps are as follows: clean the text data, remove meaningless symbols and punctuation, and perform word segmentation to ensure that the text content is structured and standardized; resize and color standardize the image data to unify the format and style of the image; split the video data by frame and retain the position information to ensure the integrity and timing of the video sequence; sample the audio data to adjust the frequency of the signal and remove noise to improve the audio quality; perform noise reduction, missing value filling, anomaly detection and filtering on the sensor data to ensure the accuracy and reliability of the data;
[0023] Step 13, perform feature encoding, extract features from the data preprocessed in step 12, and map the multimodal data into a unified vector representation. The specific feature extraction steps are as follows: for text feature extraction, a pre-trained language model is used to generate a context-sensitive feature vector; for image feature extraction, the image is cut into fixed-size tiles based on ViT (Vision Transformer), and each tile is input as a sequence into the Transformer encoder to extract the global feature vector representation. Finally, the image feature dimension is adjusted to be consistent with the text input through linear mapping; for video feature extraction, ViT is used to perform feature encoding on the key frame image and convert it into a vector representation; for audio feature extraction, an audio feature extractor is used to generate a dense vector for representing the audio content; for sensor feature extraction, feature selection and dimensionality reduction methods are used to extract in order to retain the most discriminative features. Finally, each of the above extracted features will be converted into a fixed-length numerical feature vector for subsequent analysis and processing;
[0024] Step 14: perform multimodal feature alignment, and standardize the feature vectors of different modalities into a unified representation space by sharing encoding dimensions or performing linear mapping;
[0025] Step 15, based on the feature alignment in step 14, multimodal feature fusion is performed, specifically, the Decode-Only Transformer architecture is used to further process these multimodal feature vectors, and further specifically, the context-related deep features are gradually extracted through the decoding layer of the Transformer, and the mutual influence between different modalities is captured through the self-attention mechanism of the Transformer, so as to generate a unified internal representation; Step 16, output data, specifically, output a unified multimodal semantic representation vector, which integrates text, image, and audio features, has consistency, context perception and multimodal semantic fusion capabilities, and provides deep semantic support and personalized optimization for tasks such as retrieval, classification, and generation.
[0026] As a further improvement of the present invention, the memory module comprises:
[0027] A memory storage module, which includes a short-term memory storage layer and a long-term memory storage layer, which uses a high-speed cache storage based on Attention weighting through the short-term memory storage layer to dynamically evaluate the importance of task information and only store information with high weights, and uses a vector database or a knowledge graph through the long-term memory storage layer to store long-term memory data;
[0028] Data indexing and retrieval module, which includes a semantic vectorization and attention focusing module and a context awareness and association indexing module. The module uses natural language processing technology to extract key feature words and encodes them through the Transformer model to generate semantic vectors. The context awareness and association indexing module uses the Attention model to dynamically calculate the weight of memory data in the current context, generate a context index for the memory content, and mark the task type, time, and environmental status features to quickly locate highly relevant data.
[0029] Dynamic review and reinforcement module, which includes a dynamic review module and a reinforcement module. The dynamic review module uses the self-attention mechanism to periodically review short-term memory, screen high-weight data and transfer them to the long-term memory layer. After the task is completed, the reinforcement module uses reinforcement learning combined with attention weight to evaluate the effectiveness of memory data;
[0030] A memory call module, which includes a contextual memory call module, a real-time memory call module, and a multi-task sharing and inheritance module. When the task is initialized, the contextual memory call module uses the Cross-Attention mechanism to retrieve long-term memory based on the task context, screens highly relevant experience data, and provides guidance for the current task. During the task execution process, the real-time memory call module dynamically calls short-term and long-term memories, calculates the relevance of data in real time through the self-attention mechanism, assists in decision-making, and performs multiple rounds of attention weight updates on short-term memory to ensure that the latest and relevant data are called. Through the multi-task sharing and inheritance module, memory sharing and knowledge transfer are realized through the multi-head attention mechanism in a multi-task environment. As a further improvement of the present invention, the portrait configuration module sets the personalized characteristics of the intelligent agent, and uses reinforcement learning to dynamically update the portrait configuration. The specific steps are as follows:
[0031] Step 31, perform portrait modeling and feature definition, specifically based on the pre-trained language model, combined with the label generation algorithm, directly generate the attribute feature embedding of the agent through natural language description, so as to set the basic attributes and contextual relationships of the agent to construct the characteristics of the agent;
[0032] Step 32, perform knowledge graph driven feature association, specifically, first use knowledge graph technology to associate the multi-dimensional features of the agent with its role tasks, the associated feature nodes include skills, language style, task preferences, and the associated relationship edges represent the constraints and priorities between the features, and then use graph embedding technology to dynamically calculate the matching degree between the role features of the agent and the context, to ensure that the behavior of the agent meets the scene requirements;
[0033] Step 33, generating a multimodal semantically driven portrait, specifically generating a feature portrait through semantic alignment based on multimodal clues configured based on preferences;
[0034] Step 34, optimizing the image features based on reinforcement learning, specifically using deep reinforcement learning to adjust feature priorities in real time.
[0035] As a further improvement of the present invention, the planning module adopts a pre-trained large language model and an improved topological sorting algorithm to implement the specific steps of task decomposition and execution path formulation as follows:
[0036] Step 41, inputting data from the memory module, the perception module, the portrait configuration module and the tool use module, specifically, the memory module provides historical task records and feedback, the perception module provides real-time environment status, the portrait configuration module provides personalized needs, and the tool use module provides recommended tools and their characteristics;
[0037] Step 42, task decomposition and dependency identification based on the large language model, specifically, parsing the user task through the pre-trained large language model, identifying the task objectives and requirements, and using natural language processing technology to enable the model to extract key elements of the task and convert them into executable subtasks; the specific objectives of the task are clarified through the subtasks, and the dependency relationships between the subtasks are also revealed;
[0038] Step 43, performing execution path calculation and optimization based on an improved topological sorting algorithm;
[0039] Step 44, outputting data, specifically outputting the planned task execution queue, and determining the specific subtask execution order.
[0040] As a further improvement of the present invention, the improved topological sorting algorithm in step 43 includes the following steps:
[0041] Step 431, constructing a task graph and an in-degree table, specifically by traversing all tasks, constructing a task dependency graph G, and calculating the in-degree of each task;
[0042] Step 432, find the task with in-degree 0, specifically put all the tasks with in-degree 0 in graph G into container Vec, and merge them into one large node and put it into the task execution queue;
[0043] Step 433, perform topological sorting, specifically, take out all the nodes in Vec, reduce the in-degree of their adjacent points by 1 in turn, then put all the adjacent points whose in-degree becomes 0 into Vec, and merge them into 1 large node and put them into the task execution queue, repeat step 433 until no new nodes can be added to Vec; step 434, perform loop detection, specifically, check whether there are still nodes with in-degree not 0, if so, it means there is a loop and the task cannot be completed, otherwise, the task execution queue containing all tasks can be obtained. In addition, for large nodes containing multiple tasks in the task execution queue, if the resources required for execution are sufficient, the execution order of the tasks contained in the node is not restricted, and the overall efficiency can be improved through concurrent execution, otherwise priority sorting can be performed according to the required resource status.
[0044] As a further improvement of the present invention, the specific steps of the tool usage module to achieve accurate tool matching and optimization are as follows:
[0045] Step 51, input data, specifically receiving the task execution path and related tool description data recorded in the task execution queue output by the planning module, to provide basic data for subsequent processing;
[0046] Step 52, constructing a subtask execution tool query statement, specifically, based on the generation capability of the pre-trained large language model, generating a tool matching query statement according to the task to be executed and the tool description template structure;
[0047] Step 53, matching the task execution tool based on the similarity algorithm, specifically, converting the tool matching query statement and the tool description into a vector representation through the pre-trained language model, each query and tool description will be converted into a vector of fixed dimension, and then the similarity algorithm is used to calculate the similarity between the tool query statement and the tool description, so as to achieve the matching of the subtask and the target tool;
[0048] Step 54, selecting the result based on the reinforcement learning optimization tool, specifically, after the semantic matching module gives the preliminary matching result, using reinforcement learning, optimizing the matching strategy through user feedback and historical data, and continuously adjusting the matching result based on the reward and penalty mechanism, which can be adopted in the following ways:
[0049] Weighted fusion: The similarity score of semantic matching and the reward score of reinforcement learning are weightedly combined to generate the final matching result;
[0050] Feedback-based correction: If there is a large difference between the feedback from the reinforcement learning module and the result of semantic matching, the system will adjust the weight of semantic matching to give priority to the optimization result of reinforcement learning;
[0051] Hierarchical decision-making: first perform semantic matching to obtain a batch of candidate tools, and then use the reinforcement learning model to select the best tool from them;
[0052] Finally, the matching fusion module combines the semantic matching results and the optimization feedback of the reinforcement learning module to generate the final matching tool recommendation;
[0053] Step 55, combined with the pre-trained large language model, outputs the subtask execution tool matching results, provides matching reasons and scores, helps users understand the basis of the recommendation, and generates a tool calling plan for use in the action module.
[0054] As a further improvement of the present invention, the action module generates precise tool call instructions and executes tasks, and adjusts the task execution process according to execution feedback to ensure that the task is completed as expected and provide personalized feedback. The specific steps are as follows:
[0055] Step 61, input data, which includes two parts: recommended tool information provided by the tool calling module, including the most suitable tool matched by the algorithm, as well as the basic functions and applicable scenarios of these tools, and personalized data of the portrait configuration module, including user preferences, historical behaviors, and demand characteristics;
[0056] Step 62, generating a tool calling instruction, specifically, using a pre-trained large language model to generate the tool calling instruction based on the tool information and the personalized data of the profile, further specifically, first analyzing the applicability of the recommended tool, and then generating accurate task execution instructions based on the preferences and needs in the user profile;
[0057] Step 63, calling the tool to execute the task, specifically, after the task execution instruction is generated in step 62, calling the tool according to the instruction to gradually complete the specified task, and during the tool execution, the action module monitors the task progress in real time to ensure that the tool is executed smoothly according to the instruction; if any deviation or problem occurs during the task execution, the action module will make adjustments based on the execution feedback to ensure that the task is completed as expected;
[0058] Step 64, summarize all execution results, combine with the user preference information in the portrait configuration module, and generate the final feedback content through the pre-trained large language model.
[0059] Beneficial effects of the present invention:
[0060] (1) Unified processing and integration of multimodal data
[0061] During the task execution process, users may input data in various forms (such as text, voice, image, video, etc.). Traditional architectures are difficult to effectively integrate these multimodal inputs. The present invention introduces the Decode-Only Transformer architecture through the perception module to achieve a unified representation of multimodal data, providing an efficient basis for subsequent task processing.
[0062] (2) Dynamic Information Management and Task Context Preservation
[0063] The intelligent agent needs to dynamically adjust its behavior in complex tasks by combining the current task goal with the historical context (such as user needs and task progress). To this end, the present invention uses an attention mechanism through a memory module to achieve the storage and call of short-term and long-term memory, as well as the memory review function. This dynamic information management mechanism ensures the continuity and adaptability of task execution.
[0064] (3) Personalized task configuration and dynamic adjustment
[0065] User needs are diverse and personalized (such as style preferences, specific task settings). Traditional agents are difficult to dynamically adapt to different user needs. This invention uses a portrait configuration module, based on reinforcement learning and preference settings, to achieve personalized configuration and dynamic update of agent characteristics, ensuring that its behavior and decision-making conform to preset role characteristics and user expectations.
[0066] (4) Task decomposition and path optimization
[0067] When faced with complex tasks, traditional methods are inefficient in task decomposition and execution path planning, and have difficulty handling dynamic changes. The planning module of the present invention combines a pre-trained large language model and an improved topological sorting algorithm to efficiently decompose tasks, develop optimized execution paths, and flexibly respond to possible changes in tasks, thereby improving the accuracy and efficiency of task planning.
[0068] (5) Accurate matching of tools and optimized use of resources
[0069] The execution of complex tasks often relies on a variety of external tools. How to match the right tools and use resources efficiently is one of the challenges faced by intelligent agents. This invention uses a tool usage module, combined with a pre-trained large language model and a vector similarity algorithm, to accurately match tools and ensure that the configuration of the called tool is highly consistent with the task requirements, thereby improving the success rate of task execution and resource utilization efficiency.
[0070] (6) Efficient task execution and dynamic adjustment
[0071] The actual execution of tasks may be affected by user feedback or changes in external conditions, and traditional architectures are difficult to adjust in real time. This invention uses action modules to efficiently execute tasks according to the planned path, and can dynamically adjust task strategies based on user feedback and profile configuration to ensure high-quality completion of tasks.
[0072] (7) Information feedback and system connection
[0073] The task results of the agent need to be fed back in a way that the user can understand and support docking with external systems. To this end, the present invention provides a variety of information feedback forms (such as text, charts, reports, etc.) through the output module, and supports exporting, sending or integrating into external systems to enhance the applicability of the agent. BRIEF DESCRIPTION OF THE DRAWINGS
[0074] Figure 1 It is a module block diagram of the intelligent agent architecture based on the multimodal large model of the present invention;
[0075] Figure 2 This is a flow chart of the perception module;
[0076] Figure 3 Schematic diagram of the multimodal feature fusion architecture based on Decoder-Only Transformer in the perception module;
[0077] Figure 4 This is a schematic diagram of the overall architecture of the memory module;
[0078] Figure 5 It is a flowchart of the portrait configuration module;
[0079] Figure 6 is the task dependency graph G;
[0080] Figure 7 This is a diagram of the task execution queue;
[0081] Figure 8 A flowchart of the tool usage module;
[0082] Fig. 9 The flowchart of the action module is shown in Figure 2. DETAILED DESCRIPTION
[0083] The present invention will be further described below in detail with reference to the embodiments shown in the accompanying drawings.
[0084] Reference Figure 1 As shown, the intelligent agent architecture based on the multimodal large model of this embodiment adopts a modular design and is set as 6 internal core modules and 3 external support modules. The internal core modules include perception module, memory module, portrait configuration module, planning module, tool use module and action module; the external support module consists of input module, output module and large model capability module, and the specific structure is as follows: External support module:
[0085] (1) Input module: In this architecture, the input module supports multimodal information reception, mainly including text, image, video, audio, and sensor data. The data sources include cameras, sensors, microphones, and text input.
[0086] (2) Output module: The output module is responsible for generating feedback information that users can understand and supports multiple output methods, including text, voice, image, and chart display. The output module can provide real-time feedback of deeply processed response information to meet user query, task execution, and other needs;
[0087] (3) Big model capability module: Through multimodal big model capabilities, it provides support for perception, memory, image configuration, planning, tool use, and action modules.
[0088] Internal core modules:
[0089] (1) Perception module: Based on the Decode-Only Transformer architecture and combined with the portrait configuration module, it transforms multimodal input into a unified internal representation, thereby achieving cross-modal understanding and decision-making capabilities, laying the foundation for subsequent task processing;
[0090] (2) Memory module: It realizes dynamic information storage and retrieval through the attention mechanism, supports short-term and long-term memory functions, and can also perform memory review, provide context support for task planning, and ensure the consistency and flexibility of task execution;
[0091] (3) Portrait configuration module: By setting the personalized characteristics of the intelligent agent, it provides preference information to the perception, planning, and action modules, and uses reinforcement learning to dynamically update the portrait configuration to ensure that the decision and behavior are consistent with the role characteristics of the intelligent agent;
[0092] (4) Planning module: Based on the data from the perception, memory, and image configuration modules, as well as the support from the tool usage module, the pre-trained large language model and the improved topological sorting algorithm are used to achieve task decomposition and execution path formulation;
[0093] (5) Tool usage module: Based on the subtask execution path provided by the planning module, accurate tool matching and optimization are achieved through a combination of semantic matching and reinforcement learning, ensuring that the external tools or APIs required for task execution accurately correspond to the planned task path;
[0094] (6) Action module: By combining the recommended tool information of the tool call module and the personalized data of the profile configuration module, the large language model is used to generate accurate tool call instructions and execute tasks. At the same time, the task execution process is adjusted according to the execution feedback to ensure that the task is completed as expected and provide personalized feedback.
[0095] It can be seen that this intelligent agent architecture ensures the efficiency, consistency and personalized response of task execution through multimodal input, personalized portraits and large model support, combined with precise task planning and tool calling.
[0096] The above modules are further described in detail below with reference to the accompanying drawings:
[0097] The perception module of this embodiment is as follows:
[0098] The perception module of the present invention is based on the Decode-Only Transformer architecture, focusing on the efficient parsing and internal representation transformation of multimodal inputs, and can unify the inputs of multiple data sources such as text, images, audio and video into processable internal representations, thereby achieving cross-modal understanding and decision-making capabilities. At the same time, combined with the data of the portrait configuration module, more accurate task preference information can be obtained. The specific technical route is as follows Figure 2 As shown:
[0099] (1) Input data
[0100] The input data consists of two parts: ① original data of modalities such as text, images, videos, audio, sensors, etc.; ② personalized data of the portrait configuration module, which includes user preferences, historical behaviors, demand characteristics, etc.
[0101] (2) Data preprocessing
[0102] To ensure the efficiency and consistency of subsequent processing, these data need to be preprocessed and normalized. These preprocessing and normalization steps can provide a more accurate and efficient data foundation for subsequent data analysis and task execution, as follows:
[0103] Text data: needs to be cleaned, including removing meaningless symbols and punctuation marks, and performing word segmentation to ensure that the text content is structured and standardized;
[0104] Image data: resizing and color standardization are required to unify the format and style of the image.
[0105] Video data: It needs to be split into frames and retain the position information to ensure the integrity and timing of the video sequence;
[0106] Audio data: Sampling processing is required to adjust the frequency of the signal and remove noise to improve audio quality;
[0107] Sensor data: usually expressed as numerical information such as temperature, humidity, pressure, position, acceleration, etc., and is often time series data. It needs to be processed for noise reduction, missing value filling, anomaly detection and filtering to ensure the accuracy and reliability of the data.
[0108] (3) Feature Coding
[0109] The preprocessed data needs to go through a feature extraction process to map the multimodal data into a unified vector representation:
[0110] Text feature extraction: can use but is not limited to pre-trained language models to generate context-sensitive feature vectors;
[0111] Image feature extraction: Based on ViT, the image is cut into fixed-size patches, and each patch is input into the Transformer encoder as a sequence to extract the global feature vector representation. Finally, the image feature dimension is adjusted to be consistent with the text input through linear mapping;
[0112] Video feature extraction: ViT is used to encode the key frame images and convert each frame into a high-dimensional vector representation. By adding spatial and temporal position encoding, the spatial structure information within the frame and the temporal relationship between frames are captured, and the spatiotemporal features of the video are modeled, thereby efficiently representing the dynamic and static information in complex scenes.
[0113] Audio feature extraction: For audio input, an audio feature extractor (such as MFCCs, Mel frequency cepstral coefficients) is used to generate a dense vector for representing the audio content. The audio feature extractor can convert the time series audio waveform into frequency domain features, and can extract key information such as pitch and timbre. For image processing, the extracted frequency domain features will be dimensional compressed and mapped to a common feature space, thereby providing a standardized representation for subsequent processing;
[0114] Sensor feature extraction: Sensor feature extraction usually includes statistical features (such as mean, variance, etc.), frequency domain features (such as fast Fourier transform FFT), time series features (such as autocorrelation, periodicity), and advanced features automatically extracted by deep learning models. In order to improve efficiency and reduce redundancy, feature selection and dimensionality reduction techniques, such as principal component analysis (PCA), can be used to retain the most discriminative features. Ultimately, these extracted features will be converted into fixed-length numerical feature vectors for subsequent analysis and processing.
[0115] (4) Multimodal feature alignment
[0116] In multimodal data processing, the features of each modality need to be aligned to ensure that they represent multimodal information in the same vector space. This process normalizes the feature vectors of different modalities into a unified representation space by sharing encoding dimensions or performing linear mapping. The goal of feature alignment is to enable the perception module to establish cross-modal associations and effectively fuse features of multiple data forms into the same semantic space, thus laying the foundation for subsequent processing and analysis.
[0117] (5) Multimodal feature fusion
[0118] Based on feature alignment, the Decode-Only Transformer architecture is used to further process these multimodal feature vectors. The decoding layer of the Transformer can effectively model and parse cross-modal information by gradually extracting context-related deep features. The self-attention mechanism can capture the mutual influence between different modalities to generate a unified internal representation. This representation not only contains contextual information, but also integrates the semantic content of each modality. It has a high degree of versatility and interpretability, ensuring the accurate understanding and effective integration of cross-modal data. Figure 3 As shown;
[0119] In the process of generating consistent semantic representation, by fusing the personalized background information of the portrait configuration module with multimodal data, the perception module can optimize the representation of the fused vector according to the specific needs of the user.
[0120] (6) Output data
[0121] It outputs a unified multimodal semantic representation vector, integrating features such as text, images, and audio. It has consistency, context awareness, and multimodal semantic fusion capabilities, providing deep semantic support and personalized optimization for tasks such as retrieval, classification, and generation.
[0122] The memory module of this embodiment is as follows:
[0123] The memory module is based on the Attention mechanism, which realizes the weighted storage of high-value information in short-term memory and the vectorized storage of historical experience in long-term memory. Specifically, short-term memory stores important data by weight, while long-term memory uses a vector database to save historical experience to ensure the systematicness and relevance of knowledge. The dynamic review mechanism migrates key data to long-term memory according to data weights, thereby improving the effectiveness of information storage. Data indexing is achieved through semantic vectorization and context awareness to ensure that content related to the current task can be quickly retrieved during retrieval. The call strategy provides real-time decision support and realizes multi-task memory sharing through self-attention and cross-attention mechanisms, ensuring the adaptability and efficiency of the agent in complex tasks. The following is a detailed design plan:
[0124] like Figure 4 As shown, the memory module of this embodiment is mainly composed of the following modules:
[0125] (1) Memory storage structure design
[0126] Short-term memory (STM) storage layer: Uses Attention-based weighted cache storage to dynamically evaluate the importance of task information and only store information with high weight. As the task progresses, low-weight information is gradually removed to ensure efficient use of resources. Information is temporarily stored and used in GPU / TPU video memory or RAM in the form of dynamic hidden states, Key-Value caches, attention weight matrices, etc.
[0127] Long-term memory (LTM) storage layer: Use vector databases or knowledge graphs to store long-term memory data. The encoder based on the Transformer architecture represents the data in a multimodal manner, and the Attention mechanism is used to extract the global features of the data and establish a multi-dimensional association index.
[0128] (2) Data indexing and retrieval mechanism
[0129] Semantic vectorization and attention focus: Use natural language processing (NLP) technology to extract key feature words, and encode them through the Transformer model to generate semantic vectors. Use the Attention mechanism to calculate the relevance weight of each memory and generate a semantic index to support fuzzy retrieval and semantic similarity matching;
[0130] Context awareness and association index: The Attention model dynamically calculates the weight of memory data in the current context, generates a context index for the memory content, and marks features such as task type, time, and environmental status to quickly locate highly associated data.
[0131] (3) Dynamic review and reinforcement mechanism
[0132] Dynamic review: The self-attention mechanism is used to periodically review short-term memory, filter high-weight data, and transfer them to the long-term memory layer. The review process is triggered by task completion or scheduled by time intervals to ensure that valuable information is transferred in a timely manner.
[0133] Reinforcement mechanism: After the task is completed, reinforcement learning is used in combination with attention weight to evaluate the effectiveness of memory data. Information that contributes greatly to decision-making is reinforced by increasing its attention weight to form optimized experience for future tasks.
[0134] (4) Memory call mechanism
[0135] Contextual memory call: When the task is initialized, the Cross-Attention mechanism is used to retrieve long-term memory based on the task context, filter highly relevant experience data, and provide guidance for the current task.
[0136] Real-time memory call: During task execution, short-term and long-term memory are dynamically called, and the relevance of data is calculated in real time through the self-attention mechanism to assist decision-making. Multiple rounds of attention weight updates are performed on short-term memory to ensure that the latest and relevant data are called.
[0137] Multi-task sharing and inheritance: In a multi-task environment, memory sharing and knowledge transfer are achieved through a multi-head attention mechanism. The Attention model retrieves key memories based on the semantic similarity between tasks, ensuring that new tasks quickly adapt to the environment and inherit existing experience.
[0138] The portrait configuration module of this embodiment is as follows:
[0139] The portrait configuration module is the core of building personalized behavior and decision-making patterns for the intelligent agent. By setting the basic attributes, skills, expertise, role relationships and other information of the intelligent agent, the portrait configuration module directly affects the perception, planning and execution modules of the intelligent agent, so that its decision-making and behavior are in line with the expected role characteristics. The present invention takes multimodal pre-training models and knowledge graphs as the core, and combines reinforcement learning and generative language models to achieve dynamic management and personalized optimization of intelligent agent portraits. Through flexible definition and dynamic adjustment, the intelligent agent can display preset characteristics in different scenarios, ensuring the efficiency of task execution and the naturalness of human-computer interaction. The specific technical route is as follows: Figure 5 As shown:
[0140] (1) Image modeling and feature definition
[0141] The portrait configuration module constructs the characteristics of the agent by setting the basic attributes of the agent (such as personality, professional skills, language style) and contextual relationships (such as role relationships, task roles). Based on the pre-trained language model and combined with the label generation algorithm, the attribute feature embedding of the agent is directly generated through natural language description. For example, a "technical support assistant" can have a portrait definition of "strong professionalism and concise answers", while a "customer service role" can show the attributes of "friendly and concise language".
[0142] (2) Knowledge graph drives feature association
[0143] The knowledge graph technology is used to associate the multi-dimensional characteristics of the agent with its role tasks. The characteristic nodes include skills (skill point distribution), language style (semantic embedding), task preferences, etc., and the relationship edges represent the constraints and priorities between characteristics. The matching degree between the role characteristics of the agent and the context is dynamically calculated through graph embedding technology (such as Graph Neural Networks, GNN) to ensure that the behavior of the agent meets the scene requirements.
[0144] (3) Multimodal semantic-driven portrait generation
[0145] The multimodal feature fusion architecture based on Decoder-Only Transformer supports the generation of portraits in various input forms such as text and images. Feature portraits can be generated through semantic alignment based on multimodal clues configured with preferences (such as task descriptions, related documents, and scene images). For example, in warehouse management, the agent receives warehouse layout images and operation description texts to generate feature vectors. Based on these features, the agent formulates behavioral strategies to perform tasks such as moving items or arranging cargo locations.
[0146] (4) Optimizing image features based on reinforcement learning
[0147] Through reinforcement learning (RL), the agent profile is continuously optimized during task execution. Deep Q-Learning (DQL) is used to adjust feature priorities in real time. For example, after a user reports that the agent's behavior is "not friendly enough", the RL module will adjust the weight of the language style node and update the "affinity" feature to ensure that the response is more in line with user expectations.
[0148] The planning module of this embodiment is as follows:
[0149] The planning capability module makes decisions and action plans based on user needs. This module uses the reasoning ability of the large model to dynamically analyze the input instructions to formulate response strategies and specific execution paths. Among them, the input instructions contain relevant data from the memory, perception, profiling module and tool use module. Multi-source data can enhance the accuracy and efficiency of task decomposition, dependency identification, priority setting and task execution path planning.
[0150] The specific implementation route is as follows:
[0151] (1) Input data
[0152] The input data of the planning capability module comes from the memory module, perception module, portrait configuration module and tool use module. The memory module provides historical task records and feedback, the perception module provides real-time environment status, the portrait configuration module provides personalized needs, and the tool use module provides recommended tools and their features. These multi-source data jointly support task decomposition, dependency identification, priority setting and path planning, making planning more accurate and efficient, while ensuring that the action plan formulated fits user needs and the current task environment.
[0153] (2) Task decomposition and dependency identification method based on large language model
[0154] The user tasks are parsed through a pre-trained large language model to identify the task objectives and requirements. Through natural language processing technology, the model can extract the key elements of the task and convert them into executable subtasks. These subtasks not only clarify the specific objectives of the task, but also reveal the dependencies between them. Through grammatical and semantic analysis, the model can ensure that the dependencies in the task decomposition process are reasonably identified, providing accurate guidance for subsequent execution.
[0155] (3) Execution path optimization method based on improved topological sorting algorithm
[0156] The present invention proposes an improved topological sorting algorithm for establishing a task execution path according to the dependency relationship of subtasks. Topological sorting ensures that tasks are executed in the correct order, avoiding conflicts or resource waste caused by incorrect task order. By analyzing the dependency relationship, topological sorting can determine which tasks can be executed in parallel and which tasks need to wait. This method improves the efficiency of task execution and ensures resource optimization during task execution, thereby improving the intelligent scheduling capability of the entire system.
[0157] according to Figure 6 As shown in the figure, each node represents a subtask, and the edges between nodes represent dependency relationships. For example, a->b means that task b depends on task a, that is, task b can only be executed after task a is completed. The core steps of the improved topological sorting algorithm are as follows:
[0158] ① Build the task graph and in-degree table: traverse all tasks, build the task dependency graph G (directed graph), and calculate the in-degree of each task;
[0159] ② Find the task with in-degree 0: put all the tasks with in-degree 0 (i.e. tasks that do not depend on other tasks) in graph G into a container Vec, and merge them into one large node and put it into the task execution queue;
[0160] ③ Perform topological sorting: Take out all the nodes in Vec, and reduce the in-degree of their adjacent nodes by 1. Then, put all the adjacent nodes whose in-degree becomes 0 into Vec, and merge them into 1 large node and put it into the task execution queue. Repeat ③ until no new nodes can be added to Vec;
[0161] ④ Loop detection: Finally, check whether there are any nodes with in-degree not 0. If so, it means there is a loop and the task cannot be completed. Otherwise, the result is as follows Figure 7 The task execution queue shown.
[0162] (5) Output data
[0163] The output data is the planned task execution queue, which determines the specific subtask execution order. Figure 7 As shown in the figure, the subtasks are generally executed in the order of A->B->C->D. For large nodes (such as nodes B and C) that contain multiple subtasks in the task execution queue, if the resources required for execution are sufficient, the execution order of the tasks contained in the node (such as tasks b, c, and d contained in node B) is not restricted, and the overall efficiency can be improved through concurrent execution. Otherwise, priority can be sorted according to the required resource status.
[0164] The tool usage module of this embodiment is as follows:
[0165] Based on the task execution path output by the planning module, the tool usage module is responsible for matching the external tools required to execute the task. This step needs to be completed before the specific action is executed to ensure that all resources are in place so that the action module can effectively complete the planned steps.
[0166] This invention designs a tool matching system that integrates semantic matching and reinforcement learning to achieve efficient and accurate tool matching. Figure 8 As shown:
[0167] (1) Input data
[0168] The task execution queue output by the receiving planning module records the task execution path and related tool descriptions, etc., providing basic data for subsequent processing.
[0169] (2) Constructing the query statement of the subtask execution tool
[0170] Based on the generation capability of the pre-trained large language model, tool matching query statements are generated according to the tasks to be executed and the tool description template structure.
[0171] (3) Matching task execution tools based on similarity algorithm
[0172] The tool matching query and tool description are converted into vector representations through the pre-trained language model. Each query and tool description will be converted into a fixed-dimensional vector that can retain the deep semantic information of the text. When matching, the vector similarity algorithm is used to calculate the similarity between the tool query and the tool description to achieve the matching of the subtask and the target tool.
[0173] (4) Optimization tool selection results based on reinforcement learning
[0174] After the semantic matching module gives the preliminary matching results, reinforcement learning is used to optimize the matching strategy through user feedback and historical data, and the matching results are continuously adjusted based on the reward and penalty mechanism. The strategies that can be used are as follows:
[0175] ① Weighted fusion: The similarity score of semantic matching and the reward score of reinforcement learning are weightedly combined to generate the final matching result;
[0176] ② Correction based on feedback: If there is a large difference between the feedback from the reinforcement learning module and the result of semantic matching, the system will adjust the weight of semantic matching and give priority to the optimization result of reinforcement learning;
[0177] ③Hierarchical decision-making: First, semantic matching is performed to obtain a batch of candidate tools, and then the reinforcement learning model is used to select the best tool from them.
[0178] The matching fusion module combines the results of semantic matching and the optimization feedback of the reinforcement learning module to generate the final matching tool recommendation.
[0179] (5) Output data
[0180] Combined with the pre-trained large language model, it outputs the subtask execution tool matching results, provides matching reasons and scores, and helps users understand the basis of the recommendation. It also generates a tool calling plan and provides the action module for use.
[0181] The action modules of this embodiment are as follows:
[0182] Based on the tasks and recommended tool information generated by the tool calling module and the personalized information provided by the portrait configuration module, the action module generates specific tool calling instructions and calls the tool to complete the task. Fig. 9 As shown:
[0183] (1) Input data
[0184] The input data consists of two parts: ① The recommended tool information provided by the tool call module, including the most suitable tools matched by the algorithm, as well as the basic functions and applicable scenarios of these tools; ② The personalized data of the portrait configuration module, including user preferences, historical behaviors, demand characteristics, etc. These data together provide the action module with the background information and customized requirements required to perform tasks.
[0185] (2) Generate tool call instructions
[0186] Based on the tool information and the personalized data of the profile, a pre-trained large language model is used to generate tool call instructions. The process first analyzes the applicability of the recommended tool and generates precise task execution instructions based on the preferences and needs in the user profile. These instructions not only describe the task objectives, but also include specific operational requirements such as how to use the recommended tool, call parameters, and execution order.
[0187] (3) Calling tools to execute tasks
[0188] After generating the task execution instructions, call the tool according to the instructions to gradually complete the specified task. During the tool execution, the action module monitors the task progress in real time to ensure that the tool executes smoothly according to the instructions. If any deviations or problems occur during the task execution, the action module will make adjustments based on the execution feedback to ensure that the task is completed as expected.
[0189] (4) Output data
[0190] All execution results are summarized, combined with the user preference information in the profile configuration module, and the final feedback content is generated through the pre-trained large language model. These feedback contents will not only describe the execution of the tool in detail, but also automatically adjust the format and content of the feedback according to the personalized needs of the user. The large language model can also dynamically optimize the feedback content according to the quality of task completion and changes in user preferences to make it more in line with user expectations.
[0191] In summary, the intelligent agent architecture based on the multimodal large model of this embodiment is:
[0192] Modular architecture design: By dividing the intelligent agent system into 6 internal core modules and 3 external support modules, an efficient, flexible and scalable intelligent agent architecture is innovatively realized. Each module has clear functions and cooperates with each other, ensuring the efficiency and consistency of task execution;
[0193] Multimodal information input and processing: The input module supports the reception of multimodal information, including text, images, audio, video and other data forms, and can access data from multiple devices such as cameras, sensors, microphones, etc. This innovation supports the intelligent agent's comprehensive perception of complex environmental information and strengthens the cross-modal understanding ability;
[0194] Dynamic memory and context management: The memory module uses an attention mechanism to support short-term and long-term memory, and can store and review dynamic information. This innovation provides contextual support, ensures the consistency of task execution, and enhances the ability of the intelligent agent to cope with multiple tasks and complex situations;
[0195] Personalized portrait configuration and reinforcement learning update: The portrait configuration module sets the personalized characteristics of the agent and uses reinforcement learning to dynamically adjust the portrait to ensure that the decision and behavior are consistent with the role characteristics of the agent. This innovation improves the adaptability and personalized service capabilities of the agent under different users or scenarios;
[0196] Cross-modal decision-making supported by multi-modal big models: The big model capability module improves the information flow and collaborative work between modules such as perception, memory, image configuration, and planning by introducing multi-modal big models, and promotes the realization of cross-modal decision-making capabilities. This innovation enables the intelligent agent to understand and make decisions based on data from different modalities, enhancing the intelligence of the system;
[0197] Accurate task planning and tool matching: The planning module combines a pre-trained large language model and an improved topological sorting algorithm to achieve task decomposition and execution path formulation; the tool use module is based on the execution path provided by the planning module, and optimizes tool selection through semantic matching and reinforcement learning to ensure accurate matching of external tools or APIs required for task execution. This innovation ensures the accuracy of task execution and the efficiency of tool use;
[0198] Accurate tool calling and personalized feedback: The action module combines personalized portrait data with recommended tool information, uses a pre-trained large language model to generate accurate tool calling instructions, and adjusts the task execution process based on execution feedback. This innovation provides a more flexible and personalized task execution and real-time feedback mechanism to ensure efficient execution of intelligent agents in complex tasks;
[0199] The present invention proposes a flexible and intelligent modular architecture by integrating innovative technologies such as multimodal input, personalized profiling, dynamic memory, precise planning and tool calling, which provides strong support for the application of intelligent agents in multiple scenarios and solves the problems of low task execution efficiency, poor adaptability and inaccurate feedback in traditional intelligent agent architecture.
[0200] The application scenarios of the present invention include but are not limited to the following areas:
[0201] Intelligent customer service and assisted support: In the field of customer service, intelligent agents are used to realize multimodal information understanding, user portrait customization and external tool calling, providing automated and personalized customer support, including voice recognition, sentiment analysis and problem solving functions.
[0202] Industrial monitoring and maintenance: In industrial scenarios, intelligent agents provide equipment status monitoring, anomaly detection and predictive maintenance through real-time perception and memory of multimodal information such as video and sensor data, supporting complex operating condition management and fault warning.
[0203] Education and training: As a teaching aid, the intelligent agent can realize dynamic and personalized teaching assistance through the comprehensive understanding of multimodal data (text, images, audio), adapt to the needs of different learners, and provide more targeted educational content.
[0204] Medical and health management: In medical diagnosis and health management, intelligent agents can integrate patients' multimodal data (such as medical history, imaging data, etc.), support health monitoring, disease prediction, remote consultation and medical advice, and provide auxiliary support for personalized treatment plans.
[0205] Autonomous driving and intelligent transportation: The intelligent agent serves as the decision-making core in autonomous driving and intelligent transportation systems. It realizes dynamic path planning, obstacle recognition and intelligent traffic scheduling through the perception, planning and memory of multimodal information (such as radar, camera, etc.).
[0206] Intelligent security and public safety: Intelligent agents can be used in scenarios such as surveillance video analysis and abnormal behavior recognition. By integrating multimodal information, they can realize real-time recognition and processing of potential safety hazards and enhance public safety protection capabilities.
[0207] The above is only a preferred embodiment of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions under the concept of the present invention belong to the protection scope of the present invention. It should be pointed out that for ordinary technicians in this technical field, some improvements and modifications without departing from the principle of the present invention should also be regarded as the protection scope of the present invention.
Claims
1. An intelligent agent architecture based on a multimodal large model, including an external support module and an internal core module, characterized in that: The external support module includes: Input module, used to support multi-modal information reception, including text, image, video, audio and sensor modal data. Data sources include cameras, sensors, microphones and text input; The output module is responsible for generating user-understandable feedback information, including text, voice, image and chart display; The large model capability module is used to provide support for internal core modules through multimodal large model capabilities; The internal core modules include: The perception module is based on the Decode-Only Transformer architecture and is combined with the subsequent portrait configuration module to transform multimodal input into a unified internal representation, thereby achieving cross-modal understanding and decision-making capabilities, laying the foundation for subsequent task processing; Memory module, used to store and call dynamic information through attention mechanism; The portrait configuration module is used to provide preference information to the perception module, subsequent planning module, and action module by setting the personalized characteristics of the intelligent agent, and dynamically update the portrait configuration using reinforcement learning to ensure that the decision and behavior are consistent with the role characteristics of the intelligent agent; Planning module: This module is based on the data of the perception module, memory module and portrait configuration module, as well as the support of the subsequent tool usage module. It uses a pre-trained large language model and an improved topological sorting algorithm to achieve task decomposition and execution path formulation; Tool usage module: Based on the subtask execution path provided by the planning module, the tool usage module combines semantic matching and reinforcement learning to achieve accurate tool matching and optimization, ensuring that the external tools or APIs required for task execution accurately correspond to the planned task path; The action module combines the recommended tool information of the tool calling module and the personalized data of the portrait configuration module, uses the large language model to generate accurate tool calling instructions, and executes tasks. At the same time, it adjusts the task execution process according to the execution feedback to ensure that the task is completed as expected and provide personalized feedback.
2. The agent architecture based on a multimodal large model according to claim 1, characterized in that: The perception module is based on the Decode-Only Transformer architecture and is combined with the subsequent portrait configuration module to transform multimodal input into a unified internal representation, thereby achieving cross-modal understanding and decision-making capabilities. The specific steps to lay the foundation for subsequent task processing are as follows: Step 11: Input data. The data input in this step includes two parts: ① original data including text, images, videos, audio, and sensors; ② personalized data of the portrait configuration module, including user preferences, historical behaviors, and demand characteristics; Step 12, perform data preprocessing. The specific preprocessing steps are as follows: clean the text data, remove meaningless symbols and punctuation, and perform word segmentation to ensure that the text content is structured and standardized; resize and color standardize the image data to unify the format and style of the image; split the video data by frame and retain the position information to ensure the integrity and timing of the video sequence; sample the audio data to adjust the frequency of the signal and remove noise to improve the audio quality; perform noise reduction, missing value filling, anomaly detection and filtering on the sensor data to ensure the accuracy and reliability of the data; Step 13, feature encoding is performed, and feature extraction is performed on the data preprocessed in step 12, and the multimodal data is mapped into a unified vector representation. The specific feature extraction steps are as follows: the pre-trained language model is used to generate context-sensitive feature vectors for text feature extraction; the image feature extraction is based on ViT to cut the image into fixed-size tiles, and each tile is input into the Transformer encoder as a sequence to extract the global feature vector representation. Finally, the image feature dimension is adjusted to be consistent with the text input through linear mapping; the video feature extraction is performed by ViT to feature encode the key frame image and convert it into a vector representation; the audio feature extractor is used to extract audio features to generate a dense vector for representing the audio content; the sensor feature extraction is performed by feature selection and dimensionality reduction methods to retain the most discriminative features; finally, each of the above extracted features will be converted into a fixed-length numerical feature vector for subsequent analysis and processing; Step 14: perform multimodal feature alignment, and standardize the feature vectors of different modalities into a unified representation space by sharing encoding dimensions or performing linear mapping; Step 15: Based on the feature alignment in step 14, multimodal feature fusion is performed. Specifically, the Decode-Only Transformer architecture is used to further process these multimodal feature vectors. Specifically, the context-related deep features are gradually extracted through the decoding layer of the Transformer, and the mutual influence between different modalities is captured through the self-attention mechanism of the Transformer, so as to generate a unified internal representation. Step 16, output data, specifically output a unified multimodal semantic representation vector, integrate text, image, and audio features, and have consistency, context awareness, and multimodal semantic fusion capabilities, providing deep semantic support and personalized optimization for tasks such as retrieval, classification, and generation.
3. The agent architecture based on a multimodal large model according to claim 2, characterized in that: The memory module comprises: A memory storage module, which includes a short-term memory storage layer and a long-term memory storage layer, which uses a high-speed cache storage based on Attention weighting through the short-term memory storage layer to dynamically evaluate the importance of task information and only store information with high weights, and uses a vector database or a knowledge graph through the long-term memory storage layer to store long-term memory data; Data indexing and retrieval module, which includes a semantic vectorization and attention focusing module and a context awareness and association indexing module. The module uses natural language processing technology to extract key feature words and encodes them through the Transformer model to generate semantic vectors. The context awareness and association indexing module uses the Attention model to dynamically calculate the weight of memory data in the current context, generate a context index for the memory content, and mark the task type, time, and environmental status features to quickly locate highly relevant data. Dynamic review and reinforcement module, which includes a dynamic review module and a reinforcement module. The dynamic review module uses the self-attention mechanism to periodically review short-term memory, screen high-weight data and transfer them to the long-term memory layer. After the task is completed, the reinforcement module uses reinforcement learning combined with attention weight to evaluate the effectiveness of memory data; The memory calling module includes a situational memory calling module, a real-time memory calling module and a multi-task sharing and inheritance module. When the task is initialized, the situational memory calling module uses the Cross-Attention mechanism to retrieve long-term memory based on the task context, screens highly relevant experience data, and provides guidance for the current task. During the task execution, the real-time memory calling module dynamically calls short-term and long-term memories, calculates the relevance of data in real time through the self-attention mechanism, assists decision-making, and performs multiple rounds of attention weight updates on short-term memory to ensure that the latest and relevant data are called. Through the multi-task sharing and inheritance module, memory sharing and knowledge transfer are realized through the multi-head attention mechanism in a multi-task environment.
4. The agent architecture based on a multimodal large model according to claim 3, characterized in that: The specific steps of the portrait configuration module to set the personalized characteristics of the agent and dynamically update the portrait configuration using reinforcement learning are as follows: Step 31, perform portrait modeling and feature definition, specifically based on the pre-trained language model, combined with the label generation algorithm, directly generate the attribute feature embedding of the agent through natural language description, so as to set the basic attributes and contextual relationships of the agent to construct the characteristics of the agent; Step 32, perform knowledge graph driven feature association, specifically, first use knowledge graph technology to associate the multi-dimensional features of the agent with its role tasks, the associated feature nodes include skills, language style, task preferences, and the associated relationship edges represent the constraints and priorities between the features, and then use graph embedding technology to dynamically calculate the matching degree between the role features of the agent and the context, to ensure that the behavior of the agent meets the scene requirements; Step 33, generating a multimodal semantically driven portrait, specifically generating a feature portrait through semantic alignment based on multimodal clues configured based on preferences; Step 34, optimizing the image features based on reinforcement learning, specifically using deep reinforcement learning to adjust the feature priority in real time.
5. The multimodal large model-based intelligent agent architecture according to claim 4, characterized in that: The planning module uses a pre-trained large language model and an improved topological sorting algorithm to implement task decomposition and execution path formulation in the following specific steps: Step 41, inputting data from the memory module, the perception module, the portrait configuration module and the tool use module, specifically, the memory module provides historical task records and feedback, the perception module provides real-time environment status, the portrait configuration module provides personalized needs, and the tool use module provides recommended tools and their characteristics; Step 42, performing task decomposition and dependency identification based on the large language model, specifically parsing the user task through the pre-trained large language model, identifying the task objectives and requirements, and using natural language processing technology to enable the model to extract key elements of the task and convert them into executable subtasks; Clarify the specific goals of the task through subtasks and reveal the dependencies between subtasks; Step 43, performing execution path calculation and optimization based on an improved topological sorting algorithm; Step 44, output data, specifically output the planned task execution queue and determine the specific subtask execution order.
6. The multimodal large model-based intelligent agent architecture according to claim 5, characterized in that: The improved topological sorting algorithm in step 43 includes the following steps: Step 431, constructing a task graph and an in-degree table, specifically by traversing all tasks, constructing a task dependency graph G, and calculating the in-degree of each task; Step 432, find the task with in-degree 0, specifically put all the tasks with in-degree 0 in graph G into container Vec, and merge them into one large node and put them into the task execution queue; Step 433, perform topological sorting, specifically, take out all the nodes in Vec, reduce the in-degree of their adjacent points by 1 in turn, then put all the adjacent points whose in-degree becomes 0 into Vec, and merge them into one large node and put it into the task execution queue, and repeat step 433 until no new nodes can be added to Vec; Step 434, perform loop detection, specifically check whether there are any nodes with in-degree not 0. If so, it means that there is a loop and the task cannot be completed. Otherwise, a task execution queue containing all tasks can be obtained. In addition, for large nodes containing multiple tasks in the task execution queue, if there are sufficient resources for execution, the execution order of the tasks contained in the node is not restricted, and the overall efficiency can be improved through concurrent execution. Otherwise, priority can be sorted according to the required resource status.
7. The multimodal large model-based intelligent agent architecture according to claim 6, characterized in that: The specific steps of the tool usage module to achieve accurate tool matching and optimization are as follows: Step 51, input data, specifically receiving the task execution path and related tool description data recorded in the task execution queue output by the planning module, to provide basic data for subsequent processing; Step 52, constructing a subtask execution tool query statement, specifically, based on the generation capability of the pre-trained large language model, generating a tool matching query statement according to the task to be executed and the tool description template structure; Step 53, matching the task execution tool based on the similarity algorithm, specifically, converting the tool matching query statement and tool description into a vector representation through the pre-trained language model, each query and tool description will be converted into a vector of fixed dimension, and then the vector similarity algorithm is used to calculate the similarity between the tool query statement and the tool description, so as to achieve the matching of the subtask and the target tool; Step 54, selecting the result based on the reinforcement learning optimization tool, specifically, after the semantic matching module gives the preliminary matching result, using reinforcement learning, optimizing the matching strategy through user feedback and historical data, and continuously adjusting the matching result based on the reward and penalty mechanism, which can be adopted in the following ways: Weighted fusion: The similarity score of semantic matching and the reward score of reinforcement learning are weightedly combined to generate the final matching result; Feedback-based correction: If there is a large difference between the feedback from the reinforcement learning module and the result of semantic matching, the system will adjust the weight of semantic matching and give priority to the optimization result of reinforcement learning; Hierarchical decision-making: First, semantic matching is performed to obtain a batch of candidate tools, and then the reinforcement learning model is used to select the best tool from them; Finally, the matching fusion module combines the semantic matching results and the optimization feedback of the reinforcement learning module to generate the final matching tool recommendation; Step 55, combined with the pre-trained large language model, outputs the subtask execution tool matching results, provides matching reasons and scores, helps users understand the basis of the recommendation, and generates a tool calling plan for use in the action module.
8. The multimodal large model-based intelligent agent architecture according to claim 7, characterized in that: The action module generates precise tool call instructions and executes tasks, and adjusts the task execution process according to execution feedback to ensure that the task is completed as expected and provide personalized feedback. The specific steps are as follows: Step 61, input data, which includes two parts: recommended tool information provided by the tool calling module, including the most suitable tool matched by the algorithm, as well as the basic functions and applicable scenarios of these tools, and personalized data of the portrait configuration module, including user preferences, historical behaviors, and demand characteristics; Step 62, generating a tool calling instruction, specifically, generating the tool calling instruction using a pre-trained large language model based on the tool information and the personalized data of the profile, further specifically, first analyzing the applicability of the recommended tool, and then generating accurate task execution instructions based on the preferences and needs in the user profile; Step 63, calling the tool to execute the task, specifically, after the task execution instruction is generated in step 62, calling the tool according to the instruction to gradually complete the specified task, and during the tool execution, the action module monitors the task progress in real time to ensure that the tool is executed smoothly according to the instruction; if any deviation or problem occurs during the task execution, the action module will make adjustments based on the execution feedback to ensure that the task is completed as expected; Step 64, summarize all execution results, combine with the user preference information in the portrait configuration module, and generate the final feedback content through the pre-trained large language model.
Citation Information
Cited By
Intelligent agent task training corpus construction method and device for large language model and readable medium
CN118377860A
Intelligent agent construction method and system for large model and intelligent recommendation
CN120277272A
A method and system for building an intelligent agent with large models and intelligent recommendations
CN120277272B
Abnormal behavior identification and early warning method based on cloud edge cooperation mechanism and related system
CN120337107A
Abnormal behavior identification and early warning method and related system based on cloud-edge collaboration mechanism
CN120337107B