Intelligent agent strategy optimization method based on multi-modal fusion, intelligent agent and electronic equipment

By building a fusion strategy library in the agent and selecting the most matching fusion strategy, fusion and learning of multimodal data is solved, and the agent's task execution efficiency and accuracy are achieved in complex environments, and more efficient and accurate decision-making is achieved.

CN120234759APending Publication Date: 2025-07-01GUIYANG SHIJIHENGTONG TECH
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510374596.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

The existing multimodal fusion technology is difficult to accurately extract key information in complex environments, resulting in low task execution efficiency and accuracy of agents.

Method used

By selecting the most matching fusion strategy in the pre-built fusion strategy library, the multimodal data of the target task is fused, and the task decision is generated based on the fusion data, and the decision is adjusted to adapt to the task sharing information of other agents.

Benefits of technology

It improves the agent's understanding and response speed of the environment, reduces the false alarm rate, and improves the accuracy and efficiency of task execution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120234759A_ABST
    Figure CN120234759A_ABST
Patent Text Reader

Abstract

The invention provides an agent strategy optimization method based on multi-modal fusion, an agent and electronic equipment. The method comprises the following steps: selecting a best matching fusion strategy for a target task in a pre-constructed fusion strategy library; wherein the fusion strategy library comprises a plurality of predefined fusion strategies; fusing the multi-modal data required by the target task by using a best matching fusion strategy; when the target task is executed, learning is carried out based on the fused multi-modal data corresponding to the target task, and a task decision is generated. According to the method, information loss possibly caused by simple splicing can be avoided, a foundation is laid for the accuracy of subsequent task decision making, and meanwhile, the task decision making can be dynamically adjusted according to the interaction condition with other agents, so that rapid and efficient learning is realized in a complex scene, and the accuracy and adaptability of agent decision making are greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology. Specifically, it relates to an intelligent agent strategy optimization method, an intelligent agent, and an electronic device based on multimodal fusion. Background Art

[0002] As an entity that can perceive the environment and make reasonable decisions according to environmental changes, the intelligent agent plays a crucial role in the modern information technology system. In order to enable the intelligent agent to better adapt to the complex real world, researchers have begun to introduce multimodal fusion technology into the design and implementation of intelligent agents. Multimodal fusion refers to combining information from different sensory channels (such as vision, hearing, touch, etc.) or different types (such as images, texts, audio, etc.) to form a more comprehensive understanding of the target object or scene.

[0003] When the multimodal fusion technology is applied to the intelligent agent, it can significantly improve the intelligent agent's ability to understand the environment and response speed. However, due to numerous environmental interference factors, such as complex lighting conditions, noisy background sounds, etc., the existing multimodal fusion technology is difficult to accurately extract key information, resulting in a high false alarm rate. At the same time, the existing intelligent agent interaction mechanism is mostly based on fixed communication protocols and limited information sharing methods. In a collaborative task of multiple intelligent agents, they can often only perform simple information exchanges, such as sharing location information or task progress. This interaction method is insufficient when facing tasks that require in-depth cooperation and complex strategy formulation, resulting in low task execution efficiency.

[0004] Therefore, although certain progress has been made in multimodal data processing and intelligent agent interactive learning, how to improve the task execution efficiency and accuracy of the intelligent agent is a technical problem that needs to be solved. Summary of the Invention

[0005] In view of this, the purpose of the present invention is to provide an intelligent agent strategy optimization method, an intelligent agent, and an electronic device based on multimodal fusion, which are used to improve the task execution efficiency and accuracy of the intelligent agent. To achieve the above purpose, the technical solution adopted by the present invention is as follows:

[0006] In the first aspect, the present invention provides an intelligent agent strategy optimization method based on multimodal fusion. The method includes: selecting the most matching fusion strategy for the target task in a pre-constructed fusion strategy library, where the fusion strategy library contains a variety of predefined fusion strategies; using the most matching fusion strategy to fuse the multimodal data required for the target task; when executing the target task, learning based on the fused multimodal data corresponding to the target task to generate a task decision; and adjusting the task decision according to the task sharing information of other intelligent agents executing the target task.

[0007] Second aspect, the present invention provides an agent, comprising: a multimodal fusion module and a task execution module; the multimodal fusion module is configured to select the most matching fusion strategy for a target task from a pre-constructed fusion strategy library; wherein, the fusion strategy library contains a variety of predefined fusion strategies; the multimodal fusion module is further configured to use the most matching fusion strategy to fuse the multimodal data required for the target task; the task execution module is configured to, when executing the target task, learn based on the fused multimodal data corresponding to the target task to generate a task decision; the task execution module is further configured to adjust the task decision according to the task sharing information of other agents executing the target task.

[0008] Third aspect, the present invention provides an electronic device, comprising a processor and a memory, the memory stores machine-executable instructions that can be executed by the processor, and the processor can execute the machine-executable instructions to implement the method for optimizing the agent strategy based on multimodal fusion according to any one of the foregoing embodiments.

[0009] The method for optimizing the agent strategy based on multimodal fusion, the agent and the electronic device provided by the embodiments of the present invention select the most matching fusion strategy for the target task from a pre-constructed fusion strategy library. This fusion strategy library contains a variety of predefined fusion strategies. By selecting the most suitable fusion strategy for the current target task from the fusion strategy library, it can be ensured that the fusion process is not simply splicing, but customized according to the specific requirements of the task, thus avoiding information loss that may be caused by simple splicing and laying a foundation for the accuracy of subsequent task decisions. Then, the selected most matching fusion strategy is used to fuse the multimodal data required for the target task. When executing the target task, learning is carried out based on the fused multimodal data to generate a task decision. Obtaining high-quality fused data through the fusion strategy most suitable for the target task can improve the quality of the learning process, enabling the agent to better understand the task and thus make more accurate decisions. In addition, the embodiments of the present invention can also adjust the task decision according to the task sharing information of other agents executing the target task. By introducing the shared information of other agents, the agent can adjust and optimize its own decision, thereby further improving the accuracy and adaptability of the decision. This adjustment based on shared information enables the agent to better cope with complex and changing task environments, and ultimately achieves the effect of significantly improving the decision accuracy of the agent.

[0010] To make the above objects, features and advantages of the present invention more obvious and understandable, the following specifically enumerates preferred embodiments and, in conjunction with the accompanying drawings, makes a detailed description as follows. Description of the Drawings

[0011] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the embodiments. It should be understood that the following drawings only show some embodiments of the present invention, and thus should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.

[0012] Figure 1 Schematic flowchart of the intelligent agent strategy optimization method based on multimodal fusion provided by the embodiments of the present invention;

[0013] Figure 2 Multimodal data fusion architecture diagram provided by the embodiments of the present invention;

[0014] Figure 3 An intelligent agent interactive learning model diagram provided by the embodiments of the present invention;

[0015] Figure 4 Functional module diagram of an intelligent agent provided by the embodiments of the present invention;

[0016] Figure 5 Structure block diagram of the electronic device provided by the embodiments of the present invention. Detailed implementation manners

[0017] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Usually, the components of the embodiments of the present invention described and shown in the accompanying drawings here can be arranged and designed in various different configurations.

[0018] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed present invention, but merely represents the selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.

[0019] It should be noted that relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising a..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.

[0020] Please refer to Figure 1 , Figure 1 FIG. is a schematic flowchart of an intelligent agent strategy optimization method based on multimodal fusion provided by an embodiment of the present invention. The execution subject of this method can be an electronic device, including steps S101 to S104, which are described as follows:

[0021] S101: Select the most matching fusion strategy for the target task in a pre-constructed fusion strategy library; wherein, the fusion strategy library contains a variety of predefined fusion strategies;

[0022] S102: Use the most matching fusion strategy to fuse the multimodal data required for the target task;

[0023] S103: When executing the target task, learn based on the fused multimodal data corresponding to the target task to generate a task decision;

[0024] S104: Adjust the task decision according to the task sharing information of other intelligent agents executing the target task.

[0025] In the above steps S101 to S104 provided by the embodiments of the present invention, first, the most suitable fusion strategy is selected for the target task in the pre-constructed fusion strategy library. The fusion strategy library contains a variety of predefined fusion strategies. By selecting the most suitable fusion strategy for the current target task from the fusion strategy library, it can be ensured that the fusion process is not simply splicing, but customizing the fusion method according to the specific requirements of the task, thus avoiding information loss that may be caused by simple splicing and laying a foundation for the accuracy of subsequent task decisions. Next, the selected most suitable fusion strategy is used to fuse the multi-modal data required for the target task. When the target task is executed, learning is performed based on the fused multi-modal data to generate a task decision. Obtaining high-quality fused data through the fusion strategy most suitable for the target task can improve the quality of the learning process, enabling the intelligent agent to better understand the task and thus make more accurate decisions. In addition, the embodiments of the present invention can also adjust the task decision according to the task sharing information of other intelligent agents executing the target task. By introducing the shared information of other intelligent agents, the intelligent agent can adjust and optimize its own decision, thereby further improving the accuracy and adaptability of the decision. This adjustment based on shared information enables the intelligent agent to better cope with the complex and changeable task environment and ultimately achieve the effect of significantly improving the decision accuracy of the intelligent agent.

[0026] Next, the embodiments of the present invention will clearly and detailedly describe the implementation manners of the above steps in combination with the relevant drawings.

[0027] In step S101, the embodiments of the present invention consider that: currently common multi-modal fusion methods mainly include early fusion, late fusion, and hybrid fusion. Early fusion directly stitches different modal data in the data preprocessing stage. For example, in the joint analysis task of images and texts, the image feature vector and the text word vector are simply connected and then input into the subsequent model. Its advantage is simple implementation, but it ignores the characteristic differences of each modal data at different stages. Late fusion is to first independently process each modal data and obtain their respective decision results and then perform fusion. For example, in sentiment analysis, the sentiment tendency is first obtained separately through the text analysis model and the speech intonation analysis model, and then the results of the two are combined. However, this method fails to fully utilize the interaction information of multi-modal data in the early stage. Hybrid fusion attempts to combine the advantages of the former two and perform fusion to varying degrees at different stages, but in practical applications, the decision on when and how to perform fusion lacks accuracy and dynamic adaptability.

[0028] For example, in the field of security monitoring, multi-modal fusion technology is used for personnel identification and behavior analysis. By combining surveillance video images and environmental sounds, it attempts to identify abnormal behaviors. However, due to numerous environmental interference factors, such as complex lighting conditions, noisy background sounds, etc., existing fusion technologies are difficult to accurately extract key information, resulting in a relatively high false alarm rate. Another example is in the field of medical diagnosis, where multi-modal fusion technology is used to integrate medical images (such as X-rays, CTs, etc.) and medical record text information. But current fusion methods cannot effectively handle the high-dimensional and heterogeneous problems of data, making it difficult to further improve the diagnostic accuracy.

[0029] To solve the above problems, embodiments of the present invention create a multi-modal fusion mechanism that can dynamically and precisely adapt to different task requirements and data characteristics. This mechanism can analyze the complex internal relationships of various modal data such as text, images, and audio in real time, flexibly adjust the fusion strategy according to the actual situation, avoid information loss caused by simple splicing or shallow fusion, and can significantly improve the utilization efficiency of multi-modal data and the depth of information mining to cope with complex and changeable application scenarios, laying a foundation for subsequent improvement of the learning speed and decision-making accuracy of intelligent agents.

[0030] Therefore, in step S101, embodiments of the present invention pre-construct a fusion strategy library. The fusion strategy library contains a variety of predefined fusion strategies. These fusion strategies, in addition to including the above existing fusion strategies, may also include new fusion strategies such as fusion strategies based on the attention mechanism and fusion strategies based on graph neural networks.

[0031] Based on the fusion strategy library, the implementation manner of step S101 can be: select the most matching fusion strategy from the fusion strategy library according to the task requirements of the target task and the characteristics of the required multi-modal data.

[0032] For example, assume the target task is "obstacle detection for autonomous driving vehicles". To complete this task, through task analysis, it can be known that obstacle detection requires quickly and accurately identifying the position, shape, and distance of the object in front, so the required multi-modal data may include camera image data, radar data, etc. Based on the requirements of the target task and the characteristics of the multi-modal data, embodiments of the present invention can select the most suitable fusion strategy for this task.

[0033] From the above implementation manner, it can be seen that before executing step S101, embodiments of the present invention need to first collect the multi-modal data required for the target task and perform a series of preprocessing to provide a reliable data source for subsequent generation of task decisions.

[0034] During the process of collecting multi-modal data, the data collection scope can be dynamically adjusted according to the current task objective. For example, if the task is drone obstacle recognition, radar point clouds and visible light images containing buildings and cables are preferentially collected; if the task is pilot voice command parsing, an aviation control terminology library and cockpit environment recordings need to be obtained directionally. For various modal data such as text, images, and audio, professional collection devices and software tools can be used to obtain data from different data sources. For example, text data can be collected through web crawlers, images can be captured using cameras, and audio can be recorded with microphones.

[0035] In the embodiments of the present invention, considering that the collected data often contains noise, error values, or incomplete information, the multi-modal data can also be pre-cleaned. During the data cleaning process, for text data, regular expressions can be used to remove special characters, correct spelling mistakes, and perform entity standardization (such as unifying names like "drone / UAV", etc.); for image data, noise points can be removed through image filtering algorithms, size normalization (unifying to 512×512 pixels), abnormal frame elimination (such as deleting blurred or black screen images), and image inpainting techniques can be used to fill in missing parts and correct incorrect object detection frames; for audio data, noise reduction algorithms can be used to remove background noise, clip silent segments, suppress background noise, align sampling rates, and standardize volume, etc. By these means, the quality of the data for subsequent processing is ensured, providing a basis for accurate analysis.

[0036] For the collected multi-modal data, the embodiments of the present invention can be represented by different feature expressions. For example, for text data, word embedding models (such as Word2Vec, BERT, etc.) can be used to convert the text into low-dimensional vectors and extract semantic features; for image data, convolutional neural networks (CNNs) can be used to extract visual features of the images, such as edges and textures; for audio data, it can be transformed to the frequency domain through Fourier transform and frequency features can be extracted. These feature extraction methods can transform the original data into feature vectors suitable for subsequent processing, highlighting the key information of the data and providing a data basis for subsequent feature fusion.

[0037] Based on the above-collected multi-modal data, the embodiments of the present invention provide an implementation manner for the above step S101, including steps a1 to a3, which are described as follows:

[0038] Step a1: Parse the multi-source task information of the target task to obtain a task requirement vector;

[0039] Step a2: According to the task requirement vector, quantify the feature contribution degrees of different modal features involved in the required multi-modal data to the target task;

[0040] Step a3: Select the most matching fusion strategy according to the feature contribution degrees corresponding to each modal feature.

[0041] In step a1, the multi-source task information may include, but is not limited to: task description text (such as the user instruction "generate image description"), task type tags (such as classification, generation, detection), historical task execution records (such as the modality weight distribution of similar tasks), and environmental parameters (such as the device sensor status).

[0042] In the embodiments of the present invention, the parsing process may use natural language processing techniques to extract task keywords (such as "description", "detection"), initially determine the core requirements of the task; according to the task type tags, determine the basic type of the task and the modalities that may be involved; analyze the current environmental parameters (such as light intensity, sensor availability), analyze the parameters of the current environment, such as light intensity, sensor availability, etc., understand the environmental conditions for task execution, and adjust the priority of task requirements and the applicability of modalities.

[0043] In step a2, by synthesizing the task parsing content obtained in step a1, a task requirement vector is dynamically constructed to clarify the specific requirements of the task. For example, in a drone navigation task, the task requirement vector constructed through task parsing = {task description: "low visibility navigation", task type: "navigation", environmental conditions: "low light intensity"}. It should be understood that this form of expression of the task requirement vector is only an example and is not a limitation on the task requirements in the embodiments of the present invention. For example, parameters such as priority can also be added to the task requirement vector to adapt to complex task scenarios.

[0044] In step a3, in combination with the task requirement vector obtained in step a2, the embodiments of the present invention can quantify the contribution degrees of different modality features involved in the multi-source data required for the target task to the target task, obtain a modality feature contribution degree vector, and according to the modality feature contribution degree vector, the embodiments of the present invention can determine the importance degrees of various modality features to the target task, and accordingly dynamically select the most appropriate fusion strategy to make the utilization rate of multi-modal data more efficient and accurate.

[0045] For step a3, the embodiments of the present invention comprehensively evaluate the importance of different modality features to the target task from three aspects: semantic relevance, data quality score, and task history statistics. Among them, semantic relevance refers to the matching degree between the words in the text and their corresponding modality features in a certain task, and the data quality score reflects whether the data quality of a certain modality is reliable enough; task history statistics is based on the experience of past similar tasks, analyzes the performance of each modality in the past, and thus predicts its importance in future tasks.

[0046] Therefore, as an optional implementation manner, step a3 can be implemented as follows:

[0047] Step 1: Obtain the task keywords and task type in the task requirement vector;

[0048] Step 2: Determine the semantic similarity between the task keywords and the modal features corresponding to the task keywords;

[0049] In the embodiments of the present invention, a set of all task keywords and modal features can be determined. According to the word frequency of the task keywords in the task description text, a keyword vector is constructed; according to the modal feature set, a modal feature vector is constructed, and then the distance between the keyword vector and the modal feature vector is calculated as the similarity.

[0050] For example, assume the task is to "describe a red apple". The extracted keywords are {"red", "apple"}, and the corresponding modal features are {"color", "shape"}, which are closely related to the visual modality (image). Therefore, a keyword vector can be constructed as [1, 1, 0, 0]. The modal feature vector is [0, 0, 1, 1], and then the similarity between these two vectors is calculated.

[0051] Step 3: Evaluate the data quality score of multi-modal data based on a preset sensor confidence;

[0052] In the embodiments of the present invention, the data quality score reflects whether the data quality of a certain modality is reliable enough. For example, if lidar is used to collect point cloud data, the higher the point cloud density (e.g., exceeding 90%), the more complete and less noisy the data is, so this modality should be given a higher weight. During the implementation process, the data quality can be evaluated through the confidence of the sensor. The higher the confidence, the more credible the data collected by the sensor, and the higher the importance of the modality.

[0053] Step 4: Obtain historical tasks of the same task type and count the historical feature contribution degrees of each modal feature in the historical tasks;

[0054] In the embodiments of the present invention, by combining the experience of the same type of tasks from the task history statistics, the past performance of each modal feature is analyzed to predict its importance in future tasks. For example, if in the past 100 image description tasks, the proportion of visual features (image modality) reached 78%, this indicates that the visual modality has always been dominant in such tasks. Therefore, in future similar tasks, the visual modality should also be given a greater weight.

[0055] Through statistical analysis, calculate the historical feature contribution degree of each modal feature in the historical tasks. For example, record the output results of different modalities in the tasks, calculate their influence on the final task goal, and then obtain the weight.

[0056] Step 5: Calculate the feature contribution degree of each modality feature based on semantic similarity, data quality score, and the historical feature contribution degree of each modality feature in historical tasks.

[0057] After obtaining the feature contribution degree of each modality feature through the above implementation, it is possible to determine which modality features are key features, and then select the most suitable fusion strategy from the fusion strategy library. For example, if the feature contribution degree of image features is relatively high, then a fusion strategy that can highlight image features can be selected; conversely, if text features are more important, a strategy that is more inclined to text features should be selected.

[0058] The embodiment of the present invention provides an implementation for selecting the most suitable fusion strategy: that is, based on the feature contribution degree, several candidate fusion strategies are selected from the fusion strategy library. For each candidate fusion strategy, the performance after fusion is evaluated using the validation set. By comparing the performance of different fusion strategies, the fusion strategy that performs best on the validation set is selected as the final most matching fusion strategy.

[0059] For example, in the image caption generation task, the system analyzes and determines the key nature of the visual modality in the following way: 1. Detect the number of significant objects in the image, for example, 5 high-confidence targets are identified through the YOLOv8 algorithm; 2. Compare the amount of information in the auxiliary text metadata, for example, only containing the label "outdoor scene"; 3. Calculate the mutual information between the visual features and the generated text, for example, the visual feature contribution degree reaches 82%, so as to determine that the visual feature is a key feature, and a fusion strategy that can highlight image features can be selected.

[0060] In an embodiment of the present invention, if the task requirements change, then based on the updated task requirements, the feature contribution degrees of each multi-modal feature can be dynamically adjusted to re-implement the weight allocation of different modality features.

[0061] In the embodiment of the present invention, the feature contribution degrees of each modality feature can also be dynamically evaluated during the task execution process, and the weights of each modality feature can be dynamically adjusted. For example, a multi-modal Transformer model can be used to simultaneously process image and text data through its multi-head self-attention mechanism. The model will automatically learn the complex relationships between image and text features and dynamically adjust their weights according to the task requirements.

[0062] For example, taking the video emotion analysis task as an example: when extreme expressions (such as crying) or violent body movements appear in the detected video frame, the weight of the visual modality is increased to 0.8, and at this time, the weight of the text modality (subtitle) is decreased to 0.2. Another example is that in a surveillance scenario, when a person suddenly starts running, the motion features extracted by the 3DCNN are the key features, which can activate the attention mechanism and increase the contribution degree of visual features. In a text-dominated scenario, if the video contains a large number of professional terms (such as a medical lecture) and the video frame is a static PPT, the weight of the text modality is increased to 0.7. The system detects that the frequency of keywords such as "myocardial infarction" and "thrombolytic therapy" is > 15 times / minute through the Term Frequency-Inverse Document Frequency (TF-IDF) algorithm, triggering the reinforcement of the Long Short-Term Memory (LSTM) text encoder.

[0063] In the embodiment of the present invention, for step S102, when the most matching fusion strategy is the fusion strategy based on the attention mechanism, the most matching fusion strategy is used to fuse the multi-modal data required for the target task, including:

[0064] Step b1: Extract the feature vectors corresponding to the required multi-modal data respectively;

[0065] Step b2: Calculate the correlation weights between the feature vectors through the attention mechanism;

[0066] Step b3: Perform weighted processing on the feature vectors according to the correlation weights;

[0067] Step b4: Concatenate or perform weighted summation on the weighted feature vectors to obtain the fused feature vector.

[0068] For example, when the task parsing module detects that the image modality is crucial, it automatically selects the fusion strategy based on the attention mechanism: after the visual features are extracted by ResNet-50, they interact with the text encoding through the cross-attention layer; the modality weights are dynamically generated (such as the visual Q matrix weight = 0.68, and the text K matrix weight = 0.32); the CIDEr index (an index that combines the BLEU model and the vector space model) of the output fused features is improved by 12.7% on the COCO dataset. Through the joint analysis of the task semantics and data characteristics, the embodiment of the present invention realizes the real-time optimization of the fusion strategy, and the accuracy is improved by 19% compared with the fixed fusion strategy in complex scenarios.

[0069] For the convenience of understanding the above steps S101 to S102, please refer to Figure 2 , Figure 2This is the multimodal data fusion architecture diagram provided by the embodiments of the present invention. It can be seen that the embodiments of the present invention can dynamically select and adjust the fusion method according to task requirements and data characteristics, which is the key innovation point different from the traditional fixed fusion mode. By establishing a fusion strategy library and a task analysis mechanism, the intelligence and efficiency of multimodal fusion are realized.

[0070] In the interactive learning of agents, the data after multimodal fusion plays a key role in the following aspects:

[0071] First, by dynamically adjusting the contribution degrees of the features of each modality, more accurate environmental understanding is achieved, enhancing the environmental perception ability. For example, in the scenario of a medical surgical robot, by fusing visual (endoscopic images), tactile (pressure sensors), and voice (doctor's instructions) data, the accuracy of instrument operation can be significantly improved. When abnormal tissue tremors are detected, the system will automatically increase the weight of the tactile modality to 0.75, so as to better adapt to the complex surgical environment.

[0072] Second, the fused feature vector is input into the reinforcement learning model as a state representation to drive the real-time update of the agent's strategy and support the optimization of dynamic decision-making. For example, in the scenario of industrial quality inspection, by fusing infrared imaging (for defect detection) and acoustic data (equipment abnormal sound analysis), the robotic arm can dynamically adjust the detection path according to multi-dimensional evidence. For example, when a fault occurs on the production line, the system will increase the weight of the acoustic modality from 0.3 to 0.6 to preferentially capture equipment abnormal signals, thereby improving the fault diagnosis efficiency.

[0073] Finally, based on the unified feature space constructed from the fused data, the agent can quickly transfer the learned strategy to a new scenario, achieving the effect of cross-task transfer learning. For example, after an educational robot is trained by fusing voice commands, student expressions, and answer record data, when switching to the task of elderly care, it can automatically strengthen the weight allocation of the emotion analysis modality to better meet user needs. During the training process, the system will use the fused feature vector as the input and optimize the model in combination with interactive feedback (such as user ratings, environmental reward signals). Another example is in the scenario of autonomous driving. The agent trains the decision-making network by fusing lidar, camera, and vehicle-to-everything (V2X) data, and can automatically allocate the weight of the infrared sensor feature to 0.7 when driving at night to make up for the perception limitations caused by insufficient lighting. This training method based on dynamically fused data shows excellent results in the scenario of manufacturing quality inspection, increasing the defect recognition accuracy to 98.6%, fully reflecting the practical application value of multimodal fusion technology.

[0074] Next, the embodiments of the present invention will introduce how the agent executes tasks based on the fused multi-modal data. Refer to steps S103 to S104.

[0075] In the embodiments of the present invention, please refer to Figure 3 , Figure 3 which is a diagram of an agent interactive learning model provided by the embodiments of the present invention. The agent can be a model constructed using a deep reinforcement learning framework, such as the Proximal Policy Optimization (PPO) algorithm. The agent includes a policy network and a value network. The policy network is used to generate the task decisions of the agent, and the value network is used to evaluate the value of the current state of the agent. The agent interacts with the environment, continuously learns to optimize its own network parameters, and generates and optimizes task decisions. For example, in a multi-agent collaborative logistics distribution scenario, each agent represents a delivery vehicle, and the agent needs to determine the driving route and delivery order based on information such as the current location of the goods and traffic conditions.

[0076] Considering traditional agent learning algorithms, such as Q-learning, Deep Q-Network (DQN), etc., which rely on predefined reward functions and state spaces. In simple environments, these algorithms can achieve certain effects. However, in complex and dynamic real-world environments, fixed reward functions cannot comprehensively reflect the long-term impact of agent behavior and are difficult to adapt to sudden changes in the environment. For example, in an autonomous driving scenario, the agent needs to cope with various traffic conditions and emergencies, and traditional algorithms are difficult to quickly learn and make optimal decisions.

[0077] Therefore, in the embodiments of the present invention, in step S103, during the learning process of the agent, the gradient descent algorithm is used to calculate the gradient of the preset loss function with respect to its own network parameters, and the preset network parameters are adjusted along the opposite direction of the gradient to make the loss function converge.

[0078] In one implementation, during the learning process of the agent, the gradient descent algorithm is used to update the network parameters of the policy network and the value network. By calculating the gradient of the loss function with respect to the network parameters and adjusting the network parameters along the opposite direction of the gradient, the loss function gradually decreases until it remains unchanged. For example, in the policy network, the policy gradient is calculated based on the actions of the agent and the environmental feedback to update the network parameters to improve the decision-making ability of the agent.

[0079] In the embodiments of the present invention, environmental feedback mainly refers to the response given by the environment itself through state transition and reward signals after the agent executes an action, rather than directly referring to the information of other agents. Specifically, it can be divided into two core parts: State transition feedback: When the agent executes an action, the environment will transition to a new state. For example, in an autonomous driving scenario, after the vehicle executes an acceleration action, the environmental feedback may include state parameters such as the new vehicle speed and the distance from the vehicle in front, which reflect the changes in the physical environment. Reward signal feedback: The environment generates a scalar reward value based on the current state and action. For example, in a robotic arm grasping task, a positive reward of 10 is given when the object is successfully grasped, and a negative penalty of 5 is given when a collision occurs. This signal directly guides the gradient update direction of the policy network. In a multi-agent scenario, the behaviors of other agents will be incorporated into the environmental dynamic model. For example, in the cooperation of soccer robots, the running position information of teammates will be fed back to the current agent as part of the environmental state. The core of environmental feedback is that it reflects the objective result of the interaction between the agent and the environment, rather than an active information exchange process.

[0080] In another implementation, the agent also adaptively adjusts the learning rate during the learning process, which can accelerate the model convergence and avoid getting stuck in local optima. For example, by adopting an adaptive learning rate adjustment method such as the Adam optimization algorithm, the agent can dynamically adjust the learning rate according to the first-order and second-order moment estimates of the gradient. At the beginning of training, a larger learning rate is given to enable the model to quickly update parameters; as training progresses, the learning rate is gradually decreased to enable the model to converge to the optimal solution more stably.

[0081] Compared with the traditional agent learning algorithm that relies on a fixed reward function and a state space learning method and has low learning efficiency and poor adaptability in complex environments, the embodiments of the present invention adopt a deep reinforcement learning framework combined with adaptive learning rate adjustment, which can quickly optimize the task strategy according to environmental changes and the feedback of other agents. For example, in an autonomous driving simulation experiment, when facing sudden road conditions, the agent of the embodiments of the present invention can, through the above implementation, increase the reaction speed of driving decisions by 1.8 times and shorten the time to learn the optimal driving strategy by 40%, significantly enhancing the adaptability and decision-making ability of the agent in complex dynamic environments.

[0082] The above-mentioned adaptive and dynamic adjustment characteristics provided by the embodiments of the present invention enable the agents in the embodiments of the present invention to have higher flexibility and adaptability in complex and changeable real-world environments. It can effectively improve the utilization efficiency of multi-modal data, enhance the learning speed and decision-making accuracy of the agent, and provide a better solution for solving various complex practical problems. By combining an effective gradient descent algorithm and an adaptive learning rate adjustment method, the learning speed and convergence stability of the agent can be improved, providing a guarantee for realizing efficient agent interactive learning.

[0083] In step S104, the embodiment of the present invention takes into account that existing agent interaction mechanisms are mostly based on fixed communication protocols and limited information sharing methods. In a collaborative task of multiple agents, they can often only perform simple information exchanges, such as sharing location information or task progress. This interaction method seems inadequate when facing tasks that require in-depth collaboration and complex strategy formulation. For example, when multiple robots collaborate to complete a complex assembly task, there is a lack of an effective coordination mechanism between agents, resulting in low task execution efficiency. Therefore, the embodiment of the present invention establishes a shared information space where agents can upload their own state information, task progress, etc. to the shared space and can also obtain relevant information of other agents from the shared space.

[0084] Therefore, in one embodiment, step S104 may include the following steps c1 to c3, which are described as follows:

[0085] Step c1: Obtain the state information and task progress information of other agents from the shared information space;

[0086] Step c2: Based on the state information and task progress information, the agent adjusts its own task decision;

[0087] Step c3: Upload its own state information and task progress information to the shared information space.

[0088] In the embodiment of the present invention, the shared information space can be deployed in the form of a cloud server. For example, a cloud architecture based on a trusted data space is adopted, and high-availability storage and computing are achieved through a distributed server cluster. In scenarios with high real-time requirements (such as industrial Internet of Things), a deployment mode of collaboration between edge nodes and the cloud can be adopted.

[0089] For the shared information space, a standardized API interface can be used to implement information upload. For example, an agent can upload structured / unstructured data to the shared information space through RESTful API or gRPC protocol. In the process of information upload, Quantum Key Distribution (QKD) technology can also be used to ensure transmission security. For example, when the shared information space requires data upload, the agent can encrypt the data through the national cryptographic SM9 algorithm and record the operation log through the blockchain. Another example is that the agent can also convert electromagnetic signals into semantic instructions at the physical layer and achieve lossless docking of electromagnetic modal data with the shared information space through digital coding metasurfaces. This architecture design has achieved millisecond-level data synchronization in the smart factory scenario and supports concurrent collaboration of more than 2000 agents.

[0090] Through the above embodiments, during the multi-agent collaboration process, each agent adjusts its own action strategy according to the shared information. For example, when multiple agents jointly complete a search task, when an agent discovers a clue about the target, it uploads the clue information to the shared space, and other agents adjust their search directions based on this information, improving the search efficiency.

[0091] To facilitate the overall understanding of the above strategy optimization method provided by the embodiments of the present invention, the following takes the intelligent security monitoring system and the intelligent medical diagnosis assistance system as examples for introduction.

[0092] Embodiment 1: Intelligent Security System

[0093] First of all, the intelligent security system needs to perform multi-modal data collection and preprocessing, including:

[0094] (1) Collect multi-modal data: Install multiple high-definition cameras in the monitoring area to collect image data, and at the same time configure a microphone array for the collection of audio data. The cameras capture the pictures of the monitoring scene in real time, and the microphones record the sounds of the surrounding environment, thus constructing a complete multi-modal data set.

[0095] (2) Clean and extract features from the multi-modal data: For the image data, use image denoising algorithms to eliminate noises caused by factors such as light and sensors; use edge detection algorithms to extract key features such as the outlines of people and objects in the images. For the audio data, use spectrum analysis techniques to remove background noises and extract important features such as the frequency and amplitude of the sounds. Finally, convert the processed image and audio features into vector forms suitable for subsequent processing.

[0096] Based on the above collected multi-modal data, the intelligent security system can perform adaptive multi-modal fusion, mainly including:

[0097] (1) Select the most suitable fusion strategy: According to the characteristics of the security monitoring task, the system preferentially selects a multi-modal fusion strategy based on the attention mechanism. The system will dynamically evaluate the current monitoring scene to identify whether there are abnormal events (such as personnel intrusion, illegal behaviors, etc.). If potential abnormalities are detected, the attention weights for relevant image regions and audio frequency bands will be enhanced.

[0098] (2) Fusion of data from different modalities: Send the extracted image feature vectors and audio feature vectors into the fusion module, and calculate the correlation weights between the two through the attention mechanism. Then, weight and splice the feature vectors according to the weights to generate a fused feature representation, providing more comprehensive information support for the subsequent task decision-making of the agents.

[0099] The agent conducts interactive learning and decision-making based on the fused multi-modal data, mainly including:

[0100] (1) Deploy agent models: Deploy multiple agents in the security monitoring system. Each agent is responsible for monitoring a specific area or performing specific tasks (such as target tracking, event warning, etc.). These agents adopt deep reinforcement learning models and continuously optimize their own performance during the continuous interaction with the monitoring environment.

[0101] (2) Interactive learning and decision-making: Agents achieve real-time data exchange and result sharing through a shared information space. Once an agent discovers abnormal signs, it will immediately upload the relevant information to the shared space for other agents to refer to. For example, an agent responsible for monitoring the surrounding area may expand the monitoring range to assist in checking for more abnormal conditions. Finally, based on the judgment results of all agents, the system decides whether to trigger an alarm or initiate corresponding emergency measures.

[0102] Example 2: Intelligent medical diagnosis assistance system

[0103] First, the intelligent medical diagnosis assistance system needs to perform multi-modal data collection and preprocessing, including:

[0104] (1) Collect multi-modal data: Extract the medical record text data of patients from the hospital information system, including detailed information such as symptom descriptions and examination reports. At the same time, collect the medical image data of patients, such as X-ray films, CT scan images, and other relevant medical imaging materials.

[0105] (2) Clean and extract features from multi-modal data: Perform preprocessing operations such as word segmentation and stop word removal on the medical record text data, and then use a word embedding model to convert the text into a vector representation with semantic meaning. For medical image data, use professional image segmentation algorithms to extract regions of interest (such as lesion sites), and further extract features such as texture and shape in the images through convolutional neural networks.

[0106] Based on the above collected multi-modal data, the intelligent medical diagnosis assistance system can perform adaptive multi-modal fusion, mainly including:

[0107] (1) Select the most suitable fusion strategy: For medical diagnosis tasks, the system selects a multi-modal fusion strategy based on graph neural networks according to the complexity of the patient's condition and the characteristics of the data. By analyzing the keywords in the medical record text and the features in the medical images, an association graph between different data is constructed to better reflect the internal relationship between the data.

[0108] (2)Fuse data from different modalities: Map the text feature vectors and image feature vectors into a graph structure, and utilize the message passing mechanism of the graph neural network to propagate and fuse information among nodes. This method can fully explore the potential associations between medical record texts and medical images, generating more diagnostically valuable fused feature representations.

[0109] The agent conducts interactive learning and diagnostic assistance based on the fused multi-modal data, mainly including:

[0110] (1)Deploy the agent model: In the medical diagnostic assistance system, each agent corresponds to a disease diagnosis model or an expert experience module. These agents optimize their model parameters by learning a large amount of historical case data, thereby improving the diagnostic accuracy.

[0111] (2)Interactive learning and diagnostic assistance: When faced with a new patient case, different agents independently analyze the multi-modal fusion data according to their respective diagnostic models. At the same time, the agents exchange diagnostic ideas and results through an information sharing mechanism, collaborating to provide doctors with more comprehensive and accurate diagnostic suggestions. For example, the agent responsible for diagnosing cardiovascular diseases can cooperate with the agent responsible for diagnosing respiratory diseases to jointly discuss the patient's symptoms and examination results, ensuring that the final diagnostic suggestions cover various possible causes and providing strong support for clinical decision-making.

[0112] Based on the above embodiments, it can be summarized that the intelligent agent strategy optimization method provided by the embodiments of the present invention has the following advantages:

[0113] At the multi-modal fusion level, compared with the simple splicing or shallow fusion methods of traditional technologies, the adaptive multi-modal fusion strategy provided by the embodiments of the present invention can deeply explore the complex internal connections between various modal data such as text, images, and audio. For example, in the scenario of combining medical image diagnosis with medical record text analysis, traditional methods often have difficulty effectively integrating the information of both, while the present invention can accurately capture the correlation between image features and medical record descriptions by dynamically selecting the fusion method, thereby significantly improving the accuracy of disease diagnosis. In addition, in the face of different application scenarios and task requirements, traditional fusion technologies lack flexibility and are difficult to be dynamically adjusted according to the actual situation. However, by constructing a fusion strategy library and a task analysis mechanism, the embodiments of the present invention can select the optimal solution from a variety of predefined fusion methods in real time according to the task characteristics. For example, in the security monitoring scenario, the system can quickly switch to the appropriate fusion strategy for complex situations such as light changes and personnel flow, not only improving the accuracy of abnormal behavior recognition but also effectively reducing the false alarm rate, thereby greatly enhancing the overall reliability of the security system.

[0114] In the aspect of agent interactive learning, traditional agent learning algorithms are limited by fixed reward functions and state spaces, showing problems such as low learning efficiency and poor adaptability in complex environments. In the embodiments of the present invention, by introducing a deep reinforcement learning framework and combining it with adaptive learning rate adjustment, agents can quickly optimize their own strategies according to the dynamic changes of the environment and the feedback of other agents, thus significantly improving their adaptability and decision-making ability. For example, in an autonomous driving simulation experiment, when facing sudden road conditions, the decision-making reaction speed of the agent using the method of the present invention increased by 1.8 times, and the time to learn the optimal driving strategy was shortened by 40%, fully verifying the superiority of this method in complex dynamic environments. In addition, existing agent interaction mechanisms are mostly based on fixed communication protocols and limited information sharing, resulting in low efficiency in multi-agent cooperation tasks. To address this problem, the embodiments of the present invention propose an information sharing space and cooperation mechanism to support efficient information exchange between agents. Taking the example of multi-robot cooperation to complete a complex assembly task, by sharing key data such as real-time positions and task progress, agents can achieve collaborative operations, and the task completion time is shortened by 50%. This improvement greatly enhances the cooperation efficiency and execution effect of multi-agent systems in complex tasks, demonstrating the significant advantages of the present invention in practical applications.

[0115] Generally speaking, the embodiments of the present invention can be widely applied to multiple fields such as network content security supervision, social media platform management, online customer service systems, intelligent security monitoring, intelligent medical diagnosis assistance, and intelligent education personalized tutoring. By effectively integrating multi-modal data and optimizing agent interaction, it provides a general and efficient intelligent solution for each field, promoting the overall improvement of the intelligent level of each industry.

[0116] To execute the corresponding steps in the above embodiments and each possible manner, an implementation of an agent 40 is given below. Please refer to Figure 4 , Figure 4 which is a functional module diagram of an agent provided by the embodiments of the present invention. It should be noted that the basic principle and the technical effects generated by the agent 40 provided in this embodiment are the same as those in the above embodiments. For the sake of brief description, for the parts not mentioned in this embodiment, reference can be made to the corresponding content in the above embodiments. The agent 40 includes: a multi-modal fusion module 401 and a task execution module 402;

[0117] The multi-modal fusion module 401 is used to select the most matching fusion strategy for the target task in a pre-constructed fusion strategy library; among them, the fusion strategy library contains a variety of predefined fusion strategies;

[0118] The multi-modal fusion module 401 is also used to fuse the multi-modal data required for the target task by using the most matching fusion strategy;

[0119] The task execution module 402 is used to learn based on the fused multimodal data corresponding to the target task and generate a task decision when executing the target task;

[0120] The task execution module 402 is also used to adjust the task decision according to the task sharing information of other intelligent agents that perform the target task.

[0121] It is understandable that the multimodal fusion module 401 and the task execution module 402 can collaboratively execute the various steps in the above embodiments to achieve corresponding technical effects, and the embodiments of the present invention are not described in detail herein.

[0122] Optionally, the above modules can be stored in the form of software or firmware. Figure 5 The memory shown in FIG. 1 or the operating system (OS) of the electronic device 50 may be fixed therein and may be Figure 5 Meanwhile, the data and program codes required for executing the above modules can be stored in the memory.

[0123] See also Figure 5 , Figure 5 The electronic device 50 is a block diagram of a structure of an electronic device provided in an embodiment of the present invention, and includes a memory 501, a processor 502, and a communication interface 503. The memory 501, the processor 502, and the communication interface 503 are electrically connected to each other directly or indirectly to achieve data transmission or interaction. For example, these components can be electrically connected to each other through one or more communication buses or signal lines.

[0124] Optionally, the bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 5 Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.

[0125] In an embodiment of the present invention, the processor 502 may be a general-purpose processor, a digital signal processor, an application specific integrated circuit, a field programmable gate array or other programmable logic device, discrete gate or transistor logic device, discrete hardware component, and may implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present invention may be directly embodied as being executed by a hardware processor, or may be executed by a combination of hardware and software modules in the processor. The software module may be located in the memory 501, and the processor 502 reads the program instructions in the memory 501 and combines its hardware to complete the steps of the above method.

[0126] In an embodiment of the present invention, the memory 501 may be a non-volatile memory, such as a hard disk drive (HDD) or a solid state drive (SSD), etc., or may also be a volatile memory, such as RAM. The memory may also be any other medium that can be used to carry or store desired program executable code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory in the embodiments of the present invention may also be a circuit or any other device capable of implementing a storage function, for storing instructions and / or data.

[0127] The memory 501 may be used to store software programs and modules, and may be stored in the memory 501 in the form of software or firmware, or may be solidified in the operating system (OS) of the electronic device 500. The processor 502 executes various functional applications and data processing by executing the software programs and modules stored in the memory 501. The communication interface 503 may be used for signaling or data communication with other node devices.

[0128] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the above-described devices and units may refer to the corresponding processes in the foregoing method embodiments, and will not be described herein again.

[0129] It can be understood that Figure 5 The structure shown is only schematic, and the electronic device 50 may also include more or fewer components than those shown Figure 5 herein, or may have a different configuration from that shown Figure 5 herein. Figure 5 Each of the components shown may be implemented by hardware, software or a combination thereof.

[0130] The electronic device 50 may also be, but is not limited to, a single network server, a server group composed of multiple network servers, or a cloud composed of a large number of hosts or network servers based on cloud computing.

[0131] Based on the above embodiments, the embodiments of the present invention also provide a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a computer, the computer is made to execute the intelligent agent policy optimization method based on multimodal fusion provided by the above embodiments.

[0132] Based on the above embodiments, the embodiments of the present invention also provide a computer program. When the computer program runs on a computer, the computer is made to execute the intelligent agent policy optimization method based on multimodal fusion provided by the above embodiments.

[0133] In the embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some communication interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical or other forms.

[0134] In addition, the units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0135] Furthermore, in each embodiment of the present application, the various functional modules can be integrated together to form an independent part, or each module can exist alone, or two or more modules can be integrated to form an independent part.

[0136] It should be noted that when a function is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art or a part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of various embodiments of this application. The aforementioned storage medium includes: various media that can store program codes such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs.

[0137] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. An agent strategy optimization method based on multimodal fusion, characterized in that: The method comprises: Selecting the best matching fusion strategy for the target task from a pre-built fusion strategy library; wherein the fusion strategy library includes a plurality of pre-defined fusion strategies; Using the best matching fusion strategy to fuse the multimodal data required for the target task; When executing the target task, learning is performed based on the fused multimodal data corresponding to the target task to generate a task decision; The task decision is adjusted based on task sharing information of other agents performing the target task.

2. The agent strategy optimization method based on multimodal fusion according to claim 1, characterized in that: Select the best matching fusion strategy for the target task from the pre-built fusion strategy library, including: The best matching fusion strategy is selected from the fusion strategy library according to the task requirements of the target task and the characteristics of the multimodal data.

3. The agent strategy optimization method based on multimodal fusion according to claim 2 is characterized in that: According to the task requirements of the target task and the characteristics of the multimodal data, selecting the best matching fusion strategy from the fusion strategy library includes: Performing task analysis on multi-source task information of the target task to obtain a task requirement vector; quantifying, according to the task requirement vector, the feature contribution of different modal features involved in the multimodal data to the target task; The best matching fusion strategy is selected according to the feature contribution corresponding to each modal feature.

4. The agent strategy optimization method based on multimodal fusion according to claim 3 is characterized in that: According to the task requirement vector, quantifying the feature contribution of different modal features involved in the multimodal data to the target task includes: Obtaining task keywords and task types in the task requirement vector; Determining the semantic similarity between the task keyword and the modal feature corresponding to the task keyword; Evaluating a data quality score of the multimodal data based on a preset sensor confidence level; Obtaining historical tasks of the same task type, and counting the historical feature contribution of each modal feature in the historical tasks; The feature contribution of each of the modal features is calculated according to the semantic similarity, the data quality score, and the historical feature contribution of each of the modal features in the historical tasks.

5. The agent strategy optimization method based on multimodal fusion according to claim 1, characterized in that: When the best matching fusion strategy is a fusion strategy based on an attention mechanism, the multimodal data required for the target task is fused using the best matching fusion strategy, including: Extracting feature vectors corresponding to each of the multimodal data; Calculate the correlation weights between feature vectors through the attention mechanism; Performing weighted processing on each feature vector according to the correlation weight; The weighted feature vectors are concatenated or weighted summed to obtain a fused feature vector.

6. The agent strategy optimization method based on multimodal fusion according to claim 1, characterized in that: Based on the fused multimodal data corresponding to the target task, learning is performed to generate a task decision, including: During the learning process, the gradient descent algorithm is used to calculate the gradient of the preset loss function with respect to its own network parameters, and the preset network parameters are adjusted in the opposite direction of the gradient so that the loss function converges.

7. The agent strategy optimization method based on multimodal fusion according to claim 1, characterized in that: Based on the fused multimodal data corresponding to the target task, learning is performed to generate a task decision, including: The learning rate is adaptively adjusted during the learning process.

8. The agent strategy optimization method based on multimodal fusion according to claim 1, characterized in that: Adjusting the task decision according to task sharing information of other agents performing the target task includes: Acquire the status information and task progress information of the other agents from the shared information space; Based on the state information and task progress information, the agent adjusts its own task decision; Upload its own status information and task progress information to the shared information space.

9. An intelligent agent, characterized in that: include: Multimodal fusion module and task execution module; The multimodal fusion module is used to select the most matching fusion strategy for the target task from a pre-built fusion strategy library; wherein the fusion strategy library contains a plurality of predefined fusion strategies; The multimodal fusion module is further used to fuse the multimodal data required for the target task using the best matching fusion strategy; The task execution module is used to learn based on the fused multimodal data corresponding to the target task and generate a task decision when executing the target task; The task execution module is further used to adjust the task decision according to the task sharing information of other intelligent agents that execute the target task.

10. An electronic device, characterized in that: It includes a processor and a memory, wherein the memory stores machine executable instructions that can be executed by the processor, and the processor can execute the machine executable instructions to implement the intelligent agent strategy optimization method based on multimodal fusion as described in any one of claims 1 to 8.

Citation Information

Cited By

  • Smart phone and method

    CN120856822A

  • Multi-agent decision execution method and device, electronic equipment and storage medium

    CN121052342A

  • Large model agent driven medical multi-modal knowledge fusion method

    CN121598296A

  • Multi-agent cooperation method for resisting cross-modal cognitive interference based on cognitive chain

    CN121882292A