Intelligent agent interactive learning and optimization method based on adaptive multi-modal fusion

Through adaptive multimodal fusion and deep reinforcement learning framework, the flexibility and efficiency problems in multimodal fusion and intelligent agent interaction are solved, efficient multimodal data processing and intelligent agent collaboration are achieved, and the intelligence level in multiple fields is improved.

CN120633757AInactive Publication Date: 2025-09-12天津仁爱学院
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510577850.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-09-12
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing technologies lack flexibility and adaptability in multimodal fusion. Traditional algorithms have low learning efficiency in complex environments, and the agent interaction mechanism is inefficient, making it difficult to achieve efficient multimodal data processing and agent collaboration.

Method used

By adopting an adaptive multimodal fusion strategy, combined with a deep reinforcement learning framework and adaptive learning rate adjustment, a dynamic fusion strategy and information sharing mechanism are constructed through multimodal data preprocessing, adaptive multimodal fusion, agent interactive learning and optimization methods to improve data processing and agent learning efficiency.

Benefits of technology

It achieves efficient multimodal data fusion in different scenarios and tasks, improves the accuracy of disease diagnosis, the accuracy of abnormal behavior recognition and the decision-making ability of intelligent agents in complex environments, and improves the collaborative efficiency of multi-agent systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120633757A_ABST
    Figure CN120633757A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent agent interactive learning and optimization method based on adaptive multi-modal fusion, which relates to the technical field of interactive learning among intelligent agents, comprises links such as multi-modal data preprocessing, adaptive multi-modal fusion, intelligent agent interactive learning and optimization, and realizes multi-modal data processing and intelligent agent performance improvement. Compared with a simple splicing or superficial layer fusion mode of a traditional multi-modal fusion technology, the adaptive multi-modal fusion strategy of the invention can deeply analyze complex internal relations among various modal data such as texts, images, audios and the like. For example, in a scene of combining medical image diagnosis with medical record text analysis, a traditional method is difficult to effectively integrate information of the medical image diagnosis and the medical record text analysis, however, according to the method, association between image features and medical record description is accurately captured by dynamically adjusting a fusion mode, and the accuracy of disease diagnosis is greatly improved. Researches show that after the fusion method is used, the disease misdiagnosis rate is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of interactive learning between intelligent agents, and more specifically, relates to an adaptive multimodal fusion intelligent agent interactive learning and optimization method. Background Art

[0002] In the digital age, multimodal fusion and agent-interactive learning optimization are crucial, but existing technologies are insufficient. In terms of multimodal fusion, early fusion involves simple splicing, late fusion fails to utilize early interactive information, and hybrid fusion decisions are imprecise and suffer from poor dynamic adaptability. For example, security monitoring is subject to environmental interference, and medical diagnosis faces the challenge of high-dimensional heterogeneity, resulting in poor fusion results. In the field of agent-interactive learning, traditional algorithms rely on fixed reward functions and state spaces, lacking adaptability in complex dynamic environments (such as autonomous driving). Existing interaction mechanisms are based on fixed protocols and limited sharing, resulting in low efficiency for tasks such as multi-robot collaboration. Summary of the Invention

[0003] In order to solve the above technical problems, the present invention provides an adaptive multimodal fusion intelligent agent interactive learning and optimization method to solve the above problems.

[0004] The adaptive multimodal fusion intelligent agent interactive learning and optimization method includes multimodal data preprocessing, adaptive multimodal fusion, intelligent agent interactive learning and optimization methods, etc., to achieve multimodal data processing and intelligent agent performance improvement.

[0005] Preferably, the multimodal data preprocessing collects and cleans modal data such as text, images, and audio, and uses corresponding models (such as word embedding, CNN, Fourier transform) to extract features to ensure data quality and availability. The adaptive multimodal fusion constructs a dynamic fusion strategy, selects a suitable method (such as attention mechanism, graph neural network fusion, etc.) from the fusion strategy library through the task analysis module, and implements the fusion process accordingly. The intelligent agent interactive learning uses a deep reinforcement learning framework (such as PPO) to construct an intelligent agent model, which includes a policy network and a value network, and realizes intelligent agent interaction through information sharing and collaboration mechanisms.

[0006] Preferably, the optimization method uses the gradient descent algorithm to update the agent network parameters, and adopts an adaptive learning rate adjustment method (such as the Adam optimization algorithm) to accelerate convergence and avoid local optimality. The adaptive multimodal fusion strategy can dynamically select and adjust the fusion method according to the task and data, and use the fusion strategy library and task analysis module to improve the fusion intelligence and efficiency. Different from the traditional fixed mode, the agent interaction mechanism and collaborative optimization design information sharing and collaboration mechanism enable the agent to adjust the strategy according to the information of the other party, thereby improving the efficiency of multi-agent complex task collaboration.

[0007] Preferably, the combination of gradient descent-based parameter updates and adaptive learning rate adjustments improves the learning speed and convergence stability of the intelligent agent, ensures the interactive learning effect, and the adaptive multimodal data fusion technology innovates to construct a dynamic fusion strategy to select and adjust the fusion method, explore the intrinsic connection between modalities, and improve the fusion quality and efficiency. It is the core protection point of multimodal processing. The adaptive learning algorithm of the intelligent agent uses deep reinforcement learning and adaptive learning rate adjustment to enable the intelligent agent to optimize itself according to the environment and interactive feedback. It is the key innovation protection point of artificial intelligence.

[0008] Compared with the prior art, the present invention has the following beneficial effects: Multimodal fusion level: Deeply mine information: Compared to the simple splicing or shallow fusion methods of traditional multimodal fusion technology, the adaptive multimodal fusion strategy of the present invention can deeply analyze the complex internal connections between multiple modal data such as text, images, and audio. For example, in the scenario of combining medical imaging diagnosis with medical record text analysis, traditional methods find it difficult to effectively integrate the information of the two. However, the present invention dynamically adjusts the fusion method to accurately capture the relationship between image features and medical record descriptions, greatly improving the accuracy of disease diagnosis. Studies have shown that the use of the fusion method of the present invention reduces the misdiagnosis rate of diseases.

[0009] Flexible Adaptability: Traditional fusion technologies lack flexibility and are difficult to dynamically adjust to diverse application scenarios and task requirements. The fusion strategy library and task analysis module established by this invention enable real-time selection of the optimal solution from a variety of predefined fusion methods based on task characteristics. In security monitoring scenarios, the ability to quickly switch to the appropriate fusion strategy for complex situations such as changing lighting and personnel flow improves the accuracy of abnormal behavior identification, effectively reduces false alarm rates, and enhances the reliability of the security system.

[0010] Agent interactive learning level Efficient learning and rapid adaptation: Traditional agent learning algorithms rely on fixed reward functions and state spaces, resulting in low learning efficiency and poor adaptability in complex environments. This invention uses a deep reinforcement learning framework combined with adaptive learning rate adjustment, enabling agents to rapidly optimize their strategies based on environmental changes and feedback from other agents. In autonomous driving simulation experiments, agents using this method experienced improved decision-making response speeds when faced with unexpected road conditions, shortened the time it took to learn the optimal driving strategy, and significantly enhanced the adaptability and decision-making capabilities of the agents in complex dynamic environments.

[0011] Enhanced Collaboration and Improved Efficiency: Existing agent interaction mechanisms are often based on fixed communication protocols and limited information sharing, making them inefficient in multi-agent collaborative tasks. The information sharing space and collaborative mechanisms designed in this invention enable efficient information exchange between agents. For example, by enabling multi-robot collaboration to complete a complex assembly task, by sharing information such as location and task progress in real time, agents can collaborate, shortening task completion time and significantly improving the collaborative efficiency and performance of multi-agent systems in complex tasks.

[0012] Overall application level: Broad Applicability: This technology can be widely applied in a variety of fields, including network content security monitoring, social media platform management, online customer service systems, intelligent security monitoring, intelligent medical diagnostic assistance, and personalized intelligent education tutoring. By effectively integrating multimodal data and optimizing intelligent agent interactions, it provides universal and efficient intelligent solutions for various fields, promoting the overall improvement of the intelligence level of various industries.

[0013] Driving Technological Development: This invention's innovations offer new research ideas and methods for multimodal data processing and agent learning. Its successful application is expected to drive further development of related technologies, such as promoting research into more advanced multimodal fusion algorithms and agent interaction models, and promoting the in-depth application and innovative development of artificial intelligence technology in a wider range of fields. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 This is a diagram of the architecture of the present invention adapted to multimodal fusion; Figure 2 This is a diagram of the intelligent agent interactive learning model of the present invention. DETAILED DESCRIPTION

[0015] The following embodiments of the present invention are described in further detail with reference to the accompanying drawings and examples. The following examples are used to illustrate the present invention but are not intended to limit the scope of the present invention.

[0016] See also Figure 1-Figure 2 The present invention provides an adaptive multimodal fusion intelligent agent interactive learning and optimization method, which includes multimodal data preprocessing, adaptive multimodal fusion, intelligent agent interactive learning and optimization methods, etc., to achieve multimodal data processing and intelligent agent performance improvement.

[0017] Multimodal data preprocessing collects and cleans modal data such as text, images, and audio, and uses corresponding models (such as word embedding, CNN, and Fourier transform) to extract features to ensure data quality and availability. Adaptive multimodal fusion constructs a dynamic fusion strategy, and selects appropriate methods (such as attention mechanism, graph neural network fusion, etc.) from the fusion strategy library through the task analysis module to implement the fusion process accordingly. Agent interactive learning uses a deep reinforcement learning framework (such as PPO) to construct an agent model, which includes a policy network and a value network, and realizes agent interaction through information sharing and collaboration mechanisms.

[0018] The optimization method uses the gradient descent algorithm to update the agent network parameters, and adopts an adaptive learning rate adjustment method (such as the Adam optimization algorithm) to accelerate convergence and avoid local optimality. The adaptive multimodal fusion strategy can dynamically select and adjust the fusion method according to the task and data, and use the fusion strategy library and task analysis module to improve the intelligence and efficiency of fusion. Different from the traditional fixed mode, the agent interaction mechanism and collaborative optimization design information sharing and collaboration mechanism enable the agent to adjust the strategy based on the information of the other party, thereby improving the efficiency of multi-agent complex task collaboration.

[0019] The combination of gradient descent-based parameter updates and adaptive learning rate adjustments improves the learning speed and convergence stability of intelligent agents, ensuring interactive learning effects. Adaptive multimodal data fusion technology innovates the construction of dynamic fusion strategies to select and adjust fusion methods, explore the intrinsic connections between modalities, and improve fusion quality and efficiency. It is the core protection point of multimodal processing. The adaptive learning algorithm of intelligent agents uses deep reinforcement learning and adaptive learning rate adjustment to enable intelligent agents to optimize themselves according to the environment and interactive feedback. It is a key innovation and protection point in artificial intelligence.

[0020] Multimodal data preprocessing: Data collection and cleaning: For data in multiple modalities, such as text, images, and audio, specialized collection equipment and software tools are used to acquire data from various data sources. For example, text data can be collected through web crawlers, images can be captured using cameras, and audio can be recorded using microphones. The collected data often contains noise, erroneous values, or incomplete information, requiring cleaning. For text data, regular expressions are used to remove special characters and correct spelling errors. For image data, image filtering algorithms are used to remove noise points, and image inpainting techniques are used to fill in missing parts. For audio data, noise reduction algorithms are used to remove background noise. This step ensures data quality for subsequent processing and provides a foundation for accurate analysis.

[0021] Feature extraction: Data of different modalities uses different feature representation methods. For text data, word embedding models (such as Word2Vec and BERT) are used to convert text into low-dimensional vectors to extract semantic features. For image data, convolutional neural networks (CNNs) are used to extract visual features such as edges and textures. Audio data is converted to the frequency domain through Fourier transforms to extract frequency features. These feature extraction methods can transform raw data into feature vectors suitable for subsequent processing, highlighting the key information in the data.

[0022] Adaptive multimodal fusion: Dynamic fusion strategy construction: Dynamic fusion strategies are constructed based on different task requirements and data characteristics. First, a fusion strategy library is established, containing various predefined fusion methods, such as attention-based fusion and graph neural network-based fusion. During actual runtime, a task analysis module analyzes the current task and determines its dependence on various modal data. For example, in image description generation tasks, image modal data is more critical, while in sentiment analysis tasks, text modal data carries a relatively larger weight. Based on the analysis results, the most appropriate fusion method is selected from the fusion strategy library.

[0023] Fusion process implementation: Taking attention-based fusion as an example, for text and image modalities, text feature vectors and image feature vectors are first extracted, respectively. The attention mechanism is then introduced to calculate the correlation weights between text and image features. This weighting emphasizes features relevant to the task and suppresses irrelevant information. The weighted text and image features are then concatenated or weighted summed to produce the fused feature vector. This fusion approach fully exploits the potential connections between data from different modalities and dynamically adjusts the contribution of each modality based on task requirements.

[0024] Agent interactive learning: Agent Model Construction: The agent model is constructed using a deep reinforcement learning framework, such as the Proximal Policy Optimization (PPO) algorithm. The agent consists of a policy network and a value network. The policy network generates the agent's action strategy, while the value network evaluates the value of the agent's current state. The agent continuously learns and optimizes its network parameters by interacting with the environment. For example, in a multi-agent collaborative logistics delivery scenario, each agent represents a delivery vehicle. The agent must determine the route and delivery order based on information such as the current cargo location and traffic conditions.

[0025] Interaction Mechanism Design: Agents interact through information sharing and collaboration mechanisms. A shared information space is established where agents can upload their own status information, task progress, and other information to the shared space. They can also access information about other agents from the shared space. During the collaborative process, agents adjust their action strategies based on the shared information. For example, when multiple agents collaborate on a search task, if one agent discovers a clue to the target, it uploads this clue to the shared space. Other agents then use this information to adjust their search direction, improving search efficiency.

[0026] Optimization method: Gradient descent-based parameter updates: During the agent's learning process, the parameters of the policy network and value network are updated using the gradient descent algorithm. By calculating the gradient of the loss function with respect to the network parameters, the parameters are adjusted in the opposite direction of the gradient to gradually reduce the loss function. For example, in the policy network, the policy gradient is calculated based on the agent's actions and environmental feedback, and the network weights are updated to improve the agent's decision-making ability.

[0027] Adaptive learning rate adjustment: To accelerate model convergence and avoid getting stuck in local optima, adaptive learning rate adjustment methods are used. For example, the Adam optimization algorithm dynamically adjusts the learning rate based on the first- and second-order moments of the gradient. Initially, a higher learning rate is used to enable the model to quickly update parameters. As training progresses, the learning rate is gradually reduced to ensure more stable convergence to the optimal solution.

[0028] Example: Example 1: Intelligent security monitoring system Multimodal data acquisition and preprocessing: Data Collection: Multiple high-definition cameras are deployed in the monitoring area for image acquisition, while a microphone array is set up for audio acquisition. The cameras capture the images of the monitoring scene in real time, while the microphones record the surrounding sounds.

[0029] Data cleaning and feature extraction: Images captured by the camera are preprocessed using image denoising algorithms to remove noise caused by lighting, sensor factors, and other factors. Edge detection algorithms are used to extract features such as the outlines of people and objects in the image. For audio data, spectrum analysis techniques are used to remove background noise and extract features such as sound frequency and amplitude. The processed image and audio features are converted into vector form suitable for subsequent processing.

[0030] Adaptive multimodal fusion: Fusion Strategy Selection: Based on the characteristics of security monitoring tasks, the system prioritizes a multimodal fusion strategy based on an attention mechanism. The analysis module evaluates the current monitoring scene to determine whether there are any unusual events (such as intrusions or fighting). If a possible anomaly is detected, attention is further weighted towards the relevant image regions and audio frequency bands.

[0031] Fusion implementation: The extracted image and audio feature vectors are fed into the fusion module, and the correlation weight between them is calculated using the attention mechanism. The weighted feature vectors are concatenated to obtain a fused feature representation, providing more comprehensive information for subsequent agent decision-making.

[0032] Interactive learning and decision-making of intelligent agents: Agent model deployment: Multiple agents are deployed in the security monitoring system, with each agent responsible for monitoring a specific area or performing a specific task (such as target tracking and event alarms). The agents use a deep reinforcement learning model and learn through continuous interaction with the environment (monitoring scene).

[0033] Interaction and Decision-Making: Agents exchange monitoring data and analysis results in real time through a shared information space. When an agent detects an anomaly, it uploads the relevant information to the shared space. Other agents then adjust their monitoring strategies and actions based on this shared information. For example, an agent responsible for monitoring a surrounding area can expand its monitoring range to collaboratively detect additional anomalies. Ultimately, based on the agents' comprehensive assessment, the system decides whether to trigger an alarm or take other countermeasures.

[0034] Example 2: Intelligent Medical Diagnosis Assistance System Multimodal data acquisition and preprocessing: Data collection: Obtain patient medical records from the hospital information system, including symptom descriptions, examination reports, etc. At the same time, collect patient medical imaging data, such as X-rays and CT scans.

[0035] Data cleaning and feature extraction: Medical records undergo preprocessing, including word segmentation and stop word removal. Word embedding models are then used to convert the text into semantic vectors. For medical images, specialized image segmentation algorithms are used to extract regions of interest (e.g., lesions), and convolutional neural networks are used to extract image features such as texture and shape.

[0036] Adaptive multimodal fusion: Fusion strategy selection: For medical diagnosis tasks, the system selects a multimodal fusion strategy based on graph neural networks based on the complexity of the patient's condition and the characteristics of the data. By analyzing keywords in medical records and the characteristics of medical images, a correlation graph is constructed between the data.

[0037] Fusion Implementation: Map text feature vectors and image feature vectors into a graph structure, and leverage the message passing mechanism of graph neural networks to propagate and fuse information between nodes. This approach fully exploits the potential connections between medical record text and medical images, yielding fused features with greater diagnostic value.

[0038] Intelligent agent interactive learning and diagnostic assistance: Agent model construction: In a medical diagnosis assistance system, each agent represents a disease diagnosis model or expert experience module. The agent trains its own model parameters by learning from a large amount of case data.

[0039] Interaction and Diagnostic Assistance: When a new patient case is presented, different agents analyze the multimodal fusion data based on their own diagnostic models. Agents exchange information and share diagnostic ideas and results. For example, an agent responsible for diagnosing cardiovascular disease and another responsible for diagnosing respiratory diseases can exchange insights into patient symptoms and test results, collaborating to provide doctors with more comprehensive and accurate diagnostic recommendations.

[0040] The embodiments of the present invention are presented for purposes of illustration and description and are not intended to be exhaustive or to limit the invention to the disclosed forms. Many modifications and variations will be apparent to those skilled in the art. The embodiments are chosen and described in order to better illustrate the principles of the invention and its practical application and to enable those skilled in the art to understand the invention and design various embodiments with various modifications as suited for specific applications.

Claims

1. An adaptive multimodal fusion agent interactive learning and optimization method, characterized by: It includes multimodal data preprocessing, adaptive multimodal fusion, intelligent agent interactive learning and optimization methods, etc., to achieve multimodal data processing and intelligent agent performance improvement.

2. The method for interactive learning and optimization of an intelligent agent with adaptive multimodal fusion as claimed in claim 1, characterized in that: The multimodal data preprocessing collects and cleans modal data such as text, images, and audio, and uses corresponding models (such as word embedding, CNN, and Fourier transform) to extract features to ensure data quality and availability.

3. The method for interactive learning and optimization of an intelligent agent with adaptive multimodal fusion as claimed in claim 1, characterized in that: The adaptive multimodal fusion constructs a dynamic fusion strategy, selects a suitable method (such as attention mechanism, graph neural network fusion, etc.) from the fusion strategy library through the task analysis module, and implements the fusion process accordingly.

4. The method for interactive learning and optimization of an intelligent agent with adaptive multimodal fusion as claimed in claim 1, characterized in that: The agent interactive learning adopts a deep reinforcement learning framework (such as PPO) to build an agent model, which includes a policy network and a value network, and realizes agent interaction through information sharing and collaboration mechanisms.

5. The adaptive multimodal fusion agent interactive learning and optimization method according to claim 1, characterized in that: The optimization method uses the gradient descent algorithm to update the agent network parameters, and adopts an adaptive learning rate adjustment method (such as the Adam optimization algorithm) to accelerate convergence and avoid local optimality.

6. The adaptive multimodal fusion agent interactive learning and optimization method according to claim 1, characterized in that: The adaptive multimodal fusion strategy can dynamically select and adjust the fusion method according to the task and data, and improve the fusion intelligence and efficiency through the fusion strategy library and task analysis module, which is different from the traditional fixed mode.

7. The adaptive multimodal fusion agent interactive learning and optimization method according to claim 1, characterized in that: The agent interaction mechanism and collaboration optimization design information sharing and collaboration mechanism enable agents to adjust their strategies based on information from other parties, thereby improving the efficiency of multi-agent collaboration in complex tasks.

8. The adaptive multimodal fusion agent interactive learning and optimization method according to claim 1, characterized in that: The combination of gradient descent-based parameter updates and adaptive learning rate adjustments improves the learning speed and convergence stability of the intelligent agent, ensuring interactive learning effects.

9. The adaptive multimodal fusion agent interactive learning and optimization method according to claim 1, characterized in that: Adaptive multimodal data fusion technology innovation builds dynamic fusion strategies to select and adjust fusion methods, explore the intrinsic connections between modalities, and improve fusion quality and efficiency. It is the core protection point of multimodal processing.

10. The adaptive multimodal fusion agent interactive learning and optimization method according to claim 1, characterized in that: The adaptive learning algorithm of the intelligent agent uses deep reinforcement learning and adaptive learning rate adjustment to enable the intelligent agent to optimize itself based on the environment and interactive feedback. It is a key innovation and protection point in artificial intelligence.

Citation Information

Cited By

  • Multi-modal learning data fusion analysis method and system and intelligent device

    CN120930069A