Multi-modal large model local training-based ai agent agent

By locally training an AI agent based on a multimodal large model, the problems of excessive noise and low information attention in model pre-training are solved, the effectiveness of data processing and the practicality of the agent are achieved, and the efficiency and accuracy of reasoning are improved.

CN120744485AInactive Publication Date: 2025-10-03SHANGHAI BIFANG RONGXIANG INTELLIGENT TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510726639.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-03
Publication Date
2025-10-03
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the existing technology, the direct processing of different modal data during model pre-training leads to a lot of noise, the model does not pay much attention to key multimodal information, and many hallucinations occur during the training process.

Method used

It adopts an AI agent based on local training of a multimodal large model, including modal data collection, processing and analysis, data transmission, model construction and training units, combined with computer vision, audio sensors and tactile sensors, to achieve information interaction between modalities through co-attention, use reinforcement learning and low-rank key-value compression technology to optimize the algorithm, and combine rule methods and unsupervised fine-tuning to achieve cross-domain generalization capabilities.

Benefits of technology

It improves the effectiveness and accuracy of data processing, ensures the reliability of data sources, enhances the practicality and applicability of intelligent agents, continuously improves model performance through optimized feedback and iterative updates, and improves reasoning efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120744485A_ABST
    Figure CN120744485A_ABST
Patent Text Reader

Abstract

The invention discloses an ai agent intelligent agent based on local training of a multi-modal large model, and relates to the technical field of intelligent agents. Under the action of the data processing and analyzing unit, repeated data can be deleted, and the validity of the data is ensured; the data source reliability is ensured; missing values in the data are processed, data standardization processing is carried out, abnormal value detection is also carried out, and the accuracy of the data is ensured; the model performance is continuously improved by combining the optimization feedback unit with the manually labeled feedback information, and periodic updating is performed based on the optimization feedback unit, so that the effectiveness of the intelligent agent can be ensured, and the practicability and applicability of the intelligent agent can be further ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent agent technology, and in particular to an aiagent intelligent agent based on local training of a multimodal large model. Background Art

[0002] AI agents, or artificial intelligence agents, often referred to as agents, are systems capable of perceiving their environment, making decisions, and taking actions. These systems can perform passive tasks, proactively seek solutions to problems, adapt to environmental changes, and make decisions without direct human intervention. AI agents can understand natural language, intelligently generate responses, and execute specific actions, possessing key characteristics such as autonomy, memory, planning, tools, and action.

[0003] Currently, during model pre-training, the received data of different modalities is directly processed. However, this direct processing approach results in more noise during training, the model pays less attention to key multimodal information, and the training model suffers from more hallucinations. To this end, we propose an AI agent based on local training of a large multimodal model. Summary of the Invention

[0004] The purpose of the present invention is to solve the problems mentioned in the above background technology. The present invention provides an AI agent based on local training of a multimodal large model.

[0005] In order to achieve the above-mentioned purpose, the present invention specifically adopts the following technical solutions:

[0006] An AI agent intelligent body based on local training of a multimodal large model, the AI ​​agent intelligent body comprising:

[0007] Modal data collection unit, used to realize the collection of multimodal data;

[0008] a data processing and analysis unit, configured to process and analyze the data collected by the modal data collection unit;

[0009] A data transmission unit, used for real-time data transmission;

[0010] A model building unit, which builds a basic model based on the data processed and analyzed by the data processing and analysis unit;

[0011] The model training unit trains the collected data and adjusts the model parameters based on the machine learning algorithm.

[0012] Furthermore, the data collected by the modal data collection unit includes text data, image data, audio data and video data, and the modal data collection unit is based on a combination of computer vision, audio sensors and tactile sensors.

[0013] Furthermore, the modal data collection unit is based on a multimodal pre-training model, and the multimodal pre-training model is based on a dual-stream structure.

[0014] Furthermore, the multimodal pre-training model construction includes the following steps:

[0015] Step A1: Establish pre-training tasks, including image-text matching and contrastive learning, to learn universal representations from large-scale image-text pairs;

[0016] Step A2: Build a dataset containing image pairs and text pairs for pre-training and fine-tuning.

[0017] Step A3: Architecture design: Use co-attention to achieve information interaction between modalities and improve cross-modal understanding capabilities.

[0018] Furthermore, the ai agent intelligent body also includes: an intelligent control unit, a remote control unit and a voice control unit, wherein the intelligent control unit is used to realize the manipulation of the intelligent body; the remote control unit is connected to an external remote control to realize remote control of the intelligent body; and the voice control unit is used to realize voice control of the intelligent body.

[0019] Furthermore, the ai agent intelligent body also includes: a decision generation unit and a command execution unit, wherein the decision generation unit is used to generate the intelligent body's operation decision and issue an operation command; the command execution unit is used to execute the operation command invented by the decision generation unit.

[0020] Furthermore, the decision generation unit includes a perception module, a decision mechanism module, and an execution feedback module. The perception module is responsible for collecting raw data from the environment, such as sensor readings, user input, or messages from other intelligent agents, and converting it into an internally recognizable form, integrating and analyzing the perceived multimodal data, extracting key information, and updating the internal state. The decision mechanism includes a decision model module and a goal-oriented module. The decision model module is constructed based on internal state and external input based on reinforcement learning. The goal-oriented module is used to decompose the overall goal into sub-goals and formulate an action plan. The execution feedback module is used to convert the results into specific action instructions, and continuously monitor the action effects through the perception module, comparing the actual results with the expected goals to form feedback signals. The feedback is used to adjust the decision strategy and achieve self-optimization. (Through reinforcement learning, algorithms can be optimized to improve reasoning efficiency and accuracy; combined with low-rank key-value compression technology and hybrid expert models, hardware limitations can be overcome; auxiliary loss-free load balancing technology and synthetic data strategies are used to achieve low-cost training and deployment; reinforcement learning combined with rule methods and unsupervised fine-tuning enables the model to perform outstandingly in multiple fields such as natural language processing, code generation, and image generation, and has cross-domain generalization capabilities).

[0021] Furthermore, the ai agent also includes: an optimization feedback unit and an iterative update unit. The optimization feedback unit combines manually labeled feedback information to continuously improve model performance, and the iterative update unit performs periodic updates based on the optimization feedback unit.

[0022] Furthermore, the data processing and analysis unit includes the following steps:

[0023] Step B1: Data preprocessing: converting multimodal data into a format suitable for model input, removing duplicates and filtering the collected data;

[0024] Step B2: Data optimization, data collection and integration;

[0025] Step B3: Data cleaning, processing missing values ​​in the data, performing data standardization, and performing outlier detection;

[0026] Step B4: extract features and perform feature extraction on the processed data;

[0027] Step B5: feature fusion, fusing the extracted feature data.

[0028] The beneficial effects of the present invention are as follows:

[0029] 1. The present invention can delete duplicate data under the action of the data processing and analysis unit to ensure the validity of the data; ensure the reliability of the data source; process missing values ​​in the data, perform data standardization, and perform outlier detection to ensure the accuracy of the data.

[0030] 2. The present invention continuously improves the model performance by optimizing the feedback unit in combination with manually annotated feedback information, and performs periodic updates based on the optimized feedback unit, which can ensure the effectiveness of the intelligent agent and thus ensure the practicality and applicability of the intelligent agent. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 It is a working block diagram of the present invention;

[0032] Figure 2 This is a workflow diagram for constructing a multimodal pre-training model in the present invention;

[0033] Figure 3 It is a working block diagram of the decision-making generation unit in the present invention;

[0034] Figure 4 It is a workflow diagram of the data processing and analysis unit in the present invention. DETAILED DESCRIPTION

[0035] To make the objectives, technical solutions and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.

[0036] See also Figure 1 - Figure 4 The present invention provides an AI agent based on local training of a multimodal large model, and the AI ​​agent includes:

[0037] Modal data collection unit, used to realize the collection of multimodal data;

[0038] A data processing and analysis unit, used to process and analyze the data collected by the modal data collection unit;

[0039] A data transmission unit, used for real-time data transmission;

[0040] The model building unit constructs a basic model using the data processed and analyzed by the data processing and analysis unit;

[0041] The model training unit trains the collected data and adjusts the model parameters based on the machine learning algorithm.

[0042] In this embodiment, preferably, the data collected by the modal data collection unit includes text data, image data, audio data and video data, and the modal data collection unit is based on a combination of computer vision, audio sensors and tactile sensors.

[0043] In this embodiment, preferably, the modal data collection unit is based on a multimodal pre-training model, and the multimodal pre-training model is based on a dual-stream structure.

[0044] In this embodiment, preferably, the multimodal pre-training model construction includes the following steps:

[0045] Step A1: Establish pre-training tasks, including image-text matching and contrastive learning, to learn universal representations from large-scale image-text pairs;

[0046] Step A2: Build a dataset containing image pairs and text pairs for pre-training and fine-tuning.

[0047] Step A3: Architecture design: Use co-attention to achieve information interaction between modalities and improve cross-modal understanding capabilities.

[0048] In this embodiment, preferably, the ai agent intelligent body also includes: an intelligent control unit, a remote control unit and a voice control unit, wherein the intelligent control unit is used to realize the manipulation of the intelligent body; the remote control unit is connected to an external remote control to realize remote control of the intelligent body; and the voice control unit is used to realize voice control of the intelligent body.

[0049] In this embodiment, preferably, the ai agent intelligent body also includes: a decision generation unit and a command execution unit, wherein the decision generation unit is used to generate the intelligent body's operation decision and issue the operation command; the command execution unit is used to execute the operation command invented by the decision generation unit.

[0050] In this embodiment, preferably, the decision generation unit includes a perception module, a decision mechanism module and an execution feedback module, wherein the perception module is responsible for collecting raw data from the environment, such as sensor readings, user input or messages from other intelligent agents, and converting it into an internally recognizable form, integrating and analyzing the perceived multimodal data, extracting key information and updating the internal state; the decision mechanism includes a decision model module and a goal-oriented module, wherein the decision model module is constructed based on reinforcement learning according to the internal state and external input, and the goal-oriented module is used to decompose the overall goal into sub-goals and formulate an action plan; the execution feedback module is used to convert the results into specific action instructions, and continuously monitor the action effects through the perception module, compare the actual results with the expected goals, form a feedback signal, and the feedback is used to adjust the decision strategy and achieve self-optimization.

[0051] Reinforcement learning can optimize algorithms and improve reasoning efficiency and accuracy; combined with low-rank key-value compression technology and hybrid expert models, it can break through hardware limitations; auxiliary loss-free load balancing technology and synthetic data strategies are used to achieve low-cost training and deployment; reinforcement learning combined with rule-based methods and unsupervised fine-tuning enables the model to perform outstandingly in multiple fields such as natural language processing, code generation, image generation, etc., and has cross-domain generalization capabilities.

[0052] In this embodiment, preferably, the ai agent also includes: an optimization feedback unit and an iterative update unit. The optimization feedback unit combines the manually labeled feedback information to continuously improve the model performance, and the iterative update unit performs periodic updates based on the optimization feedback unit.

[0053] In this embodiment, preferably, the data processing and analysis unit includes the following steps:

[0054] Step B1: Data preprocessing: converting multimodal data into a format suitable for model input, removing duplicates and filtering the collected data;

[0055] Step B2: Data optimization, data collection and integration;

[0056] Step B3: Data cleaning, processing missing values ​​in the data, performing data standardization, and performing outlier detection;

[0057] Step B4: extract features and perform feature extraction on the processed data;

[0058] Step B5: feature fusion, fusing the extracted feature data.

[0059] The working principle and use process of the present invention:

[0060] A modal data collection unit is used to collect multimodal data. The data collected by the modal data collection unit includes text data, image data, audio data, and video data. The modal data collection unit is based on a combination of computer vision, audio sensors, and tactile sensors. The modal data collection unit is based on a multimodal pre-training model, and the multimodal pre-training model is based on a dual-stream structure.

[0061] The construction of a multimodal pre-training model includes the following steps:

[0062] Step A1: Establish pre-training tasks, including image-text matching and contrastive learning, to learn universal representations from large-scale image-text pairs;

[0063] Step A2: Build a dataset containing image pairs and text pairs for pre-training and fine-tuning.

[0064] Step A3: Architecture design: Use co-attention to achieve information interaction between modalities and improve cross-modal understanding capabilities.

[0065] A data processing and analysis unit, used to process and analyze the data collected by the modal data collection unit;

[0066] The data processing and analysis unit includes the following steps:

[0067] Step B1: Data preprocessing: converting multimodal data into a format suitable for model input, deduplicating and filtering the collected data, removing duplicate data, and ensuring data validity;

[0068] Step B2: Data optimization: collect and integrate data to ensure the data source is reliable;

[0069] Step B3: Data cleaning, processing missing values ​​in the data, performing data standardization, and performing outlier detection to ensure data accuracy;

[0070] Step B4: extract features and perform feature extraction on the processed data;

[0071] Step B5: feature fusion, fusing the extracted feature data.

[0072] A data transmission unit, used for real-time data transmission;

[0073] The model building unit constructs a basic model using the data processed and analyzed by the data processing and analysis unit;

[0074] The model training unit trains the collected data and adjusts the model parameters based on the machine learning algorithm.

[0075] Intelligent control unit, used to realize the control of intelligent body.

[0076] Remote control unit, an external remote control is used to remotely control the intelligent agent.

[0077] Voice control unit, used to implement voice-controlled intelligent agents.

[0078] The decision generation unit is used to generate the operation decision of the intelligent agent and issue the operation command;

[0079] The decision generation unit includes a perception module, a decision mechanism module, and an execution feedback module. The perception module is responsible for collecting raw data from the environment, such as sensor readings, user input, or messages from other agents, and converting it into an internally recognizable form. It integrates and analyzes the perceived multimodal data (such as text, images, and sounds), extracts key information, and updates the internal state.

[0080] The decision-making mechanism includes a decision model module and a goal-oriented module. The decision model module is constructed based on reinforcement learning according to internal states and external inputs. Reinforcement learning can optimize the algorithm and improve reasoning efficiency and accuracy. Combining low-rank key-value compression technology and hybrid expert models can break through hardware limitations. Auxiliary loss-free load balancing technology and synthetic data strategies are used to achieve low-cost training and deployment. Reinforcement learning is combined with rule methods and unsupervised fine-tuning to enable the model to perform outstandingly in multiple fields such as natural language processing, code generation, and image generation, and has cross-domain generalization capabilities.

[0081] The goal-oriented module is used to decompose the overall goal into sub-goals and formulate an action plan; the execution feedback module is used to convert the results into specific action instructions, such as turning the mobile robot, sending messages or adjusting parameters, and drive the intelligent agent to execute, and continuously monitor the action effect through the perception module, compare the actual results with the expected goals, and form feedback signals. The feedback is used to adjust the decision-making strategy and achieve self-optimization.

[0082] The command execution unit is used to execute the operation command invented by the decision generation unit.

[0083] Optimize the feedback unit and combine it with manually annotated feedback information to continuously improve model performance.

[0084] The iterative update unit performs periodic updates based on the optimized feedback unit.

[0085] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not limited to the embodiments shown herein but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. An AI agent based on local training of a multimodal large model, characterized by: The AI ​​agent includes: Modal data collection unit, used to realize the collection of multimodal data; a data processing and analysis unit, configured to process and analyze the data collected by the modal data collection unit; A data transmission unit, used for real-time data transmission; A model building unit, which builds a basic model based on the data processed and analyzed by the data processing and analysis unit; The model training unit trains the collected data and adjusts the model parameters based on the machine learning algorithm.

2. The AI ​​agent based on local training of a multimodal large model according to claim 1, characterized in that: The data collected by the modal data collection unit includes text data, image data, audio data and video data, and the modal data collection unit is based on a combination of computer vision, audio sensors and tactile sensors.

3. The AI ​​agent based on local training of a multimodal large model according to claim 1, characterized in that: The modal data collection unit is based on a multimodal pre-training model, and the multimodal pre-training model is based on a dual-stream structure.

4. The AI ​​agent based on local training of a multimodal large model according to claim 3 is characterized in that: The multimodal pre-training model construction includes the following steps: Step A1: Establish pre-training tasks, including image-text matching and contrastive learning, to learn universal representations from large-scale image-text pairs; Step A2: Build a dataset containing image pairs and text pairs for pre-training and fine-tuning. Step A3: Architecture design: Use co-attention to achieve information interaction between modalities and improve cross-modal understanding capabilities.

5. The AI ​​agent based on local training of a multimodal large model according to claim 1, characterized in that: The ai agent also includes: an intelligent control unit, a remote control unit and a voice control unit, wherein the intelligent control unit is used to realize the manipulation of the agent; the remote control unit is connected to an external remote controller to realize remote control of the agent; and the voice control unit is used to realize voice control of the agent.

6. The AI ​​agent based on local training of a multimodal large model according to claim 1, characterized in that: The ai agent also includes: a decision generation unit and a command execution unit, wherein the decision generation unit is used to generate the agent's operation decision and issue an operation command; the command execution unit is used to execute the operation command invented by the decision generation unit.

7. The AI ​​agent based on local training of a multimodal large model according to claim 6, characterized in that: The decision generation unit includes a perception module, a decision mechanism module and an execution feedback module, wherein the perception module is responsible for collecting raw data from the environment, such as sensor readings, user input or messages from other intelligent agents, and converting it into an internally recognizable form, integrating and analyzing the perceived multimodal data, extracting key information and updating the internal state; the decision mechanism includes a decision model module and a goal-oriented module, wherein the decision model module is constructed based on reinforcement learning according to the internal state and external input, and the goal-oriented module is used to decompose the overall goal into sub-goals and formulate an action plan; the execution feedback module is used to convert the results into specific action instructions, and continuously monitor the action effects through the perception module, compare the actual results with the expected goals, and form a feedback signal. The feedback is used to adjust the decision strategy and achieve self-optimization.

8. The AI ​​agent based on local training of a multimodal large model according to claim 1, characterized in that: The AI ​​agent also includes: an optimization feedback unit and an iterative update unit. The optimization feedback unit combines manually labeled feedback information to continuously improve model performance, and the iterative update unit performs periodic updates based on the optimization feedback unit.

9. The AI ​​agent based on local training of a multimodal large model according to claim 1, characterized in that: The data processing and analysis unit includes the following steps: Step B1: Data preprocessing: converting multimodal data into a format suitable for model input, removing duplicates and filtering the collected data; Step B2: Data optimization, data collection and integration; Step B3: Data cleaning, processing missing values ​​in the data, performing data standardization, and performing outlier detection; Step B4: extract features and perform feature extraction on the processed data; Step B5: feature fusion, fusing the extracted feature data.

Citation Information

Cited By

  • Intelligent agent continuous training and effect evaluation closed-loop method and system fusing work order feedback, and medium

    CN121835735A