Large model Agent intelligent decision-making method and system fusing multi-modal data
Through the combination of multimodal data fusion and deep learning and reinforcement learning, the problem of insufficient adaptability of intelligent decision-making systems in dynamic environments is solved, and efficient and accurate decision-making results and performance optimization are achieved.
Patent Information
- Application Number
- CN202510459338.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-07-22
AI Technical Summary
Existing intelligent decision-making systems are difficult to make full use of the rich information of multimodal data, and lack adaptability and flexibility in dynamic environments, making it difficult to adjust decision strategies in real time.
Multimodal data fusion technology is used to integrate text, image, and audio data, and decision-making results are generated through deep learning models and reinforcement learning algorithms, and data changes are monitored in real time, model parameters and strategies are dynamically adjusted, and model performance is optimized in combination with feedback.
Improves the accuracy and adaptability of intelligent decision-making, maintains optimal performance in complex tasks and dynamic environments, simplifies system maintenance and reduces maintenance costs.
Smart Images

Figure CN120354943A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of artificial intelligence, multimodal data processing, deep learning, reinforcement learning and intelligent decision-making technology, and specifically to a large-model Agent intelligent decision-making method and system integrating multimodal data. Background Art
[0002] With the rapid development of artificial intelligence technology, intelligent decision-making systems play an increasingly important role in dealing with complex tasks and dynamic environments. However, traditional intelligent decision-making systems can usually only process single-modal data, such as text or images, and it is difficult to fully utilize the rich information of multimodal data. In addition, existing systems often lack sufficient adaptability and flexibility when facing dynamically changing environments, making it difficult to adjust decision-making strategies in real time.
[0003] Therefore, how to improve the performance and adaptability of intelligent decision-making in dealing with complex tasks and dynamic environments is a technical problem that needs to be solved urgently. Summary of the invention
[0004] The technical task of the present invention is to provide a large-model Agent intelligent decision-making method and system that integrates multimodal data to solve the problem of how to improve the performance and adaptability of intelligent decision-making in processing complex tasks and dynamic environments.
[0005] The technical task of the present invention is achieved in the following way: a large model agent intelligent decision-making method integrating multimodal data, the method is as follows:
[0006] Multimodal data fusion: Integrate text, image, and audio data from different modalities, and generate a unified feature representation through feature extraction and feature fusion technology;
[0007] Intelligent decision-making: Decision reasoning is performed based on the fused feature representation, and the final decision result is generated using deep learning models (such as Transformer, BERT, etc.) and reinforcement learning algorithms (such as DQN, PPO, etc.);
[0008] Adaptive learning: real-time monitoring of data changes and decision-making effects, and dynamic adjustment of deep learning model parameters and strategies;
[0009] Feedback optimization: Further optimize the performance of deep learning models by collecting feedback information on decision results.
[0010] As a preferred embodiment, multimodal data fusion is specifically as follows:
[0011] Data collection and preprocessing: Collect multimodal data in real time or in batches from social media, sensor networks, and public databases. Collect data for different data types through API interfaces or crawler tools, and preprocess the collected data of different types to obtain preprocessed data.
[0012] Feature extraction: For data of different modes, corresponding feature extraction technology is used to extract features and obtain corresponding features;
[0013] Feature fusion: Fuse features from different modalities to generate a unified feature representation; and adopt multiple fusion strategies such as early fusion, mid-term fusion and late fusion to ensure the comprehensiveness and effectiveness of the features.
[0014] Preferably, the different types of collected data are preprocessed as follows:
[0015] For text data: Apply natural language processing (NLP) technology (such as NLTK, spaCy) to perform word segmentation, part-of-speech tagging, and named entity recognition operations, and use the BERT model for deep semantic understanding;
[0016] For image data: use OpenCV to perform basic operations such as resizing, cropping, and rotation, as well as object detection and classification based on deep learning methods (such as YOLO and SSD) to provide high-quality input for subsequent feature extraction;
[0017] For audio data: Librosa library is used to perform preprocessing operations such as sampling rate conversion, denoising, and volume normalization. Mel spectrogram conversion technology is used to convert audio signals into a form suitable for machine learning model processing.
[0018] For data of different modes, corresponding feature extraction technology is used to extract features, and the corresponding features are obtained as follows:
[0019] Text feature extraction: In addition to using the BERT model, we also use TF-IDF and Word2Vec traditional NLP methods to supplement feature representation and capture more contextual information.
[0020] Image feature extraction: In addition to ResNet, multiple convolutional neural network models such as Inception-V3 and VGG16 are introduced to select the most suitable architecture according to different scene requirements to achieve more accurate feature capture;
[0021] Audio feature extraction: In addition to the MFCC algorithm, explore the use of Perceptual Linear Prediction (PLP) advanced audio feature extraction technology to improve the ability to understand speech signals.
[0022] Preferably, the deep learning model is as follows:
[0023] Model selection and optimization: The Transformer architecture and its variants (BERT, RoBERTa, T5) are used to process the fused multi-modal feature representations; the deep learning model performs well in natural language processing tasks and is extended to the processing of image and audio data, specifically: ViT (Vision Transformer) is used to process image features, the pre-trained BERT model is used to understand text content, and WaveNet is used to process audio signals;
[0024] Feature interaction and enhancement: To better capture the correlations between different modalities, a cross-modal attention mechanism is introduced, allowing the deep learning model to dynamically adjust the importance weights of different modalities according to the context; in addition, graph neural networks (GNNs) are applied to construct the correlation map between features to further enhance the feature expression ability;
[0025] The reinforcement learning algorithm is as follows:
[0026] Parameter adjustment strategy: In addition to gradient descent algorithms (such as Adam, RMSprop), adaptive learning rate methods (such as AdaGrad, Adadelta) and the latest optimizer such as LAMB (Layer-wise Adaptive Moments for Batch training) are also considered to improve the efficiency and effectiveness of large-scale distributed training;
[0027] Transfer learning and fine-tuning: For specific tasks, the pre-trained model is used as a starting point, and the fine-tuning strategy is used to quickly adapt to the new task requirements. Adversarial training technology is used to reduce the risk of overfitting and improve the generalization ability;
[0028] Dynamic algorithm selection: Based on task requirements and data characteristics, the agent can dynamically select the most suitable learning algorithm from a rich algorithm library; support vector machines (SVM) or random forests (RandomForests) are preferred when processing structured data, while generative adversarial networks (GANs) or diffusion models (Diffusion Models) are used to generate high-quality image or video data for data augmentation or anomaly detection when facing unstructured data.
[0029] Preferably, the adaptive learning is as follows:
[0030] Real-time monitoring and analysis: Use system monitoring tools (Prometheus, Grafana) to collect data changes and decision-making effect indicators in real time; use time series analysis methods (ARIMA model) to predict future trends and provide a basis for subsequent model adjustments; use Grafana or other visualization tools to create real-time dashboards to intuitively display changes in key performance indicators (KPIs), and regularly generate detailed performance analysis reports to help developers understand the operating status of the system and make corresponding adjustments; decision-making effect indicators include hardware performance indicators such as CPU usage, memory usage, network latency, and disk I / O, as well as model performance indicators such as accuracy, recall, and F1 score;
[0031] Dynamic adjustment: Parameter optimization based on monitoring results, that is, based on monitoring results, hyperparameter optimization techniques (such as Baye optimization and random search) are applied to automatically adjust key parameters in the deep learning model; in addition, PPO and GRPO are used to optimize the model training process;
[0032] Anomaly detection: Use integrated machine learning algorithms (Autoencoders) to identify abnormal patterns in data and adjust decision-making strategies in a timely manner; once an anomaly is detected, the corresponding response mechanism is immediately initiated, and an alert is sent through Prometheus to notify relevant personnel and suspend current operations.
[0033] As a preferred embodiment, feedback optimization is as follows:
[0034] Feedback collection: Provide multimodal feedback channels, collect ratings (1-5 stars), check boxes (correct / incorrect marks) and text reviews (NLP sentiment analysis to extract satisfaction) through the user interface (UI); and process real-time feedback streams, use Kafka message queues to receive feedback events, and use Flink to deduplicate (based on UUID), normalize (map to the 0-1 interval) and multimodal association; at the same time, evaluate feedback confidence, use GAN to detect false feedback (such as score-brushing behavior), and filter abnormal data;
[0035] Performance evaluation: Use quantitative indicators such as accuracy, recall, F1 score, and mean square error to evaluate model performance, and use Jenkins to trigger the LMMs-Eval framework regularly to generate a multi-dimensional evaluation report;
[0036] Optimization and adjustment: According to the feedback information and performance evaluation results, the key parameters of the model are adjusted in a targeted manner, and advanced optimization methods such as Bayesian optimization, genetic algorithm, and NAS algorithm are used to find the global optimal solution; when it is found that the existing model performs poorly in certain specific scenarios, the GRPO algorithm is used to continuously improve the decision-making effect through interaction with the environment.
[0037] A large model Agent intelligent decision-making system that integrates multi-modal data. This system uses a centralized configuration management system to uniformly manage the configuration parameters of each module and implements a unified logging and monitoring system. The system includes:
[0038] A multi-modal data fusion module for integrating text, image, and audio data from different modalities and generating a unified feature representation through feature extraction and feature fusion technologies.
[0039] An intelligent decision-making engine for making decision inferences based on the fused feature representation, and using deep learning models (such as Transformer, BERT, etc.) and reinforcement learning algorithms (such as DQN, PPO, etc.) to generate optimal decisions.
[0040] An adaptive learning module for real-time monitoring of data changes and decision-making effects, and dynamically adjusting model parameters and strategies.
[0041] A feedback optimization module for further optimizing the performance of the model by collecting feedback information on decision results.
[0042] Among them, the multi-modal data fusion module adopts various fusion strategies such as early fusion, mid-term fusion, and late fusion; the intelligent decision-making engine combines supervised learning and reinforcement learning, using supervised learning to quickly update the model in the case of labeled data, and using reinforcement learning to optimize strategies in the case of unlabeled data or exploratory tasks; the adaptive learning module uses Prometheus and Grafana for data collection and visualization; the feedback optimization module uses the LMMs-Eval automated testing framework for regular evaluation.
[0043] Preferably, the multi-modal data fusion module includes:
[0044] A data collection and preprocessing sub-module for collecting multi-modal data from different data sources and performing preprocessing.
[0045] A feature extraction sub-module for using specialized feature extraction technologies for data of different modalities.
[0046] A feature fusion sub-module for fusing features of different modalities and generating a unified feature representation.
[0047] The intelligent decision-making engine includes:
[0048] A deep learning model sub-module for processing the fused feature representation using Transformer and BERT models.
[0049] A reinforcement learning algorithm sub-module for optimizing decision-making strategies in combination with DQN and PPO algorithms.
[0050] A decision-making generation sub-module, which is used to generate a final decision result by integrating the outputs of a deep learning model and a reinforcement learning algorithm;
[0051] The adaptive learning module includes:
[0052] A real-time monitoring and analysis sub-module, which is used to collect data changes and decision-making effect indicators in real time through system monitoring tools;
[0053] A dynamic adjustment sub-module, which is used to dynamically adjust model parameters and strategies according to the monitoring results;
[0054] An anomaly detection sub-module, which is used to identify abnormal patterns in data using an integrated machine learning algorithm;
[0055] The feedback optimization module includes:
[0056] A feedback collection sub-module, which is used to collect feedback information from users or the environment on the decision result;
[0057] A performance evaluation sub-module, which is used to evaluate the model performance using accuracy, recall rate, F1 score and other quantitative indicators;
[0058] An optimization and adjustment sub-module, which is used to trigger the retraining or parameter adjustment of the model according to the feedback information and performance evaluation results.
[0059] An electronic device, including: a memory and at least one processor;
[0060] Wherein, a computer program is stored on the memory;
[0061] The at least one processor executes the computer program stored in the memory, so that the at least one processor executes the large model Agent intelligent decision-making method for fusing multi-modal data as described above.
[0062] A computer-readable storage medium, in which a computer program is stored, and the computer program can be executed by a processor to implement the large model Agent intelligent decision-making method for fusing multi-modal data as described above.
[0063] The large model Agent intelligent decision-making method and system for fusing multi-modal data of the present invention have the following advantages:
[0064] (1) Multi-modal data processing ability: The present invention can effectively integrate and process data from different modalities, make full use of the rich information of multi-modal data, and improve the accuracy of decision-making;
[0065] (2) Adaptability and flexibility: Through real-time monitoring and dynamic adjustment, the present invention can maintain the best performance in a dynamic environment and adapt to new data and task requirements;
[0066] (3) High efficiency and stability: The present invention combines a deep learning model and a reinforcement learning algorithm, which can quickly generate accurate decision results and continuously improve performance through a feedback optimization mechanism;
[0067] (4) Reducing maintenance costs: The centralized configuration management, unified logging, and monitoring system of the present invention simplify the system maintenance work; and the automated test framework and real-time performance monitoring reduce the workload of manual monitoring and troubleshooting;
[0068] (5) The present invention improves the performance and adaptability of intelligent decision-making in dealing with complex tasks and dynamic environments through multi-modal data processing, deep learning models, and reinforcement learning algorithms, and improves the accuracy and efficiency of intelligent decision-making;
[0069] (6) Compared with the prior art, the present invention can effectively process multi-modal data, improve the accuracy and adaptability of decision-making, and is applicable to complex and changeable environments and task requirements. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] The present invention will be further described below with reference to the accompanying drawings.
[0071] Attached Figure 1 is a flowchart of the intelligent decision-making method of the large model Agent for fusing multi-modal data;
[0072] Attached Figure 2 is a schematic diagram of the process of intelligent decision-making;
[0073] Attached Figure 3 is a schematic diagram of the structure of the large model Agent intelligent decision-making system for fusing multi-modal data. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0074] The intelligent decision-making method and system of the large model Agent for fusing multi-modal data of the present invention will be described in detail below with reference to the accompanying drawings of the specification and specific embodiments.
[0075] Example 1:
[0076] As shown in the attached Figure 1 figure, this embodiment provides an intelligent decision-making method of the large model Agent for fusing multi-modal data, and the method is as follows:
[0077] S1. Multi-modal data fusion: Integrate text, image, and audio data from different modalities, and generate a unified feature representation through feature extraction and feature fusion techniques;
[0078] S2, Intelligent Decision-making: Decision reasoning is performed based on the fused feature representation, and the final decision result is generated using deep learning models (such as Transformer, BERT, etc.) and reinforcement learning algorithms (such as DQN, PPO, etc.);
[0079] S3, Adaptive learning: real-time monitoring of data changes and decision-making effects, and dynamic adjustment of deep learning model parameters and strategies;
[0080] S4. Feedback optimization: Further optimize the performance of the deep learning model by collecting feedback information on decision results.
[0081] The multimodal data fusion in step S1 of this embodiment is specifically as follows:
[0082] S101. Data collection and preprocessing: Collect multimodal data in real time or in batches from social media, sensor networks, and public databases. Collect data for different data types through API interfaces or crawler tools, and preprocess the collected data of different types to obtain preprocessed data.
[0083] S102, feature extraction: for data of different modes, corresponding feature extraction technology is used to extract features and obtain corresponding features;
[0084] S103, feature fusion: fuse features of different modalities to generate a unified feature representation; and adopt multiple fusion strategies such as early fusion, mid-term fusion and late fusion to ensure the comprehensiveness and effectiveness of the features.
[0085] In step S101 of this embodiment, the data preprocessing of different types of collected data is specifically as follows:
[0086] S10101. For text data: Apply natural language processing (NLP) technology (such as NLTK, spaCy) to perform word segmentation, part-of-speech tagging, and named entity recognition, and use the BERT model for deep semantic understanding;
[0087] S10102. For image data: use OpenCV to perform basic operations such as resizing, cropping, and rotation, as well as object detection and classification based on deep learning methods (such as YOLO and SSD) to provide high-quality input for subsequent feature extraction;
[0088] S10103. For audio data: Use the Librosa library to perform preprocessing operations such as sampling rate conversion, denoising, and volume normalization. At the same time, use Mel spectrum conversion technology to convert audio signals into a form suitable for machine learning model processing.
[0089] For the data of different modalities in step S102 of this embodiment, corresponding feature extraction techniques are used for feature extraction, and the corresponding features are obtained as follows:
[0090] S10201. Text feature extraction: In addition to using the BERT model, traditional NLP methods such as TF-IDF and Word2Vec are also combined to supplement the feature representation to capture more context information;
[0091] S10202. Image feature extraction: In addition to ResNet, multiple convolutional neural network models such as Inception-V3 and VGG16 are introduced, and the most suitable architecture is selected according to different scenario requirements to achieve more accurate feature capture;
[0092] S10203. Audio feature extraction: In addition to the MFCC algorithm, the use of the Perceptual Linear Prediction (PLP) advanced audio feature extraction technology is explored to improve the understanding ability of speech signals.
[0093] As shown in the appendix Figure 2 The deep learning model in step S2 of this embodiment is as follows:
[0094] S2-101. Model selection and optimization: The Transformer architecture and its variants (BERT, RoBERTa, T5) are used to process the fused multi-modal feature representation; The deep learning model performs well in natural language processing tasks and is extended to the processing of image and audio data. Specifically, ViT (Vision Transformer) is used to process image features, the pre-trained BERT model is used to understand text content, and WaveNet is used to process audio signals;
[0095] S2-102. Feature interaction and enhancement: In order to better capture the correlation between different modalities, a cross-modal attention mechanism is introduced to allow the deep learning model to dynamically adjust the importance weights of different modalities according to the context; In addition, graph neural networks (GNNs) are also applied to construct a correlation map between features to further enhance the feature expression ability.
[0096] The reinforcement learning algorithm in step S2 of this embodiment is as follows:
[0097] S2-201, Parameter Adjustment Strategy: In addition to gradient descent algorithms (such as Adam, RMSprop), adaptive learning rate methods (such as AdaGrad, Adadelta) and the latest optimizers such as LAMB (Layer-wise Adaptive Moments for Batch training) are also considered to improve the efficiency and effectiveness of large-scale distributed training;
[0098] S2-202, Transfer Learning and Fine-tuning: For specific tasks, use the pre-trained model as a starting point and quickly adapt to the new task requirements through the fine-tuning strategy. Use adversarial training techniques to reduce the risk of overfitting and improve the generalization ability;
[0099] S2-203, Dynamic Algorithm Selection: Based on task requirements and data characteristics, the agent can dynamically select the most suitable learning algorithm from a rich algorithm library; when dealing with structured data, support vector machine (SVM) or random forests are preferred, while when facing unstructured data, generative adversarial networks (GANs) or diffusion models (Diffusion Models) are used to generate high-quality image or video data for data augmentation or anomaly detection.
[0100] The adaptive learning in step S3 of this embodiment is specifically as follows:
[0101] S301, Real-time Monitoring and Analysis: Use system monitoring tools (Prometheus, Grafana) to collect data changes and decision-making effect indicators in real time; and use time series analysis methods (ARIMA model) to predict future trends to provide a basis for subsequent model adjustment; at the same time, create a real-time dashboard with Grafana or other visualization tools to intuitively display the changes of key performance indicators (KPIs), and generate detailed performance analysis reports regularly to help developers understand the running state of the system and make corresponding adjustments; among them, the decision-making effect indicators include hardware performance indicators such as CPU usage, memory occupancy, network latency, and disk I / O, as well as model performance indicators such as accuracy, recall rate, and F1 score;
[0102] S302, Dynamic Adjustment: According to the monitoring results, optimize the parameters, that is, based on the monitoring results, apply hyperparameter optimization techniques (such as Bayesian optimization, random search) to automatically adjust the key parameters in the deep learning model; in addition, PPO and GRPO are also used to optimize the model training process;
[0103] S303, Anomaly Detection: Use integrated machine learning algorithms (Autoencoders) to identify abnormal patterns in data and adjust decision-making strategies in a timely manner; once an anomaly is detected, immediately initiate the corresponding response mechanism, send an alert through Prometheus to notify relevant personnel and suspend the current operation.
[0104] The feedback optimization in step S4 of this embodiment is specifically as follows:
[0105] S401, Feedback Collection: Provide multimodal feedback channels, collect ratings (1-5 stars), check boxes (correct / incorrect marks) and text reviews (NLP sentiment analysis to extract satisfaction) through the user interface (UI); and process real-time feedback streams, use Kafka message queues to receive feedback events, and use Flink to deduplicate (based on UUID), normalize (map to the 0-1 interval) and multimodal association; at the same time, evaluate feedback confidence, use GAN to detect false feedback (such as score-brushing behavior), and filter abnormal data;
[0106] S402, Performance evaluation: Use quantitative indicators such as accuracy, recall, F1 score, and mean square error to evaluate model performance, and use Jenkins to trigger the LMMs-Eval framework regularly to generate a multi-dimensional evaluation report;
[0107] S403, Optimization and Adjustment: According to the feedback information and performance evaluation results, the key parameters of the model are adjusted in a targeted manner, and advanced optimization methods such as Bayesian optimization, genetic algorithm, and NAS algorithm are used to find the global optimal solution; when it is found that the existing model performs poorly in certain specific scenarios, the GRPO algorithm is used to continuously improve the decision-making effect through interaction with the environment.
[0108] Embodiment 2:
[0109] As attached Figure 3 As shown, this embodiment provides a large model Agent intelligent decision-making system integrating multimodal data. The system adopts a centralized configuration management system to uniformly manage the configuration parameters of each module and implement a unified log recording and monitoring system; the system includes:
[0110] Multimodal data fusion module, which is used to integrate text, image, and audio data from different modalities and generate a unified feature representation through feature extraction and feature fusion technology;
[0111] Intelligent decision engine, which is used to make decision reasoning based on the fused feature representation, and uses deep learning models (such as Transformer, BERT, etc.) and reinforcement learning algorithms (such as DQN, PPO, etc.) to generate optimal decisions;
[0112] An adaptive learning module for real-time monitoring of data changes and decision-making effects, and dynamically adjusting model parameters and strategies;
[0113] A feedback optimization module for further optimizing the performance of the model by collecting feedback information on decision-making results;
[0114] Among them, the multi-modal data fusion module adopts various fusion strategies such as early fusion, mid-term fusion, and late fusion; the intelligent decision-making engine combines supervised learning and reinforcement learning, uses supervised learning to quickly update the model in the case of labeled data, and uses reinforcement learning to optimize strategies in the case of unlabeled data or exploratory tasks; the adaptive learning module uses Prometheus and Grafana for data collection and visualization; the feedback optimization module uses the LMMs-Eval automated test framework for regular evaluation.
[0115] The multi-modal data fusion module in this embodiment includes:
[0116] A data collection and preprocessing sub-module for collecting multi-modal data from different data sources and performing preprocessing;
[0117] A feature extraction sub-module for using specialized feature extraction techniques for data of different modalities;
[0118] A feature fusion sub-module for fusing features of different modalities to generate a unified feature representation.
[0119] The intelligent decision-making engine in this embodiment includes:
[0120] A deep learning model sub-module for processing the fused feature representation using Transformer and BERT models;
[0121] A reinforcement learning algorithm sub-module for optimizing decision-making strategies in combination with DQN and PPO algorithms;
[0122] A decision generation sub-module for synthesizing the outputs of the deep learning model and the reinforcement learning algorithm to generate the final decision result.
[0123] The adaptive learning module in this embodiment includes:
[0124] A real-time monitoring and analysis sub-module for real-time collecting data change and decision-making effect indicators through system monitoring tools;
[0125] A dynamic adjustment sub-module for dynamically adjusting model parameters and strategies according to the monitoring results;
[0126] An anomaly detection sub-module for identifying abnormal patterns in data using integrated machine learning algorithms.
[0127] The feedback optimization module in this embodiment includes:
[0128] A feedback collection sub-module, configured to collect feedback information on the decision-making result from users or the environment;
[0129] A performance evaluation sub-module, configured to evaluate the model performance using accuracy, recall rate, and F1-score quantification metrics;
[0130] An optimization and adjustment sub-module, configured to trigger retraining of the model or parameter adjustment according to the feedback information and performance evaluation results.
[0131] Example 3:
[0132] This example also provides an electronic device, including: a memory and a processor;
[0133] Wherein, the memory stores computer-executable instructions;
[0134] The processor executes the computer-executable instructions stored in the memory, so that the processor executes the large model Agent intelligent decision-making method for fusing multi-modal data in any embodiment of the present invention.
[0135] The processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0136] The memory can be used to store computer programs and / or modules. The processor realizes various functions of the electronic device by running or executing the computer programs and / or modules stored in the memory, and by calling the data stored in the memory. The memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function, etc.; the data storage area can store data created according to the use of the terminal, etc. In addition, the memory may further include high-speed random access memory, and may also include non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash memory card, at least one magnetic disk storage period, a flash memory device, or other volatile solid-state storage devices.
[0137] Example 4:
[0138] This embodiment also provides a computer-readable storage medium, which stores multiple instructions that are loaded by a processor to cause the processor to execute the large model Agent intelligent decision-making method for fusing multi-modal data in any embodiment of the present invention. Specifically, a system or device equipped with a storage medium can be provided, on which software program codes for implementing the functions of any one of the above embodiments are stored, and the computer (or CPU or MPU) of the system or device is caused to read and execute the program codes stored in the storage medium.
[0139] In this case, the program code read from the storage medium itself can implement the functions of any one of the above embodiments, so the program code and the storage medium storing the program code constitute a part of the present invention.
[0140] Examples of the storage medium for providing the program code include floppy disks, hard disks, magneto-optical disks, optical disks (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RYM, DVD-RW, DVD+RW), magnetic tapes, non-volatile memory cards, and ROMs. Optionally, the program code can be downloaded from a server computer via a communication network.
[0141] Furthermore, it should be clear that not only can the functions of any one of the above embodiments be realized by executing the program code read by the computer, but also by causing an operating system or the like operating on the computer based on the instructions of the program code to complete part or all of the actual operations.
[0142] In addition, it can be understood that the program code read from the storage medium is written into the memory provided in the expansion board inserted into the computer or into the memory provided in the expansion unit connected to the computer, and then based on the instructions of the program code, the CPU or the like installed on the expansion board or the expansion unit is caused to execute part and all of the actual operations, thereby realizing the functions of any one of the above embodiments.
[0143] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for intelligent decision-making of a large model Agent that integrates multi-modal data, characterized in that, The method is as follows: Multimodal data fusion: Integrate text, image, and audio data from different modalities, and generate a unified feature representation through feature extraction and feature fusion technology; Intelligent decision-making: Decision reasoning is performed based on the fused feature representation, and the final decision result is generated using deep learning models and reinforcement learning algorithms; Adaptive learning: real-time monitoring of data changes and decision-making effects, and dynamic adjustment of deep learning model parameters and strategies; Feedback optimization: Further optimize the performance of deep learning models by collecting feedback information on decision results.
2. The large model Agent intelligent decision-making method for fusing multi-modal data according to claim 1, wherein, The details of multimodal data fusion are as follows: Data collection and preprocessing: Collect multimodal data in real time or in batches from social media, sensor networks, and public databases. Collect data for different data types through API interfaces or crawler tools, and preprocess the collected data of different types to obtain preprocessed data. Feature extraction: For data of different modes, corresponding feature extraction technology is used to extract features and obtain corresponding features; feature Fusion: Fusion the features of different modalities to generate a unified feature representation; and adopt multiple fusion strategies such as early fusion, mid-term fusion and late fusion to ensure the comprehensiveness and effectiveness of the features.
3. The intelligent decision-making method of the large model Agent for fusing multi-modal data according to claim 2, wherein The data preprocessing of different types of collected data is as follows: For text data: Apply natural language processing technology to perform word segmentation, part-of-speech tagging, and named entity recognition, and use the BERT model for deep semantic understanding; For image data: use OpenCV to perform basic operations such as resizing, cropping, and rotation, as well as object detection and classification based on deep learning methods to provide high-quality input for subsequent feature extraction; For audio data: Librosa library is used to perform preprocessing operations such as sampling rate conversion, denoising, and volume normalization. Mel spectrogram conversion technology is used to convert audio signals into a form suitable for machine learning model processing. For data of different modes, corresponding feature extraction technology is used to extract features, and the corresponding features are obtained as follows: Text feature extraction: In addition to using the BERT model, we also use TF-IDF and Word2Vec traditional NLP methods to supplement feature representation and capture more contextual information. Image feature extraction: In addition to ResNet, multiple convolutional neural network models such as Inception-V3 and VGG16 are introduced to select the most suitable architecture according to different scene requirements to achieve more accurate feature capture; Audio feature extraction: In addition to the MFCC algorithm, explore the use of Perceptual Linear Prediction advanced audio feature extraction technology to improve the ability to understand speech signals.
4. The large model Agent intelligent decision-making method for fusing multi-modal data according to claim 1, wherein, The deep learning model is as follows: Model Selection and Optimization: The Transformer architecture and its variants are adopted to process the fused multi-modal feature representations; deep learning models perform excellently in natural language processing tasks and are extended to the processing of image and audio data. Specifically, ViT is used to process image features, the pre-trained BERT model is used to understand text content, and WaveNet is used to process audio signals. Feature Interaction and Enhancement: The cross-modal attention mechanism is introduced to allow deep learning models to dynamically adjust the importance weights of different modalities according to the context; in addition, graph neural networks are applied to construct the association graph between features to further enhance the feature expression ability. The reinforcement learning algorithm is as follows: Parameter Adjustment Strategy: In addition to the gradient descent algorithm, adaptive learning rate methods and the latest optimizers such as LAMB are also considered. Transfer Learning and Fine-tuning: For specific tasks, the pre-trained model is used as a starting point, and the fine-tuning strategy is used to quickly adapt to the new task requirements, using adversarial training techniques. Dynamic Algorithm Selection: Based on task requirements and data characteristics, the agent can dynamically select the most suitable learning algorithm from a rich algorithm library; support vector machines or random forests are preferred when processing structured data, while generative adversarial networks or diffusion models are used to generate high-quality image or video data for data augmentation or anomaly detection when facing unstructured data.
5. The large model Agent intelligent decision-making method for fusing multi-modal data according to claim 1, characterized in that Adaptive Learning is as follows: Real-time Monitoring and Analysis: System monitoring tools are used to collect data changes and decision-making effect indicators in real time; time series analysis methods are adopted to predict future trends; at the same time, Grafana or other visualization tools are used to create a real-time dashboard to intuitively display the changes of key performance indicators and generate detailed performance analysis reports regularly; among them, the decision-making effect indicators include hardware performance indicators such as CPU usage, memory occupancy, network latency, and disk I / O, as well as model performance indicators such as accuracy, recall rate, and F1 score. Dynamic Adjustment: According to the monitoring results, parameter optimization is carried out, that is, based on the monitoring results, hyperparameter optimization techniques are applied to automatically adjust the key parameters in the deep learning model; in addition, PPO and GRPO are used to optimize the model training process. Anomaly Detection: Integrated machine learning algorithms are used to identify abnormal patterns in data and adjust decision-making strategies in a timely manner; once an anomaly is detected, the corresponding response mechanism is immediately activated, and alerts are sent through Prometheus to notify relevant personnel and suspend the current operation.
6. The large model Agent intelligent decision-making method for fusing multi-modal data according to claim 1, wherein, Feedback Optimization is as follows: Feedback Collection: Multimodal feedback channels are provided to collect ratings, checkboxes, and text evaluations through the user interface; real-time feedback stream processing is carried out, and Kafka message queues are used to receive feedback events, which are de-duplicated, normalized, and multimodally associated through Flink; at the same time, feedback confidence evaluation is carried out, and GAN is used to detect false feedback and filter abnormal data. Performance Evaluation: Quantification indicators such as accuracy, recall rate, F1 score, and mean squared error are used to evaluate the model performance, and the LMMs-Eval framework is triggered regularly through Jenkins to generate multi-dimensional evaluation reports. Optimization and adjustment: According to the feedback information and performance evaluation results, the key parameters of the model are adjusted specifically. Advanced optimization methods such as Bayesian optimization, genetic algorithms, and NAS algorithms are used to find the global optimal solution.
7. A large model Agent intelligent decision-making system that integrates multi-modal data, characterized in that, This system uses a centralized configuration management system to uniformly manage the configuration parameters of each module and implements a unified logging and monitoring system. The system includes: A multi-modal data fusion module for integrating text, image, and audio data from different modalities and generating a unified feature representation through feature extraction and feature fusion techniques. An intelligent decision-making engine for making decision inferences based on the fused feature representation and generating optimal decisions using deep learning models and reinforcement learning algorithms. An adaptive learning module for real-time monitoring of data changes and decision-making effects and dynamically adjusting model parameters and strategies. A feedback optimization module for further optimizing the performance of the model by collecting feedback information on decision results. Among them, the multi-modal data fusion module adopts various fusion strategies such as early fusion, mid-term fusion, and late fusion; the intelligent decision-making engine combines supervised learning and reinforcement learning, using supervised learning to quickly update the model in the case of labeled data and using reinforcement learning to optimize strategies in the case of unlabeled data or exploratory tasks; the adaptive learning module uses Prometheus and Grafana for data collection and visualization; the feedback optimization module uses the LMMs-Eval automated test framework for regular evaluation.
8. The large model Agent intelligent decision-making system for fusing multi-modal data according to claim 7, characterized in that, The multi-modal data fusion module includes: A data collection and preprocessing sub-module for collecting multi-modal data from different data sources and performing preprocessing. A feature extraction sub-module for using specialized feature extraction techniques for data of different modalities. A feature fusion sub-module for fusing features of different modalities and generating a unified feature representation. The intelligent decision-making engine includes: A deep learning model sub-module for processing the fused feature representation using Transformer and BERT models. A reinforcement learning algorithm sub-module for optimizing decision-making strategies by combining DQN and PPO algorithms. A decision generation sub-module for synthesizing the outputs of the deep learning model and the reinforcement learning algorithm to generate the final decision result. The adaptive learning module includes: A real-time monitoring and analysis sub-module for real-time collecting data change and decision-making effect indicators through system monitoring tools. A dynamic adjustment sub-module for dynamically adjusting model parameters and strategies according to the monitoring results. An anomaly detection sub-module for identifying abnormal patterns in data using integrated machine learning algorithms. The feedback optimization module includes: A feedback collection sub-module for collecting feedback information from users or the environment on decision results. A performance evaluation sub-module for evaluating model performance using accuracy, recall rate, and F1 score quantization metrics. An optimization and adjustment sub-module for triggering the retraining or parameter adjustment of the model according to the feedback information and performance evaluation results.
9. An electronic device, characterized in that, Includes: A memory and at least one processor; Among them, a computer program is stored on the memory. The at least one processor executes the computer program stored in the memory, such that the at least one processor executes the large model Agent intelligent decision-making method for fusing multimodal data according to any one of claims 1 to 6.
10. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, and the computer program can be executed by a processor to implement the large model Agent intelligent decision-making method for fusing multimodal data according to any one of claims 1 to 6.
Citation Information
Cited By
Interpretable cross-media agent decision path tracking method
CN120632598A
Water affair decision-making system based on multi-modal model fusion
CN120634051A
Automatic case self-healing method and system for operating system
CN120849300A
Emergency data fusion and decision-making method under cross-modal dynamic routing mechanism
CN120995384A
Multimodal deep neural network model, system and method based on continuous learning
CN120996965A