Method and system for collecting decision data by intelligent agents driven by multimodal large models
Through the multimodal large model-driven intelligent body method, intelligent trolleys can effectively integrate multimodal information and optimize decision data, solving the efficiency and accuracy challenges of decision-making systems in the existing technology in complex environments, and achieving efficient and accurate decision-making and task execution.
Patent Information
- Application Number
- CN202510287600.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-03-12
AI Technical Summary
The existing smart car decision-making system has efficiency and accuracy challenges in converging multimodal information and optimizing decision data, especially in complex environments and multi-round interactive scenarios.
The multimodal large model-driven agent method is adopted to optimize the decision distribution by obtaining the information entropy and normalized mass scores of sensor data, and to optimize the decision model and feedback learning are performed by applying time-sensitive multimodal alignment processing and Bayesian neural networks.
It realizes the efficient decision-making process of smart cars in complex environments, improves the accuracy and reliability of decision-making, enhances the adaptability and learning ability of the system, and ensures the efficiency and accuracy of task execution.
Smart Images

Figure CN119808006B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of intelligent transportation and automation, and in particular to a method and system for collecting decision data by an intelligent agent driven by a multi-modal large model. Background Art
[0002] With the continuous development of artificial intelligence technology, multimodal large models are increasingly used in various intelligent systems, especially in the fields of smart car decision-making and autonomous driving. The decision-making system of a smart car usually relies on a variety of sensor data, such as images, radar, temperature, sound, etc., which can provide the smart car with a comprehensive perception of its environment. In recent years, intelligent agents based on vision-language models (VLMs) have gradually become one of the core technologies for improving the intelligence level of automated systems. In particular, the powerful capabilities of large models in visual understanding, natural language processing, and decision reasoning have been proven to significantly improve the decision-making efficiency and accuracy of smart cars in complex environments.
[0003] However, there are two common problems in the existing intelligent car decision-making system: one is how to reasonably integrate multimodal information to make the best decision in different environments; the other is how to effectively collect and use decision data for further optimization and learning in multiple rounds of interaction. Although multimodal information fusion has a certain research foundation, how to ensure efficient integration and intelligent reasoning of various types of information while ensuring system efficiency is still a challenge to be solved. Summary of the invention
[0004] The purpose of the invention is to provide a method and system for collecting decision data by an intelligent agent driven by a multimodal large model, so as to solve the above-mentioned problems existing in the prior art.
[0005] The technical solution is a method for collecting decision data by an intelligent agent driven by a multimodal large model, comprising the following steps:
[0006] Obtain the raw sensor data of the predetermined sensors of the smart car, calculate the information entropy and normalized quality score of each raw sensor data; perform adaptive multimodal information fusion based on the information entropy and normalized quality score to generate a fusion feature vector;
[0007] Apply time-sensitive multimodal alignment processing to the fused feature vector and generate time-aligned features through timestamp embedding and time-series self-attention mechanism;
[0008] Based on the time series alignment features, the decision distribution is generated through the Bayesian neural network; based on the decision distribution, the epistemic uncertainty and random uncertainty are calculated and risk assessment is performed to generate the final decision and uncertainty indicator set;
[0009] The final decision, the uncertainty indicator set and the pre-stored task description are combined to evaluate the task criticality score; based on the task criticality score and the pre-stored system resource status, the task execution plan is calculated;
[0010] Read the task execution plan and initialize the decision model through meta-learning; identify data gap areas and generate synthetic feedback data; combine the priority execution comparison feedback in the task execution plan to learn and optimize the decision model, and store the execution experience in the task experience library.
[0011] A system for collecting decision data driven by a multimodal large model-driven intelligent agent, including:
[0012] at least one processor; and,
[0013] a memory communicatively connected to at least one of the processors; wherein,
[0014] The memory stores instructions that can be executed by the processor, and the instructions are used to be executed by the processor to implement the method of collecting decision data for an intelligent agent driven by a multimodal large model.
[0015] Beneficial effect: The present invention makes full use of multimodal input to realize an efficient decision-making process of the smart car; at the same time, through multiple rounds of interaction between the smart car and the environment, the decision-making strategy is continuously optimized, and detailed decision data is collected for subsequent learning and analysis, thereby further improving the execution capability and task completion accuracy of the smart car. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 A flowchart of the steps of a method for collecting decision data by an intelligent agent driven by a multimodal large model provided in an embodiment of the present application.
[0017] Figure 2 A flowchart of the steps of adaptive multimodal information fusion provided in an embodiment of the present application.
[0018] Figure 3 A flowchart of the steps for generating temporal alignment features through timestamp embedding and temporal self-attention mechanism provided in an embodiment of the present application.
[0019] Figure 4 A flowchart of the steps for generating a decision distribution through a Bayesian network and generating a final decision and uncertainty indicator set is provided in an embodiment of the present application. DETAILED DESCRIPTION
[0020] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.
[0021] It should be noted that in order to clearly show the steps of this application, serial numbers are marked for each step in the specification. These serial numbers are only used for the convenience of explanation and do not limit the order of execution of the steps. In actual operation, according to the technical requirements of the specific implementation scenario, the steps can be executed in a different order than that shown in the specification, and in some cases, parallel processing between steps can be achieved.
[0022] like Figure 1 As shown, the present application proposes a method for collecting decision data by an intelligent agent driven by a multimodal large model, comprising the following steps:
[0023] S1. Obtaining raw sensor data of predetermined sensors of the smart car, calculating information entropy and normalized quality score of each raw sensor data; performing adaptive multimodal information fusion based on information entropy and normalized quality score to generate a fusion feature vector;
[0024] S2, applying time-sensitive multimodal alignment processing to the fused feature vector, and generating time-aligned features through timestamp embedding and time-series self-attention mechanism;
[0025] S3. Generate decision distribution through Bayesian neural network based on time series alignment features; calculate epistemic uncertainty and random uncertainty based on decision distribution and perform risk assessment to generate the final decision and uncertainty indicator set;
[0026] S4, combining the final decision, the uncertainty indicator set and the pre-stored task description to evaluate the task criticality score; and calculating the task execution plan based on the task criticality score and the pre-stored system resource status;
[0027] S5. Read the task execution plan and initialize the decision model through meta-learning; identify data gap areas and generate synthetic feedback data; optimize the decision model by combining the priority execution comparison feedback learning in the task execution plan, and store the execution experience in the task experience library.
[0028] In this embodiment, the original sensor data of multiple sensors of the smart car are obtained, and the quality of each sensor data is evaluated to obtain a normalized quality score; based on the normalized quality score and information entropy, the low-quality mode is identified, and the compensation data is generated by historical data compensation; the original sensor data of the high-quality mode and the compensation data of the low-quality mode are input into the corresponding feature extraction network to obtain the feature vector of each mode; finally, the fusion weight is calculated based on the normalized quality score and information entropy, and weighted fusion is performed to obtain the fused feature vector.
[0029] The time difference matrix is calculated based on the timestamp of the sensor data of each sensor and converted into a time embedding vector; the time embedding vector is added to the fused feature vector to generate time series enhancement features; for sensor data with inconsistent sampling rates, progressive feature interpolation is applied to obtain the feature time window set; the query, key and value matrices of time series attention are generated from multi-scale time series features, and the time difference weights are introduced to calculate the time series sensitive attention score; the multi-head time series attention mechanism is applied to process the multi-scale time series features to obtain the time series self-attention features; the reconciliation mechanism is applied to the time series self-attention features of different time scales, and the time series relationship reasoning is performed through the graph attention network to finally generate the time series alignment features.
[0030] Define the decision space according to the task type and construct a Bayesian neural network as the decision model; extract the prior rule set from the domain knowledge base and convert it into a prior probability distribution; perform multiple sampling of the Bayesian neural network weights, use the time series alignment features as input to generate multiple decision samples, and estimate the decision probability distribution (decision distribution); calculate the cognitive uncertainty caused by insufficient knowledge and the random uncertainty caused by environmental randomness respectively; determine the dynamic decision threshold according to the safety level of the task and the current system status; generate the final decision and the related uncertainty indicator set based on uncertainty and risk assessment, which will be used for mission criticality assessment and resource allocation in subsequent steps.
[0031] Receive the final decision and uncertainty indicator set, and evaluate the criticality and resource requirements of the current decision in combination with the task description; extract key attributes from the task description, evaluate the correlation between the task and safety to obtain a safety correlation score; calculate the time urgency of the task, and calculate the task criticality score by integrating multiple factors; monitor the resource status of the smart car and calculate resource availability; estimate the required computing resources based on the decision type and uncertainty level; establish a computing demand mapping table, and select the appropriate model complexity level based on the task criticality score and resource availability; analyze the potential contribution of each modality to reducing decision uncertainty, evaluate the importance of each modality to the current task, generate a modal importance vector, and calculate the activation level of each modality; allocate computing resources according to the activation level, and optimize the task execution plan based on the task criticality score and dependency, and pass the resource allocation plan and task execution plan to the next step to guide the feedback learning process.
[0032] Receive resource allocation plans and task execution plans, and adjust the scope and depth of meta-learning and knowledge transfer based on resource constraints; extract task representation vectors from task descriptions and environment state vectors, and retrieve similar tasks from the task experience library; use resource-aware meta-learning algorithms to initialize model parameters for the current task based on experience data from similar tasks; adjust data analysis granularity based on available computing resources, analyze the feedback data coverage of the current task, and identify data gap areas; train conditional generative adversarial networks to generate synthetic feedback data for data gaps; select a subset of feedback data to be processed based on the priority in the task execution plan, classify the feedback data into positive and negative feedback, and evaluate the reliability score of each feedback; construct a contrastive learning batch to calculate feature contrast loss, and apply the policy gradient method to update the decision model; store the execution experience of the current task in the task experience library, and periodically analyze the data in the library to extract experience patterns and integrate them into the domain knowledge base.
[0033] This embodiment can effectively integrate data from different types of sensors by calculating information entropy and normalized quality scores and then performing adaptive multimodal information fusion, thereby reducing the limitations and uncertainties of a single sensor, and making the generated fusion feature vector more comprehensive and accurate in reflecting the environment and state of the smart car, providing a richer and more reliable basis for subsequent decision-making. Multimodal alignment processing is performed, and time-series alignment features are generated with the help of timestamp embedding and time-series self-attention mechanisms, which can not only capture the dynamic changes of data in the time dimension, but also strengthen the temporal correlation of different modal data, which helps the model to better understand the temporal laws of data and improve the perception of complex environments and dynamic scenes. The decision distribution is generated by the Bayesian neural network. The probabilistic characteristics of the Bayesian neural network enable the model to quantify the uncertainty of the decision results, provide a more reliable decision-making basis, and avoid misjudgments caused by a single deterministic decision. Risk assessment is performed by calculating cognitive uncertainty and random uncertainty, which enables the smart car to fully consider various uncertainty factors in the decision-making process, make more reasonable decisions according to the degree of risk, and improve the ability to deal with complex and unknown situations. By evaluating the criticality score of the task and calculating the task execution plan according to the system resource status, resources can be reasonably allocated according to the importance of the task and the actual situation of the system resources to ensure efficient and orderly execution of the task. Through meta-learning and comparative feedback learning, the model can continuously learn and adapt to new environments and tasks, and continuously improve decision-making performance and generalization capabilities. This embodiment improves the environmental perception ability of the smart car, the accuracy and reliability of decision-making, the ability to deal with uncertainty, and the efficiency of task execution through multimodal information fusion, time series feature processing, uncertainty assessment, task planning, and model optimization. At the same time, it enables the decision-making model to have the ability of self-learning and optimization to adapt to changing task requirements and environmental conditions.
[0034] According to one aspect of the present application, step S1 further comprises:
[0035] S11. Obtain raw sensor data such as camera data, lidar data, microphone data, temperature sensor data, and inertial measurement unit (IMU) inertial data from the sensor array of the smart car, and record the sensor data timestamp of each sensor data; construct an environment state vector based on the raw sensor data, and retrieve scenarios similar to the current environment state vector from historical records to obtain environmental influencing factors; calculate the signal-to-noise ratio and data integrity of each sensor data, and analyze the timing stability index and spectrum consistency of the sensor data; integrate the above indicators, calculate the comprehensive quality score of each sensor data, and normalize it to the [0, 1] interval to obtain the normalized quality score.
[0036] S12. Calculate the information entropy of the original sensor data of each sensor to measure the uncertainty of the data; based on the normalized quality score and information entropy, identify low-quality modes through threshold judgment and set a low-quality flag; for modes marked as low-quality, retrieve historical records similar to the current environmental state vector from the historical database to obtain similar historical data; calculate the compensation coefficient, the larger the value, the lower the quality of the original data, and more historical data is needed for compensation; weightedly fuse the original sensor data and similar historical data according to the compensation coefficient to generate compensated data; apply the corresponding feature extraction network to the original data (high-quality mode) or compensated data (low-quality mode) of each sensor to obtain the feature vector of each mode; calculate the fusion weight based on the normalized quality score and information entropy, perform weighted fusion on the feature vectors of each mode, and generate a fused feature vector; at the same time, record the fusion weight and modal quality index of each mode for subsequent analysis.
[0037] like Figure 2 As shown, according to one aspect of the present application, the step of adaptive multimodal information fusion includes:
[0038] Based on the raw sensor data, the environmental state vector is constructed, and the sensor data quality is evaluated. The signal-to-noise ratio, data integrity, timing stability index and spectrum consistency are calculated to obtain the normalized quality score; the information entropy of the raw sensor data is calculated; for the raw sensor data with a normalized quality score lower than the preset threshold or an information entropy higher than the preset threshold, similar scene data is retrieved from the historical database and compensation data is generated; the raw sensor data and the compensation data are input into the corresponding feature extraction network to obtain the feature vectors of each modality; the fusion weight is calculated based on the normalized quality score and information entropy, and the feature vectors of each modality are weightedly fused to generate a fused feature vector.
[0039] In one embodiment of the present application, raw data is obtained from the sensor array of the smart car (including camera data X_vis, laser radar data X_lidar, microphone data X_audio, temperature sensor data X_temp, IMU inertial data X_imu, etc.). Each sensor data is timestamped and recorded to obtain the sensor data timestamp T_i (i=1, 2, ..., N, where N is the total number of sensors).
[0040] Construct the environmental state vector E_state based on the current sensor data, including factors such as lighting conditions, weather conditions, and movement speed. Retrieve scenarios similar to the current environmental state vector E_state from historical records to obtain the environmental impact factor E_impact, which represents the degree of influence of various environmental factors on the performance of different sensors. Calculate the signal-to-noise ratio SNR_i of each sensor data: SNR_i = 10log10(P_signal_i / P_noise_i); where P_signal_i is the effective signal power of sensor i, and P_noise_i is the noise power. Detect data integrity C_i and identify missing values, outliers, and discontinuities: C_i = 1 (N_missing + w_a N_anomaly) / N_total; where N_missing is the number of missing data points, N_anomaly is the number of abnormal data points, w_a is the weight coefficient of the abnormal points, and N_total is the total number of data points. Calculate the temporal stability index S_i of the sensor data: S_i = 1 avg(|X_i(t) X_i(t-1)|) / range(X_i); where X_i(t) represents the data of sensor i at time t, and range(X_i) represents the value range of the sensor i data. Evaluate the spectral consistency F_i of the data and detect sudden changes or abnormal frequency components through time-frequency analysis.
[0041] Integrate the above indicators and calculate the comprehensive quality score Q_i of each sensor data: Q_i = w_snr SNR_i +w_c C_i + w_s S_i + w_f F_i w_e E_impact_i; where w_snr, w_c, w_s, w_f, w_e are the weight coefficients of each indicator, which are obtained through historical data optimization. Normalize the comprehensive quality score Q_i to the interval [0, 1] to obtain the normalized quality score Q_norm_i.
[0042] The information entropy H_i is calculated for each sensor data: H_i = -∑(p(x_i)log(p(x_i))), where p(x_i) is the probability distribution of the sensor i data value x_i, estimated by the data histogram. The lower the entropy value, the higher the data certainty; the higher the entropy value, the higher the data uncertainty. Based on the normalized quality score Q_norm_i and the information entropy H_i, identify low-quality modes: LQ_flag_i = (Q_norm_i < τ_q) || (H_i > τ_h), where τ_q is the quality score threshold and τ_h is the entropy value threshold. These two thresholds are dynamically adjusted according to task requirements. For modes marked as low quality, similar scene data are retrieved from the historical database DB_hist: X_similar = RetrieveTopK(DB_hist, E_state, K), where the RetrieveTopK function retrieves the K most similar historical records based on the current environment state vector E_state. Calculate the compensation coefficient α_i: α_i = 1 Q_norm_i; the larger the α_i value, the lower the quality of the original data, and more historical data is needed for compensation. Generate compensated data X_comp_i: X_comp_i = (1 α_i)X_i + α_iX_similar_i; where X_similar_i is the retrieved similar historical data. Apply the feature extraction network to the original data (high-quality modality) or compensated data (low-quality modality) of each sensor: F_i = Encoder_i(X_i_final); where X_i_final is the final sensor data (original or compensated), Encoder_i is the feature extractor of the corresponding modality (such as CNN for vision, Transformer for audio, etc.), and obtain the feature vector F_i of each modality, with the dimension unified as d dimensions. The fusion weight is calculated based on the normalized quality score Q_norm_i and the information entropy H_i: W_i = Q_norm_iexp(-βH_i) / ∑(Q_norm_jexp(-βH_j)), where β is the entropy weight coefficient, which controls the influence of entropy on the weight distribution. Perform weighted fusion to generate the fusion feature vector F_fusion: F_fusion = ∑(W_iF_i). Record the fusion weight W_i and modal quality index M_quality (including the normalized quality score Q_norm_i and information entropy H_i) of each modality for subsequent analysis.
[0043] This embodiment can measure the quality of sensor data from different angles by calculating multi-dimensional indicators such as signal-to-noise ratio, data integrity, timing stability index and spectrum consistency. The normalized quality score obtained by combining these indicators can comprehensively and accurately reflect the quality of each sensor data. The uncertainty of the data is judged based on the information entropy, and the low-quality mode is identified in combination with the normalized quality score. For low-quality data, similar data is retrieved from the historical database for weighted fusion compensation, which effectively reduces the data error caused by factors such as sensor failure and environmental interference, and improves the accuracy and reliability of the data. Obtaining raw data from multiple sensors of the smart car and constructing an environmental state vector can more comprehensively describe the environmental information of the smart car. By retrieving similar situations to obtain environmental influencing factors, the smart car can better understand the current environmental conditions and improve its perception and adaptability to complex environments. Feature extraction is performed on the original data or compensated data, and then the fusion weight is calculated based on the normalized quality score and information entropy. The feature vectors of each mode are weighted fused to generate a fused feature vector. This multimodal feature fusion method can combine the advantages of different sensors, more comprehensively and accurately reflect environmental characteristics, and improve the environmental perception accuracy of the smart car. This embodiment provides a rich basis for subsequent data analysis, algorithm optimization, and decision-making of the smart car, which helps to continuously improve the performance and behavior of the smart car.
[0044] According to one aspect of the present application, step S2 is further:
[0045] S21. Based on the timestamp of the sensor data of each sensor, the time difference matrix between each pair of modes is calculated, and it is normalized to obtain the normalized time difference; the normalized time difference is converted into a time embedding vector using the sinusoidal position encoding method; the time embedding vector is added to the feature vector of each mode to obtain the time series enhancement feature; the sampling rate of each sensor is identified, a sampling rate ratio vector is constructed, and a feature space interpolation function is applied to the sensor data with a lower sampling rate to generate a progressive feature interpolation; a feature time window set containing windows with different time spans is constructed, and the time series enhancement features are aggregated for each time window to generate multi-scale time series features.
[0046] S22. Generate query, key and value matrices from multi-scale temporal features as the input of the attention mechanism; introduce temporal difference weights, adjust the attention score according to the absolute value of the temporal difference, and calculate the temporal sensitive attention score; implement a multi-head temporal attention mechanism, divide the features into multiple heads, each head independently calculates attention and processes information, and then merges the results to obtain a multi-head temporal attention; apply the multi-head temporal attention mechanism to process multi-scale temporal features, combine residual connections and layer normalization, and obtain temporal self-attention features.
[0047] S23. Apply a reconciliation mechanism to temporal self-attention features of different time scales, dynamically adjust the weights of each time window according to task requirements, and generate temporal feature reconciliation; construct a temporal graph structure in which nodes are features and edges are temporal relationships, and apply a graph attention network to perform temporal relationship reasoning to obtain temporal relationship enhanced features; introduce a task description embedding representing the current task goal, and screen and strengthen the temporal relationship enhanced features through a cross-attention mechanism to obtain task goal-oriented features; process the task goal-oriented features through a feedforward neural network to obtain the final temporal alignment features; and store the temporal alignment features and the temporal relationship data during their generation for subsequent analysis and optimization.
[0048] like Figure 3 As shown in Figure 2, the steps of generating time-series alignment features through timestamp embedding and time-series self-attention mechanism include:
[0049] The time difference matrix is calculated based on the timestamp of the original sensor data, and a time embedding vector is generated; the time embedding vector is added to the fused feature vector to generate a time series enhancement feature; the time series enhancement feature is aggregated over time windows to generate multi-scale time series features; the time difference weighted self-attention mechanism is applied to the multi-scale time series features to generate time series self-attention features; a time series graph structure is constructed based on the time series self-attention features and a graph attention network is applied to perform time series relationship reasoning, and the time series alignment features are generated by combining the pre-stored task description embedding.
[0050] In one embodiment of the present application, based on the sensor data timestamp T_i of each sensor, the time difference matrix ΔT_ij between each pair of modalities is calculated: ΔT_ij=T_i T_j(i, j=1, 2, ..., N). The time difference matrix ΔT_ij is normalized to obtain the normalized time difference ΔT_norm_ij, and the range is controlled within [-1, 1]. The normalized time difference ΔT_norm_ij is converted into a time embedding vector T_emb_ij using the sinusoidal position encoding method: T_emb_ij[2k]=sin(ΔT_norm_ij / 10000 (2k / d_model) );T_emb_ij[2k+1]=cos(ΔT_norm_ij / 10000 (2k / d_model) ), where k is the position index and d_model is the embedding dimension. The time embedding vector T_emb_i is added to each modal feature vector F_i to obtain the temporal enhancement feature F_t_i: F_t_i = F_i + T_emb_i.
[0051] Identify the sampling rate of each sensor and construct the sampling rate ratio vector R_sample. For sensor data with lower sampling rates, apply progressive feature interpolation F_interp_i: F_interp_i(t') = Interpolate(F_t_i(t1), F_t_i(t2), t'); where t1 and t2 are two sampling time points near t', and Interpolate is the feature space interpolation function. Construct a feature time window set W_t, which contains windows of different time spans (such as 100ms, 300ms, 1s, etc.). For each time window w, aggregate the temporal enhancement features F_t_i to generate multi-scale temporal features F_ms_i: F_ms_i(w)=TemporalAggregate(F_t_i, w); where TemporalAggregate is a temporal aggregation function, which can be implemented using convolution, pooling, or attention mechanisms.
[0052] Generate query, key and value matrices from multi-scale temporal features F_ms_i: Q_i = W_q F_ms_i; K_i = W_kF_ms_i; V_i = W_v F_ms_i; where W_q, W_k, W_v are learnable parameter matrices. Introduce the time difference weight W_t_ij to adjust the attention score: W_t_ij = exp(-γ|ΔT_norm_ij|), where γ is the time sensitivity parameter that controls the degree of influence of time difference on attention. Calculate the time-sensitive attention score A_ij: A_ij = softmax((Q_iK_j T / sqrt(d_k))W_t_ij); where d_k is the dimension of the key vector, T Represents transposition. Implement the multi-head temporal attention mechanism MHTA and divide the features into h heads: head_l = Attention(Q_i l , K_j l , V_j l , W_t_ij); MHTA = Concat(head_1,head_2,...,head_h)W_o; where l is the head index and W_o is the output projection matrix. Apply the multi-head temporal attention mechanism MHTA to process the multi-scale temporal features F_ms_i and obtain the temporal self-attention features F_sa: F_sa = LayerNorm(F_ms + MHTA(F_ms)); where LayerNorm is the layer normalization operation and F_ms is the set of multi-scale temporal features F_ms_i of all modalities.
[0053] A reconciliation mechanism is applied to the temporal self-attention features F_sa(w) of different time scales: F_harmonized = ∑(α_wF_sa(w)), where α_w is the weight of the time window w, and F_harmonized is the reconciled feature, which is dynamically adjusted according to the task requirements. A temporal graph structure G_t is constructed, where nodes are features and edges are temporal relationships: G_t = BuildTemporalGraph(F_harmonized, ΔT_norm_ij); where BuildTemporalGraph is a temporal graph construction function; the graph attention network GAT is applied to perform temporal relationship reasoning, and the temporal relationship enhanced feature F_tre is obtained: F_tre = GAT(G_t).
[0054] The task description embedding T_emb is introduced to represent the current task goal. Through the cross-attention mechanism, the temporal relationship enhancement feature F_tre is screened and enhanced based on the task description embedding T_emb: F_task=CrossAttention(T_emb, F_tre, F_tre), where T_emb is the query, F_tre is the key and value, and CrossAttention is the cross-attention function. The task goal-oriented feature F_task is processed through a feedforward neural network to obtain the final temporal alignment feature F_aligned: F_aligned = FFN(F_task), where FFN is a two-layer feedforward neural network with a ReLU activation function. The temporal alignment feature F_aligned and the temporal relationship data R_temporal during its generation process are stored for subsequent analysis and optimization.
[0055] This embodiment calculates the time difference matrix between each pair of modalities and converts it into a time embedding vector, which can accurately capture the time difference between different modal data, integrate time information into feature representation, enhance the time series perception ability of features, and enable the model to better understand the time context of data. By constructing a feature time window set to analyze data from multiple time scales, the time series patterns and trends of different granularities are captured, and the model's modeling ability for complex time series information is improved. The time difference weight is introduced to adjust the attention score, so that when the model calculates attention, it can assign different weights according to the absolute value of the time difference, pay more attention to time-related data, and enhance the sensitivity of the attention mechanism to time series information. By implementing a multi-head time series attention mechanism, the expressive power of the model is increased, time series information can be captured from different angles, the modeling ability of complex time series relationships is improved, and the model is more flexible and robust when processing multi-scale time series features. The reconciliation mechanism is applied to the time series self-attention features of different time scales and the graph attention network is applied, so that the model can better understand and utilize the time series relationships in the data and mine hidden time series patterns and dependencies. The cross-attention mechanism ensures that the features extracted by the model are closely centered around the current task objectives, improves the pertinence and effectiveness of the features, and enables the model to better serve specific tasks.
[0056] According to one aspect of the present application, step S3 is further:
[0057] S31. According to the task type, define a decision space containing all possible decision options, define a decision parameter space for continuous control tasks, and define a decision category space for discrete decision tasks; construct a Bayesian neural network whose weights and biases are probability distributions rather than fixed values; extract a priori rule set related to the current task from the domain knowledge base and convert it into a priori probability distribution for initializing the weight distribution of the Bayesian neural network; sample the weights of the Bayesian neural network multiple times to obtain multiple model instances, and use time series alignment features as input to generate multiple decision samples; based on these decision samples, estimate the decision probability distribution, and calculate the mean and variance of the decision distribution.
[0058] S32. Uncertainty decomposition and risk assessment: Estimate the epistemic uncertainty caused by insufficient knowledge or data, and calculate the variance of the decision expectation of each model instance and the overall expectation; estimate the random uncertainty caused by environmental randomness, and calculate the decision variance within each model instance; define the possible decision result set and its corresponding cost function, and calculate the risk score for each decision option; calculate the epistemic uncertainty threshold and the total uncertainty threshold based on the safety level of the task and the current system status; based on uncertainty and risk assessment, generate the final decision and its related set of uncertainty indicators, and take conservative decisions or request more information when the uncertainty is too high or the risk is too great.
[0059] like Figure 4 As shown in FIG, the steps of generating a decision distribution through a Bayesian network and generating a final decision and uncertainty indicator set include:
[0060] Construct a Bayesian neural network as a decision model; extract a priori rule set from the domain knowledge base and convert it into a priori probability distribution, and initialize the Bayesian neural network; perform a predetermined number of sampling on the weights of the Bayesian neural network to obtain a predetermined number of model instances; and estimate the decision distribution in combination with the time series alignment feature; calculate the cognitive uncertainty caused by insufficient knowledge or data and the random uncertainty caused by environmental randomness based on the decision distribution: perform risk assessment based on cognitive uncertainty and random uncertainty to obtain a risk score; determine the dynamic decision threshold based on the pre-stored task safety level and system status, and generate the final decision and uncertainty indicator set according to the type of uncertainty and the risk score; the dynamic decision threshold includes the total uncertainty threshold and the cognitive uncertainty threshold.
[0061] In one embodiment of the present application, a decision space Ω_D is defined according to the task type, which contains all possible decision options. For continuous control tasks (such as steering, acceleration), a decision parameter space Θ_D is defined. For discrete decision tasks (such as path selection, action type), a decision category space C_D is defined. A Bayesian neural network BNN is constructed, whose weights and biases are probability distributions rather than fixed values: W ~ N(μ_w, σ_w 2 ); b ~ N(μ_b, σ_b 2 ), where N represents the normal distribution, μ and σ are the mean and standard deviation of the distribution respectively. The prior rule set R_prior related to the current task is extracted from the domain knowledge base KB_domain. The prior rule set R_prior is converted into a prior probability distribution P_prior(Θ) to initialize the weight distribution of the Bayesian neural network BNN.
[0062] The weights of the Bayesian neural network BNN are sampled M times to obtain M model instances: Θ_m ~ P(Θ) (m= 1, 2, ..., M). For each model instance, the time-series alignment feature F_aligned is used as input to generate a decision sample: D_m = BNN(F_aligned; Θ_m), where D_m represents the decision generated by the mth model instance. Based on M decision samples, the decision probability distribution P(D|F_aligned) is estimated: P(D|F_aligned) ≈ (1 / M)∑Δ(D D_m); where Δ is the Dirac function, which can be replaced by kernel density estimation or histogram in actual calculations. Calculate the mean μ_D and variance σ_D of the decision distribution 2 :μ_D = (1 / M)∑D_m;σ_D 2= (1 / M)∑(D_m μ_D) 2 .
[0063] Estimate epistemic uncertainty U_epistemic due to insufficient knowledge or data: U_epistemic = (1 / M)∑(E[D|F_aligned,Θ_m] E[D|F_aligned]) 2 ; where U_epistemic is epistemic uncertainty, which quantifies the degree of uncertainty caused by insufficient knowledge or data; M is the number of sampling times of the Bayesian neural network, that is, the total number of model instances obtained; E[D|F_aligned, Θ_m] represents the expected value of decision D given the time alignment feature F_aligned and the mth model parameter instance Θ_m; E[D|F_aligned] represents the overall expectation of decision D given the time alignment feature F_aligned, usually the average of all model instance predictions; ∑ represents the summation symbol for all M model instances, (E[D|F_aligned, Θ_m] E[D|F_aligned]) 2 It represents the square of the difference between the expected value predicted by each model instance and the overall expectation, and measures the dispersion of the model prediction.
[0064] Estimate the random uncertainty U_aleatoric caused by the randomness of the environment: U_aleatoric = (1 / M)∑Var[D|F_aligned, Θ_m]; where U_aleatoric represents random uncertainty, quantifying the degree of uncertainty caused by the intrinsic randomness of the environment; Var[D|F_aligned, Θ_m] represents the variance of decision D given the time alignment feature F_aligned and the mth model parameter instance Θ_m; define the possible decision result set O_D and its corresponding cost function C(o). For each decision option d, calculate its risk score R(d): R(d) = ∑P(o|d, F_aligned)C(o); where P(o|d, F_aligned) represents the probability of the result o given the decision d and the current feature state. According to the safety level S_level of the task and the current system state Sys_state, calculate the epistemic uncertainty threshold τ_epistemic and the total uncertainty threshold τ_total. Safety-critical tasks use lower uncertainty thresholds to ensure the reliability of decisions: τ_epistemic = base_τ_e(1 S_level)f(Sys_state); τ_total = base_τ_t(1 S_level)f(Sys_state); where base_τ_e is the epistemic uncertainty baseline threshold, base_τ_t is the total uncertainty baseline threshold, and f(Sys_state) is the system state adjustment function, which takes into account factors such as battery level and processor load. Based on uncertainty and risk assessment, generate the final decision: D_final = {μ_D, if U_total < τ_total; RequestMoreInfo(), if U_epistemic > τ_epistemic; SafeDefaultDecision(), if R(μ_D) > τ_risk; μ_D withcaution flag, otherwise}; Output the final decision D_final and its related uncertainty indicator set U_metrics (including U_epistemic, U_aleatoric, R(μ_D), etc.). The generation rules of the final decision process are: when the total uncertainty is less than the total uncertainty threshold, the mean of the decision distribution is used; when the epistemic uncertainty is greater than the epistemic uncertainty threshold, more information is requested; when the risk score is greater than the preset risk threshold, a safe default decision is used; in other cases, the mean of the decision distribution with warning signs is used. The total uncertainty includes epistemic uncertainty and random uncertainty.
[0065] This embodiment defines a decision space to comprehensively cover all possible decision options and ensure the scientific nature of the decision. Using a Bayesian neural network and initializing the weight distribution in combination with a priori rule set, the model has domain knowledge in the initial stage, improving decision efficiency and accuracy. The Bayesian neural network weights are sampled multiple times to generate multiple model instances, which can effectively capture different potential decision distributions and enhance the robustness of the system. By quantifying epistemic uncertainty (lack of knowledge or data) and stochastic uncertainty (environmental randomness), a clear decomposition of the sources of uncertainty is provided to make decisions more transparent and controllable. According to the uncertainty index, a conservative decision plan can be selected to ensure system safety when the risk is high or the uncertainty is large. Combined with the cost function, the risk score of each decision option is quantitatively evaluated to support the selection of the optimal decision based on the importance of the task and the system state. In high-risk scenarios, it is possible to request more information or take conservative decisions, thereby reducing the potential negative consequences of the system. This embodiment is particularly suitable for task scenarios that require high reliability and safety.
[0066] According to one aspect of the present application, step S4 is further:
[0067] S41. Receive the final decision, uncertainty indicator set and task description, and extract the final decision type; establish a decision type mapping table to define the basic criticality of different decision types; adjust the criticality based on uncertainty indicators and risk scores; for a given decision type and uncertainty level, estimate the required computing resources, and adjust the resource requirements according to time constraints; extract key attributes from the task description, including task type, deadline and priority, and decompose the composite task to obtain a subtask set; evaluate the relevance of the task to safety and generate a safety relevance score; calculate the time urgency of the task to reflect the time pressure of task execution; comprehensively consider the basic criticality of the decision type, uncertainty factors and risk factors, and calculate the task criticality score; monitor the computing resource status, storage resource status and energy status of the smart car, calculate resource availability, and calculate resource allocation priority based on the task criticality score and resource requirements.
[0068] S42. Select model complexity based on task criticality score and uncertainty indicator set, and give priority to more complex models for high uncertainty decisions; establish computing requirement mapping table for different model complexity and precision levels, and select appropriate model complexity level based on task criticality score and resource availability; evaluate the potential contribution of each modality to reducing uncertainty based on the source of decision uncertainty, combine the uncertainty contribution with task relevance, evaluate the importance of each modality, and generate a modal importance vector, while considering the modal characteristics and task relevance and data quality; calculate the activation level of each modality based on the modal importance vector and model complexity level, and the value in the range of [0, 1] represents the degree of modal activation; allocate computing resources according to the activation level, and use low-precision calculation or skip processing for modalities with lower activation levels; for multi-task scenarios, build a task priority queue based on task criticality score and dependency relationship, and apply resource-constrained task scheduling algorithm to generate task execution plan.
[0069] According to one aspect of the present application, the step of calculating the task execution plan according to the task criticality score and the pre-stored system resource status includes:
[0070] Extract key attributes from pre-stored task descriptions, including task type, deadline, and priority; evaluate the safety association score and time urgency of the task based on the key attributes, and comprehensively calculate the task criticality score; monitor the computing resources, storage resources, and energy status of the smart car, and calculate resource availability; select the model complexity level based on the task criticality score and resource availability; calculate the modal importance vector of each mode to the current task, and calculate the activation level of the mode in combination with the model complexity level; allocate computing resources according to the activation level, reduce the processing accuracy or skip the processing of the mode with an activation level below the preset threshold, and generate a resource allocation plan; based on the resource allocation plan, build a task priority queue for multi-task scenarios and generate a resource-constrained task execution plan.
[0071] In one embodiment of the present application, the final decision D_final and the uncertainty indicator set U_metrics are received, and this information is combined with the task description T_desc to evaluate the criticality and resource requirements of the current decision. Key attributes are extracted from the task description T_desc, including the task type T_type, the deadline T_deadline, and the priority T_priority. For composite tasks, the task is decomposed to generate a subtask set T_sub and its dependencies. The relevance of the task to safety is evaluated, and a safety association score S_score is generated: S_score = AssessSafetyImpact(T_type, Env_state), where the AssessSafetyImpact function evaluates the safety impact based on the task type and the environmental state. The time urgency T_urgency of the task is calculated: T_urgency = 1 (T_deadline CurrentTime()) / MaxTimeWindow, where MaxTimeWindow is the maximum allowed time window for the task, CurrentTime() is the current time point of the system when the task is executed, and Env_state is the current state of the task environment. Taking the above factors into consideration, the criticality score C_task is calculated as follows: C_task = w_s S_score + w_p T_priority + w_u T_urgency, where w_s, w_p, and w_u are weight coefficients that can be dynamically optimized through reinforcement learning. Monitor the computing resource status R_compute, storage resource status R_memory, and energy status R_energy of the smart car. Computing resource availability R_availability: R_availability=min(R_compute / R_compute_max, R_memory / R_memory_max, R_energy / R_energy_max).
[0072] For different model complexity and accuracy levels, a computing requirement mapping table M_compute is established. Based on the task criticality score C_task and resource availability R_availability, an appropriate model complexity level L_complexity is selected: L_complexity = SelectComplexityLevel(C_task, R_availability), where SelectComplexityLevel is a decision function for selecting the model complexity level. For the current task, the importance of each modality is evaluated and a modality importance vector I_modality is generated: I_modality_i = Similarity(F_i, T_emb)Q_norm_i, where the Similarity function calculates the similarity between the modal feature and the task embedding, and Q_norm_i is the modality quality score. Based on the modality importance vector I_modality and the model complexity level L_complexity, the activation level A_level_i of each modality is calculated: A_level_i = SoftActivationFunction(I_modality_i, L_complexity), where SoftActivationFunction is a function for adjusting the modality activation level, and SoftActivationFunction returns a value in the interval [0, 1], indicating the modality activation level. Computing resources are allocated according to the activation level A_level_i: R_allocated_i = R_total A_level_i / ∑A_level_j, where R_total is the total computing resources that can be allocated. For modalities with activation levels lower than the threshold, low-precision calculations are applied or processing is skipped: ProcessingMode_i = {HighPrecision, if A_level_i > 0.8; MediumPrecision, if 0.4< A_level_i ≤ 0.8; LowPrecision, if 0.1 < A_level_i ≤ 0.4; Skip, if A_level_i≤ 0.1}. The activation level of the modal is divided as follows: when the activation level is greater than 0.8, the high-precision processing mode is used; when the activation level is between 0.4 and 0.8, the medium-precision processing mode is used; when the activation level is between 0.1 and 0.4, the low-precision processing mode is used; when the activation level is less than or equal to 0.1, the modal processing is skipped.
[0073] For multi-task scenarios, a task priority queue Q_task is constructed based on the task criticality score C_task and dependencies. The resource-constrained task scheduling algorithm is applied to generate a task execution plan P_execution: P_execution = ScheduleOptimization(Q_task, R_availability, TaskDependencies), where ScheduleOptimization is the task scheduling optimization function and TaskDependencies is the dependency between tasks. The output resource allocation scheme R_allocation and the task execution plan P_execution are used for system execution.
[0074] In another embodiment of the present application, the final decision D_final, the uncertainty indicator set U_metrics and the task description T_desc are received, and the final decision type D_type (such as obstacle avoidance, path planning, speed control, etc.) is extracted; the epistemic uncertainty U_epistemic, the random uncertainty U_aleatoric and the overall risk score R_total are extracted from the uncertainty indicator set U_metrics. A decision type mapping table M_decision is established to define the basic criticality of different decision types: BaseImportance(D_type) = M_decision[D_type]. For example, the basic criticality of the obstacle avoidance decision is higher than that of the speed optimization decision. Adjust the criticality based on the uncertainty indicator: UncertaintyFactor = 1 + w_ep * U_epistemic + w_al * U_aleatoric, where w_ep and w_al are weight coefficients, and high uncertainty usually means that a higher processing criticality is required. Further adjustment based on the risk score: RiskFactor = 1 + w_risk * R_total, the higher the risk, the more critical the decision. For a given decision type and uncertainty level, estimate the required computing resources: R_required = BaseResource(D_type) * UncertaintyFactor, where BaseResource is the basic resource requirement for each decision type. Adjust the resource requirements according to the decision time constraint: R_adjusted = R_required *(1 + w_time * (1 - AvailableTime / RequiredTime)); where w_time is the weight coefficient for adjusting the resource requirements based on time urgency; AvailableTime is the remaining time from the current moment to the task deadline; RequiredTime is the estimated time required to complete the task. The higher the time urgency, the higher the resource requirements.
[0075] Taking the above factors into consideration, calculate the revised task criticality score C_task: C_task = BaseImportance(D_type) * UncertaintyFactor * RiskFactor, where BaseImportance represents the basic criticality, UncertaintyFactor is the uncertainty factor, and RiskFactor is the risk factor. Normalize the task criticality to the interval [0, 1]: C_task_norm = C_task / max(C_task_history ∪ {C_task}), where C_task_history is the historical task criticality record. Monitor the system resource status, including computing resources R_compute, memory resources R_memory, and energy level R_energy. Calculate the resource allocation priority based on the task criticality score C_task_norm and resource demand R_adjusted: P_resource = C_task_norm * R_adjusted / R_availability, where R_availability is the resource availability.
[0076] Select model complexity based on task criticality score C_task_norm and uncertainty metric set U_metrics: L_complexity = SelectComplexityLevel(C_task_norm, U_metrics, R_availability). For high uncertainty decisions, prefer more complex models: if U_epistemic > τ_high: L_complexity += 1; Increase model complexity. Based on the sources of decision uncertainty, evaluate the potential contribution of each modality to reducing uncertainty: I_uncertainty_i = ContributionToUncertaintyReduction(F_i, U_metrics), where ContributionToUncertaintyReduction is a function used to evaluate the potential contribution of each modality in reducing decision uncertainty. Combine uncertainty contributions with task relevance and recalculate the modality importance vector I_modality: I_modality_i = w_sim * Similarity(F_i, T_emb) + w_unc * I_uncertainty_i, where w_sim and w_unc are weight coefficients.
[0077] This embodiment comprehensively evaluates the criticality score of the task in combination with the basic criticality, uncertainty factor and risk factor of the decision type. By dynamically adjusting the criticality, high-risk, high-uncertainty and high-priority tasks can be prioritized to improve the scientificity and rationality of the decision. The computing resources, storage resources and energy status of the smart car are monitored, and the resource requirements are adjusted in combination with the task criticality score and time urgency. By constructing a computing demand mapping table and a task priority queue, limited resources are reasonably allocated, which not only ensures the smooth completion of high-priority tasks, but also reduces the waste of system resources. Based on uncertainty contribution and task relevance, the importance of different modes is evaluated and the modal importance vector is generated. By calculating the modal activation level, the computing resource allocation is dynamically adjusted, the key modal data is efficiently utilized, and the resource consumption of low-contribution modes is reduced, thereby enhancing the adaptability and stability of the system. In high-uncertainty scenarios, complex models are preferred to improve decision accuracy, while simplified models are selected in resource-constrained or low-priority tasks. By accurately controlling the model complexity, not only the computing performance is optimized, but also the accuracy and efficiency are effectively balanced. A task scheduling algorithm based on task criticality score and dependency is adopted to generate the optimal execution plan under resource constraints. It not only ensures the logical dependencies and time constraints between tasks, but also improves the overall efficiency of multi-tasking. This embodiment achieves highly intelligent task decision-making and resource scheduling by comprehensively evaluating task priority, uncertainty and resource status. It is suitable for complex system scenarios with multiple tasks and limited resources, especially in tasks that require high reliability and dynamic adaptability, and can improve the decision-making efficiency, security and stability of the system.
[0078] According to one aspect of the present application, step S5 is further:
[0079] S51. Receive the resource allocation plan and task execution plan, extract the computing resource quota and storage resource quota allocated to the learning task, and determine the time window of the learning task; adjust the feature extraction depth according to the available resources; extract the task representation vector from the task description and the environment state vector to characterize the characteristics of the current task; based on the storage resource limit, adjust the retrieval scope, use the approximate nearest neighbor search to retrieve the historical tasks similar to the current task in the task experience library, and obtain relevant experience; based on the computing resource quota and the time window, determine the number of meta-learning iterations, execute the resource-aware model-independent meta-learning (MAML) algorithm, initialize the model parameters of the current task based on the experience data of similar tasks, and accelerate the adaptation of new tasks; based on the modality activation level, select the modality for knowledge distillation, perform lightweight knowledge distillation on the selected modality, extract knowledge from the more complex teacher model, and fine-tune the model parameters obtained from the meta-learning initialization to better adapt to the specific needs of the current task; monitor the resource usage of the learning process and dynamically adjust the resource allocation of the remaining learning steps.
[0080] S52. According to the available computing resources, adjust the data analysis granularity, analyze the feedback data coverage of the current task, evaluate the coverage of the decision space by the existing data, and identify data gap areas; train conditional generative adversarial networks based on existing data, and learn data distribution through adversarial training of the generator and the discriminator; according to the priority in the task execution plan, use the trained conditional generative adversarial network to generate synthetic feedback data for data gap areas and fill in the sparse areas of feedback data; evaluate the quality of the generated synthetic data, calculate the consistency of the synthetic data, and check the consistency of the synthetic data with the prior rules; integrate the highly consistent synthetic data with the real feedback data to form an enhanced feedback data set to provide more comprehensive training data for the learning algorithm; add resource monitoring and dynamic adjustment mechanisms to ensure execution within resource constraints.
[0081] S53. According to the priority in the task execution plan, select a subset of feedback data to be processed, and increase the feedback processing depth for high-priority tasks; classify the feedback data into positive feedback and negative feedback, and divide them based on the reward value of the feedback; evaluate the reliability score of each feedback data, taking into account the credibility, internal consistency and historical accuracy of the feedback source; construct a comparative learning batch, including anchor samples, positive samples and negative samples, and calculate the feature contrast loss; apply the reliability-weighted policy gradient method to update the decision model, and feedback with high reliability has a greater weight; evaluate the updated policy on the verification scenario set, calculate the verification performance index, and decide whether to accept the policy update; store the execution experience of the current task in the task experience library, and update the modal quality history record; periodically analyze the task experience library, use the frequent pattern mining algorithm to extract common experience patterns, and integrate them into the domain knowledge base for prior knowledge of future tasks; based on the results of feedback learning, update the parameter distribution of the Bayesian neural network, and pass the updated model parameters to step S3 for the next round of decision generation, forming a closed loop of decision-making-optimization-learning.
[0082] According to one aspect of the present application, the step of executing the comparative feedback learning optimization decision model includes:
[0083] Extract the task representation vector from the task description and the environment state vector; retrieve similar tasks in the task experience library based on the task representation vector and the task execution plan, and initialize the decision model parameters of the current task through the meta-learning algorithm; analyze the feedback data coverage of the current task based on the initialized decision model parameters and identify data gap areas; construct a conditional generative adversarial network to generate synthetic feedback data for data gap areas; classify the synthetic feedback data into positive feedback and negative feedback based on the priority in the task execution plan, and evaluate the feedback reliability score; construct a contrastive learning batch based on the feedback reliability score, update the decision model through contrastive loss and policy gradient method, and generate execution experience; store the execution experience in the task experience library, and extract the experience pattern to update the domain knowledge base.
[0084] In one embodiment of the present application, a resource allocation scheme R_allocation and a task execution plan P_execution are received, the scope and depth of meta-learning and knowledge transfer are adjusted based on resource constraints, and a task representation vector T_repr is extracted from the task description T_desc and the environment state vector E_state: T_repr = TaskEncoder(T_desc, E_state), where TaskEncoder is a task encoder network. Historical tasks similar to the current task are retrieved from the task experience database DB_task: T_similar = RetrieveTopK(DB_task, T_repr, K_task), where K_task is the number of similar tasks retrieved. Using a model-agnostic meta-learning algorithm (such as MAML), the model parameters of the current task are initialized based on the experience data D_similar of similar tasks: Θ_init = MAML(Θ_base, D_similar), where Θ_base is the basic model parameter, and MAML performs a meta-learning optimization process to obtain Θ_init. Extract knowledge from a more complex teacher model and apply it to the student model of the current task: L_distill = KL(P_student(D|F_aligned), P_teacher(D|F_aligned)), where KL is the KL divergence, P_student and P_teacher are the output probabilities of the student model and the teacher model, respectively. Fine-tune the model parameters initialized from meta-learning to adapt to the current task: Θ_adapted = Θ_init α▽ΘL(D_available, Θ_init), where α is the learning rate, L is the task loss function, and D_available is the currently available task data.
[0085] Analyze the feedback data coverage C_feedback of the current task and identify data gaps: C_feedback = Coverage(D_available, decision_space), where the Coverage function evaluates the coverage of the decision space by the existing data. Identify the data gap area G_data, that is, the part of the decision space that is not covered by the feedback data. Train the conditional generative adversarial network CGAN based on the existing data: min_G max_D E[log(D(F, D))] + E[log(1-D(F, G(F, z)))], where G is the generator, D is the discriminator, F is the feature input, and z is random noise. The generated synthetic data fills the sparse area of the feedback data. Use the trained conditional generative adversarial network CGAN to generate synthetic feedback data for the data gap area G_data: D_synthetic = G(F_template, z), where F_template is the template feature, representing the target scene condition. Evaluate the quality of the generated synthetic data and calculate the consistency of the synthetic data C_synthetic: C_synthetic =ConsistencyCheck(D_synthetic, R_prior), where the ConsistencyCheck function checks the consistency of the synthetic data with the prior rules. Keep the synthetic data with high consistency to form the verified synthetic data set D_synthetic_valid. Integrate the verified synthetic data set D_synthetic_valid with the real feedback data D_real to form the enhanced feedback data set D_enhanced: D_enhanced=D_real∪(w_synthetic D_synthetic_valid), where w_synthetic is the synthetic data weight coefficient, reflecting the trust in the synthetic data.
[0086] Feedback data is classified into positive feedback D_positive and negative feedback D_negative: D_positive = {d ∈ D_enhanced | Reward(d) > τ_reward}; D_negative = {d ∈ D_enhanced | Reward(d) ≤ τ_reward}, where the Reward function evaluates the reward value of the decision feedback and τ_reward is the reward threshold. Evaluate the reliability score R_feedback of each feedback data: R_feedback = w_source S_source + w_consist C_internal + w_hist A_historical, where S_source is the credibility of the feedback source, C_internal is the internal consistency measure, and A_historical is the historical accuracy. Normalize the reliability score R_feedback to the interval [0, 1] to obtain the normalized reliability score R_norm. Construct a contrastive learning batch B_contrast, including anchor samples, positive samples and negative samples: B_contrast = {(a, p, n) | a ∈ D_enhanced, p ∈ D_positive, n ∈ D_negative}; Calculate the feature contrast loss L_contrast: L_contrast = -log(exp(sim(F_a, F_p) / τ) / (exp(sim(F_a, F_p) / τ) + ∑exp(sim(F_a, F_n) / τ))), where sim is the similarity function, τ is the temperature parameter, and F_a, F_p, and F_n are the feature representations of anchor points, positive samples, and negative samples, respectively. Apply the reliability-weighted policy gradient method to update the decision model: ▽ΘJ(Θ) = E[R_norm▽Θlog(πΘ(D|F_aligned))], where πΘ is the policy function with parameter Θ, which represents the probability of taking a specific decision under a given feature state. Perform gradient update to obtain the updated policy parameters Θ_updated: Θ_updated = Θ_current + η▽ΘJ(Θ), where η is the learning rate, which can be adjusted dynamically according to the complexity of the task. Evaluate the updated policy on the validation scenario set S_validation and calculate the validation performance index P_validation: P_validation = EvaluatePolicy(πΘ_updated, S_validation), where EvaluatePolicy represents a function used to evaluate the performance of the updated policy on the validation scenario set.If the validation performance indicator P_validation is better than the previous policy, accept the update; otherwise, roll back to the previous policy or reduce the learning rate and try again.
[0087] Store the execution experience of the current task in the task experience library DB_task: DB_task = DB_task ∪ {(T_repr, F_aligned, D_final, R_outcome, U_metrics)}, where R_outcome is the task execution result, indicating the actual effect of the decision. Update the modal quality history record H_modality to record the performance quality of each modality under different environmental conditions: H_modality = UpdateHistory(H_modality, E_state, M_quality), where UpdateHistory is the update history function. Periodically analyze the task experience library DB_task and extract common experience patterns: K_patterns = ExtractPatterns(DB_task, min_support, min_confidence), where the ExtractPatterns function uses a frequent pattern mining algorithm to identify recurring decision patterns. Integrate the extracted experience patterns K_patterns into the domain knowledge base KB_domain for prior knowledge of future tasks: KB_domain = UpdateKnowledgeBase(KB_domain, K_patterns), UpdateKnowledgeBase is a function for updating prior knowledge. Based on the results of feedback learning, update the parameter distribution of the Bayesian neural network BNN: P_updated(Θ) = UpdatePosterior(P(Θ), D_enhanced), where the UpdatePosterior function is used to update the parameter distribution of the Bayesian neural network (BNN) based on the results of feedback learning. Pass the updated model parameters to step S3 for the next round of decision generation: BNN_next=UpdateModel(BNN_current, Θ_updated), where BNN_next is the updated Bayesian neural network, the UpdateModel function is used to apply new policy parameters to the existing Bayesian neural network, BNN_current is the currently used Bayesian neural network, and Θ_updated is the updated policy parameters.
[0088] In another embodiment of the present application, a resource allocation scheme R_allocation and a task execution plan P_execution are received, the computing resource quota R_learning and the storage resource quota R_storage allocated to the learning task are extracted, and the time window T_window of the learning task is determined based on the task execution plan P_execution. The feature extraction depth is adjusted according to the available resources: extraction_depth = AdjustDepth(R_learning), where AdjustDepth is the extraction depth function. The task representation vector T_repr is extracted using a resource-aware encoder: T_repr = TaskEncoder(T_desc, E_state, extraction_depth). If resources are extremely limited, lightweight feature extraction is used: if R_learning < min_threshold: T_repr = LightweightEncoder(T_desc), where min_threshold is the minimum threshold for resource usage, and LightweightEncoder is a lightweight feature extraction function. Based on storage resource limitations, the search scope is adjusted: search_scope = AdjustScope(R_storage, DB_task.size), where AdjustScope is a function for adjusting the search scope. Use approximate nearest neighbor search to improve retrieval efficiency: T_similar =ApproxNearestNeighbor(DB_task[:search_scope], T_repr, K_task), where K_task is the number of retrieval tasks that can be dynamically adjusted according to resources.
[0089] Based on the computing resource quota R_learning and the time window T_window, determine the number of meta-learning iterations: max_iterations = ComputeMaxIterations(R_learning, T_window), where the ComputeMaxIterations function is used to dynamically calculate the maximum number of iterations of the meta-learning algorithm based on the allocated computing resource quota and the time window of the learning task. Execute the resource-aware MAML algorithm: Θ_init = ResourceConstrainedMAML(Θ_base, D_similar, max_iterations), where ResourceConstrainedMAML is the resource-aware MAML algorithm. If resources are insufficient, degenerate to simple transfer learning: if max_iterations < min_MAML_iterations: Θ_init = SimpleFinetuning(Θ_base, D_similar.mean()), where SimpleFinetuning is the transfer learning function, min_MAML_iterations is the minimum iteration threshold of the meta-learning algorithm MAML, and D_similar.mean() is the average feature vector of similar task datasets. Based on the modality activation level A_level_i, select the modality for knowledge distillation: modalities_to_distill = {i | A_level_i > τ_distill}; Perform lightweight knowledge distillation on the selected modality: for iin modalities_to_distill: L_distill_i = KL(P_student_i(D|F_i), P_teacher_i(D|F_i)).
[0090] Weighted distillation loss according to modality importance: L_distill_total = ∑(I_modality_i * L_distill_i) for i in modalities_to_distill. Monitor resource usage during learning: R_used = MonitorResourceUsage(), MonitorResourceUsage() represents the monitoring resource usage function, dynamically adjust the resource allocation of the remaining learning steps: if R_used > R_learning * progress_ratio: reduce the resource usage of subsequent steps, AdjustRemainingSteps(remaining_steps, R_learning - R_used), where progress_ratio is the progress ratio of the learning process, AdjustRemainingSteps is the dynamic adjustment resource allocation function, and remaining_steps is the number of steps remaining in the current learning task or computing task. Adjust data analysis granularity according to available computing resources: analysis_granularity = AdjustGranularity(R_learning), where AdjustGranularity is the dynamic adjustment data analysis granularity function; use the granularity adjusted method to calculate the feedback data coverage C_feedback. The basic logic of the remaining steps remains unchanged, but they all require the addition of resource monitoring and dynamic adjustment mechanisms to ensure execution within resource constraints.
[0091] According to the priority in the task execution plan P_execution, select the subset of feedback data to be processed: D_priority = SelectFeedbackByPriority(D_enhanced, P_execution), where SelectFeedbackByPriority is the function for selecting a subset of feedback data. For high-priority tasks, the feedback processing depth is increased, and the remaining steps also need to add resource monitoring and dynamic adjustment functions to ensure that the feedback learning process can be completed efficiently within resource constraints.
[0092] This embodiment can efficiently utilize available resources according to the allocation of computing and storage resources by dynamically adjusting the feature extraction depth and the number of meta-learning iterations. The historical task experience retrieval based on approximate nearest neighbor search speeds up the model initialization speed, provides better initial parameters for new tasks, and accelerates the adaptation process of new tasks. The model parameters are fine-tuned using a lightweight knowledge distillation method to effectively extract the knowledge of the teacher model while reducing the computational overhead. Real-time monitoring of resource usage and dynamic adjustment of subsequent steps ensure that learning tasks are completed within resource constraints and reduce computational waste. Dynamic adjustment of analysis granularity based on computing resources can analyze feedback data at the best resolution, evaluate data coverage, and identify data gap areas. Conditional generative adversarial network (GAN) generates high-quality synthetic feedback data to fill data sparse areas and improve the comprehensiveness and diversity of training data. Through synthetic data quality assessment and consistency check, the high consistency between generated data and real data is ensured, which helps to optimize the learning algorithm. Add resource monitoring and dynamic adjustment mechanisms to avoid resource overload while ensuring the continuity of task execution plans. Increase the feedback processing depth for high-priority tasks to ensure the learning quality of key tasks while ensuring reasonable resource allocation for low-priority tasks. The reliability score is calculated by comprehensively considering the credibility and consistency of the feedback source, thereby improving the accuracy of the model update. Combining contrastive learning and reliability-weighted policy gradient methods, the model's adaptability to feedback data is further improved. Periodically analyze the task experience library, extract common experience patterns and integrate them into the domain knowledge base to promote cross-task knowledge sharing. Form a closed loop of decision-making-optimization-learning, and further improve the model performance through parameter updates of the Bayesian neural network. This embodiment ensures efficient learning under resource constraints through resource monitoring and dynamic adjustment. Combining meta-learning, data generation and adversarial training, knowledge distillation, contrastive learning and other technologies, the task execution efficiency, model adaptability and data utilization efficiency are comprehensively improved. Ultimately, the task completion time is shortened; the efficiency of computing resources and storage resources is improved; the generalization ability of the learning algorithm is enhanced; and the utilization rate of feedback data is improved.
[0093] In another embodiment of the present application, a system for collecting decision data based on an intelligent agent driven by a multimodal large model is composed of four modules, namely a decision-making intelligent agent based on a multimodal large model, a dynamic feature fusion module, a task adaptability adjustment module, and a multi-round feedback optimization module. For any input sensor data, the processing process of the model is mainly divided into two branches: decision generation and feedback optimization. Its purpose is to retain the decision details and information in the decision-making process to the maximum extent while ensuring the system response speed, so as to achieve efficient decision-making and execution of the smart car in a complex environment. The specific steps of the method for collecting decision data based on an intelligent agent driven by a multimodal large model are as follows:
[0094] Step 1: Introduce a multimodal large model architecture.
[0095] Step 1.1: Preprocessing of sensor data. Obtain raw data from the smart car's sensors (such as cameras, microphones, temperature sensors, lidar, etc.) and preprocess them, including data cleaning, denoising, normalization, etc. Assume there are N sensors, and the data of each sensor is represented by X; (where i = 1, 2, ..., N).
[0096] Step 1.2, feature extraction. The preprocessed sensor data is processed by their respective feature extraction modules. For example, image data is processed by convolutional neural network (CNN) to extract visual features, sound data is processed by sound feature extraction model (such as MFCC), and temperature data is processed by MLP. The feature vector of each sensor can be expressed as: F i =T i (X i ), where T i is the feature extraction function of sensor i, F i represents the feature vector of sensor i.
[0097] Step 1.3: Multimodal feature fusion. Fuse the features E of all sensors. Use the weighted concatenation method to combine the features of different modalities into a fused feature vector: F fusion =∑ i=1 N w i ·F i , where w i is the weight of the sensor i feature, which is dynamically adjusted by the adaptive mechanism during the training process.
[0098] Step 2: Multimodal alignment text module.
[0099] Step 2.1: Self-Attention Mechanism. The multimodal fusion feature F fusion Input to the self-attention module, which will feature F fusion It is processed as a learnable query. The self-attention mechanism captures the relationship between different modal features and generates enhanced polymorphic representations.
[0100] Step 2.2, Cross-Attention Mechanism. After the self-attention module, the task target text T is processed jointly with the feature input cross-attention module after the self-attention. The task target text T is used as an additional input to help the model focus on the multimodal features related to the current task, so as to generate a representation that meets the task requirements.
[0101] Step 2.3: Generate multimodal descriptions. After the cross-attention operation, the generated features are further processed by a feed-forward neural network (FFN), and finally a description of the multimodal data F is generated by linear projection. desc ,This description will serve as the input of the subsequent decision-making module, providing richer contextual information.
[0102] Step 3: Decision generation.
[0103] Step 3.1, input decision agent. Multimodal description F desc The input is sent to the decision-making agent (the reasoning module based on the large language model) for reasoning, and the agent generates a preliminary decision D based on the current state of the environment. raw The reasoning process of the decision-making agent can be expressed as: raw =M(F desc , T), where M is the inference function of the large language model and T is the context information of the current task.
[0104] Step 3.2: Decision verification and adjustment. raw Verify to ensure it meets the mission objectives. If not, proceed to the adjustment phase: D adjusted = D raw +△D, where △D is the adjustment item, which is dynamically updated based on feedback information and environmental changes.
[0105] Step 4: Task adaptability adjustment module.
[0106] Step 4.1: Environmental perception and weight adjustment. According to the latest sensor data collected by the environmental perception module, dynamically adjust the weight w of each modality. i , so that the data of important modalities can get more attention. For example, if image data is more important in the current task, the weight of image data can be increased image :w image =e α·Sim(Fimage,T) / ∑ i=1 N e α·Sim(Fi,T) , where Sim(F image , T) represents the similarity between the image features and the task target T, and α is the adjustment parameter.
[0107] Step 4.2: Dynamic adjustment process. According to the requirements of the current task, the smart car will automatically select the optimal combination of modal features to make decisions. When executing the task, the model dynamically selects task-related features for weighted fusion: F selected =w vis ·F vis + w sound ·F sound , where Fselected is the weighted fusion result, w vis is the weight of the visual modality (e.g., image data), F vis is the feature data of the visual modality, w sound is the weight of the acoustic mode (e.g., sound data), F sound is the characteristic data of the acoustic mode.
[0108] Step 5: Multi-round feedback optimization module.
[0109] Step 5.1: Feedback data collection. During the task execution process, the smart car collects execution data and feedback information in real time. These feedback data include task execution status, environmental changes, sensor data, etc.
[0110] Step 5.2: Feedback optimization and decision adjustment. Based on real-time feedback, the decision agent optimizes future decisions. The optimization process is as follows: optimized = D adjusted +△D feedback , where △D feedback It is a further adjustment of the decision based on feedback information.
[0111] Step 5.3, task strategy update. Through multiple rounds of interaction and feedback, the smart car continuously optimizes its decision-making strategy, so that the decision-making ability is improved in similar task scenarios. By analyzing the decision history, the model adjusts the weight of the decision strategy, thereby enhancing the stability and accuracy of the decision.
[0112] Furthermore, the step 1 comprises:
[0113] Step 11: Preprocessing of sensor data. Input various sensor data such as images, sounds, temperature, etc. into the multimodal large model architecture, and perform preliminary preprocessing on the data of each sensor to ensure that the input data format is consistent to facilitate subsequent feature extraction. In this step, raw data is collected from multiple sensors of the smart car, and each data is formatted and normalized to ensure data consistency and accuracy. For example, image data may need to be adjusted to the same resolution, sound data needs to be preprocessed to extract features (such as MFCC), and temperature data can be directly standardized. The goal of preprocessing is to ensure that the data of all sensors can be smoothly input and fused during subsequent processing.
[0114] Step 12: Multimodal feature fusion. In order to optimize the data fusion effect, a cross-modal feature weighted fusion method is used to automatically adjust the weights of different modalities under different task requirements to ensure that key information can be extracted first. The feature vectors of different sensors are merged by weighted splicing, and the features of each modality are weighted and adjusted according to the task requirements. Specifically, the weights will be dynamically adjusted according to the requirements of the current task to ensure that the modal data related to the task is given priority in the decision-making process. For example, if the current task relies more on image information, the system will assign a higher weight to the image data.
[0115] Furthermore, the step 2 comprises:
[0116] Step 21, decision agent input. The fused multimodal features are input into the decision agent (reasoning module based on the large language model) for reasoning, and the agent generates a preliminary decision based on the current environmental state and task objectives. The purpose of this process is to use the reasoning ability of the large language model to generate a decision output that meets the current task requirements based on multimodal data (such as vision, sound, temperature information) and the task target text T. The agent's reasoning ability can help generate specific behavioral guidance, such as path planning, action selection, etc.
[0117] Step 22: Decision verification and adjustment. The generated preliminary decision will be verified to ensure that it meets the task objectives. If the preliminary decision fails to meet expectations, the system will enter the adjustment phase to revise and optimize the decision. The goal of this step is to verify the generated decision through feedback information and task objectives to ensure the correctness of the decision. If the generated decision fails to achieve the expected effect, the system will make dynamic adjustments based on environmental changes or task feedback to ensure the smooth execution of subsequent tasks.
[0118] Furthermore, the step 3 comprises:
[0119] Step 31: Dynamic weight adjustment of task requirements. According to the different requirements of the task, the system will dynamically adjust the weight distribution of modal features to ensure that the key information related to the current task is prioritized in the multimodal information. For example, if the current task requires more reliance on visual data (such as target detection), the system will increase the weight of visual data to ensure that visual information occupies a more important position in the decision-making process. The purpose of this step is to dynamically weight different modalities so that the decision-making process can be more in line with task requirements and avoid interference from irrelevant data.
[0120] Step 32: Adaptive optimization based on environmental feedback. During the dynamic adjustment process, the system will use real-time environmental feedback information and user input (such as task priority, progress update, etc.) to optimize the task execution strategy. This adaptive adjustment mechanism enables the system to flexibly adapt to different environmental changes and task requirements, ensuring that the decision-making system has higher flexibility and adaptability in complex and dynamic environments. The system adjusts the task execution strategy in real time and optimizes the fusion mechanism of multimodal features to better cope with environmental changes and task challenges.
[0121] Further, the step 4 comprises:
[0122] Step 41: Feedback data collection. During the task execution, the smart car will continue to collect feedback data, including task execution status, environmental changes, sensor data, etc. This feedback information will provide a basis for subsequent decision optimization. By analyzing the feedback data during the task execution process, the system can detect potential problems in decision execution and provide effective information for the adjustment of subsequent decisions.
[0123] Step 42: Feedback analysis and decision adjustment. The collected decision data will undergo multiple rounds of analysis to identify potential problems or bottlenecks in task execution. By analyzing these feedbacks, the system can adjust the execution strategy of subsequent tasks, optimize the decision-making process, and ensure the efficiency and accuracy of subsequent task execution. Decision adjustments can include goal re-evaluation, path adjustment, or new task strategy generation.
[0124] Step 43: Optimize the decision strategy. Based on the feedback during the execution process, the system will further optimize the decision strategy. Through continuous learning and adjustment, the decision agent can continuously improve the accuracy of decision-making, thereby achieving more efficient and accurate execution in future tasks. This process enables the system to continuously enhance its ability to respond to environmental changes in multiple rounds of feedback and task execution, ensuring that the decision strategy can adapt and produce good results in different scenarios.
[0125] Further, the step 5 comprises:
[0126] Step 51: Multimodal alignment text module. In this step, the fused multimodal features F fusion It is input into a multimodal text alignment module, which is similar to QFormer, a module used to bridge the relationship between multimodal features and task objectives. fusionAs a learnable query, it is input into the self-attention module, which captures the relationship between the features of each modality and takes the task target text T as the input of the cross-attention operation. The purpose of this module is to ensure that the decision generation process can better consider the task goal through the precise alignment of multimodal data and task goals, thereby generating more accurate decisions.
[0127] Step 52, self-attention and cross-attention processing. In this module, the system deeply integrates multimodal features with the task target text through self-attention and cross-attention mechanisms. Through the self-attention mechanism, the system can capture the relationship between the features of each modality, while the cross-attention mechanism ensures that the task target text can effectively guide the processing of multimodal features. This process helps the system generate a multimodal description that is highly aligned with the task target, so that the decision-making agent can generate decisions that are more in line with the task requirements.
[0128] Step 53: Generate multimodal data description. After self-attention and cross-attention processing, the model will generate the final multimodal data description F through feedforward neural network (FFN) and linear projection. desc , which will be used as input for subsequent decision generation. Through this process, the system can better integrate multimodal information and generate expected decisions guided by task goals.
[0129] This embodiment proposes a decision data collection system based on a multimodal large model. By introducing multimodal data fusion, dynamic weight adjustment and multi-round feedback optimization mechanism, the accuracy of smart car decision-making and the efficiency of task execution are improved. The synergy of multimodal information in the decision-making process is improved, while ensuring the adaptability and execution ability in different task scenarios. Through the fusion mechanism of the multimodal large model, data from different sensors are effectively integrated to improve the perception and understanding of complex environments. The decision strategy is dynamically adjusted according to real-time environmental changes and task requirements to ensure the flexibility and efficiency of task execution. Through the interaction process between the smart car and the environment, the decision strategy is optimized in real time, thereby improving the accuracy of task execution and the responsiveness of the system. With the support of multimodal information, the decision-making agent can make more accurate and reasonable decisions, thereby optimizing the behavior and task completion of the smart car. The method implemented in this embodiment makes full use of multimodal inputs such as vision, sound, and temperature through dynamic fusion modules, cross-modal attention mechanisms, and knowledge feedback mechanisms to achieve an efficient decision-making process for the smart car. At the same time, through multiple rounds of interaction between the smart car and the environment, the decision-making strategy can be continuously optimized, and detailed decision-making data can be collected for subsequent learning and analysis, thereby further improving the execution capability and task completion accuracy of the smart car.
[0130] The present invention belongs to the field of intelligent transportation and automation, and in particular, relates to a method for collecting decision data of an intelligent car based on a multimodal large model intelligent agent. In the field of intelligent transportation systems, it can be applied to decision optimization and data collection of vehicles such as intelligent cars and automatic driving systems. Through the fusion of multimodal data (such as sensor data such as images, sounds, and temperature) and the reasoning ability of large language models, the decision accuracy and execution efficiency of intelligent cars in complex environments can be improved, ensuring that vehicles can make efficient and safe decisions under different traffic conditions, thereby improving traffic management and driving safety. In the field of automatic driving, it can be applied to the perception and decision-making module of automatic driving cars. Through the fusion of multi-sensor data and the reasoning ability of intelligent agents, the system can automatically generate decisions in complex and dynamic road environments, optimize path planning and behavior selection. This method can improve the adaptability of the automatic driving system under changing lighting, weather and different road conditions, and enhance the safety and reliability of the system. In the field of logistics and warehousing management, it can be applied to the path planning and task execution of automated warehouses and intelligent logistics vehicles. Through real-time processing and decision optimization of multimodal data, the system can achieve efficient cargo identification, path selection and task scheduling, improve the automation and accuracy of the logistics management system, reduce operating costs and improve transportation efficiency. In the field of intelligent monitoring and security systems, it can be applied to mobile monitoring systems based on small vehicles. Through data fusion and intelligent decision analysis of multimodal sensors, real-time data collection and event recognition can be performed in various environments, and the response speed and accuracy of the monitoring system in complex scenarios can be improved, thereby enhancing the effect and quality of security monitoring. In the field of smart agriculture, it can be applied to automated agricultural equipment, such as smart agricultural vehicles or drones. By collecting and analyzing data from multiple sensors in real time (such as ambient temperature and humidity, soil conditions, etc.) and making decisions through intelligent agents, agricultural equipment can be better adapted to different agricultural tasks (such as sowing, spraying, etc.), improve agricultural production efficiency, and promote the development of precision agriculture. In the field of environmental monitoring and emergency response, it can be applied to mobile environmental monitoring equipment. In the process of post-disaster rescue or environmental monitoring, through real-time data collection and intelligent decision analysis of multimodal sensors, the system can quickly respond and optimize decisions, improve the speed and accuracy of environmental monitoring and emergency response, and provide effective support for post-disaster rescue and environmental protection.
[0131] According to one aspect of the present application, a system for collecting decision data by an intelligent agent driven by a multimodal large model includes:
[0132] at least one processor; and,
[0133] a memory communicatively connected to at least one of the processors; wherein,
[0134] The memory stores instructions that can be executed by the processor, and the instructions are used to be executed by the processor to implement the method for collecting decision data by an intelligent agent driven by a multimodal large model as described in any of the above embodiments.
[0135] The present invention collects the original data of the predetermined sensor, calculates the information entropy and the normalized quality score, and generates a fusion feature vector through adaptive multimodal information fusion; performs time-sensitive multimodal alignment processing on it, and combines timestamp embedding with self-attention mechanism to obtain time-series alignment features; based on this, a Bayesian neural network is used to generate decision distribution, evaluate uncertainty and generate the final decision and indicator set; the task score is evaluated in combination with the task description, and an execution plan is formulated according to the resource status; the model is initialized, synthetic data is generated, the model is optimized and experience is stored. Combining large-scale language models and multimodal perception technology, a new framework for cross-modal decision data collection is proposed. Through this framework, the smart car can make real-time decisions in a constantly changing environment and adjust future behaviors according to historical decision data to better adapt to the needs of different task scenarios. Based on the decision type, uncertainty and risk score, the system can more accurately evaluate the mission criticality and resource requirements to ensure that important decisions are fully supported by calculation. Through meta-learning, knowledge transfer and feedback learning processes, its depth and scope can be dynamically adjusted according to available resources to ensure effective learning even in resource-constrained situations. By analyzing the potential contribution of each mode to uncertainty, the system can more intelligently select which modes should be activated to achieve the maximum decision benefit under resource constraints. The present invention completes the whole process from multimodal data acquisition, timing alignment, decision generation, task adaptation optimization to feedback learning. The system is specifically designed for practical challenges faced by smart cars, such as unstable data quality, timing differences, decision uncertainty, limited resources, and sparse feedback data, to ensure that the system can continuously and efficiently make decisions and learn in a complex and changing environment.
[0136] The preferred embodiments of the present invention are described in detail above; however, the present invention is not limited to the specific details in the above embodiments. Within the technical concept of the present invention, various equivalent transformations can be made to the technical solutions of the present invention, and these equivalent transformations all belong to the protection scope of the present invention.
Claims
1. A method for collecting decision data by an intelligent agent driven by a multimodal large model, characterized in that: The following steps are involved: Obtain the raw sensor data of the predetermined sensors of the smart car, and calculate the information entropy and normalized quality score; Generate a fusion feature vector through adaptive multi-modal information fusion; Apply multimodal alignment processing to the fused feature vector and generate temporal alignment features through timestamp embedding and temporal self-attention mechanism; Based on the time series alignment features, the decision distribution is generated through the Bayesian neural network; Calculate epistemic uncertainty and random uncertainty based on decision distribution and conduct risk assessment to generate the final decision and uncertainty indicator set; Evaluate the task criticality score in combination with the pre-stored task description; The task execution plan is calculated based on the task criticality score and the pre-stored system resource status; Read the task execution plan and initialize the decision model through meta-learning; identify data gap areas and generate synthetic feedback data; Combine the priority execution comparison feedback in the task execution plan to learn and optimize the decision model, and store the execution experience in the task experience library; The steps of generating decision distribution through Bayesian network and generating final decision and uncertainty indicator set include: Construct a Bayesian neural network as a decision model; Extract the prior rule set from the domain knowledge base and convert it into a prior probability distribution to initialize the Bayesian neural network; The weights of the Bayesian neural network are sampled a predetermined number of times to obtain a predetermined number of model instances; and the decision distribution is estimated in combination with the time series alignment feature; Calculate epistemic uncertainty caused by insufficient knowledge or data and stochastic uncertainty caused by environmental randomness based on decision distribution; Conduct risk assessment based on epistemic uncertainty and random uncertainty to obtain a risk score; Determine the dynamic decision threshold based on the pre-stored task safety level and system status, and generate the final decision and uncertainty indicator set according to the uncertainty type and risk score; the dynamic decision threshold includes the total uncertainty threshold and the cognitive uncertainty threshold; The steps of calculating the task execution plan according to the task criticality score and the pre-stored system resource status include: Extract key attributes from pre-stored task descriptions, including task type, deadline, and priority; Evaluate the safety relevance score and time urgency of the task based on key attributes, and calculate the task criticality score comprehensively; Monitor the computing resources, storage resources and energy status of the smart car, and the availability of computing resources; Select the model complexity level based on mission criticality score and resource availability; Calculate the modal importance vector of each modality for the current task, and calculate the activation level of the modality based on the model complexity level; Allocate computing resources according to the activation level, reduce processing accuracy or skip processing for modalities whose activation level is lower than a preset threshold, and generate a resource allocation plan; Based on the resource allocation scheme, a task priority queue is constructed for multi-task scenarios and a resource-constrained task execution plan is generated.
2. The method according to claim 1, characterized in that The steps of adaptively fusing multimodal information to generate a fused feature vector include: Construct an environmental state vector based on the raw sensor data, evaluate the sensor data quality, calculate the signal-to-noise ratio, data integrity, timing stability index, and spectrum consistency, and obtain a normalized quality score; Calculate the information entropy of raw sensor data; For raw sensor data whose normalized quality score is lower than a preset threshold or whose information entropy is higher than a preset threshold, similar scene data are retrieved from a historical database and compensation data is generated; Input the original sensor data and the compensated data into the corresponding feature extraction network to obtain the feature vectors of each mode; The fusion weight is calculated based on the normalized quality score and information entropy, and the feature vectors of each modality are weighted and fused to generate a fused feature vector.
3. The method according to claim 2, characterized in that The calculation formula of fusion weight is: W_i = Q_norm_i exp(-βH_i) / ∑(Q_norm_j exp(-βH_j)); Among them, Wi is the fusion weight of modality i, Qnorm is the normalized quality score, Hi is the information entropy, and β is the entropy weight coefficient.
4. The method according to claim 1, characterized in that: The steps of generating time-aligned features through timestamp embedding and time-series self-attention mechanism include: Calculate the time difference matrix based on the timestamps of the original sensor data and generate a time embedding vector; Add the time embedding vector to the fused feature vector to generate time series enhanced features; Aggregate the time series enhancement features in time windows to generate multi-scale time series features; Apply the time difference weighted self-attention mechanism to the multi-scale temporal features to generate temporal self-attention features; Based on the temporal self-attention features, a temporal graph structure is constructed and a graph attention network is applied to perform temporal relationship reasoning, and the temporal alignment features are generated by combining the pre-stored task description embedding.
5. The method according to claim 4, characterized in that The calculation formula of temporal self-attention feature is: W_t_ij = exp(-γ|ΔT_norm_ij|); A_ij = softmax((Q_iK_j T / sqrt(d_k))W_t_ij); Among them, W_t_ij is the time difference weight, ΔT_norm_ij is the normalized time difference, γ is the time sensitivity parameter, A_ij is the attention score, Q_i and K_j are the query and key matrices respectively, d_k is the key vector dimension, T Indicates transpose.
6. The method according to claim 1, characterized in that Epistemic uncertainty is: U_epistemic = (1 / M)∑(E[D|F_aligned,Θ_m] E[D|F_aligned]) 2 ; Where U_epistemic is epistemic uncertainty, M is the number of samplings, E[D|F_aligned,Θ_m] represents the expectation of model instance m for decision making, and E[D|F_aligned] represents the average expectation of all models; The random uncertainty is: U_aleatoric = (1 / M)∑Var[D|F_aligned, Θ_m]; Where U_aleatoric is the random uncertainty, and Var[D|F_aligned, Θ_m] represents the variance of the decision made by model instance m.
7. The method according to claim 1, characterized in that In the process of generating the final decision based on the type of uncertainty and the results of risk assessment, the generation rules are: When the total uncertainty is less than the total uncertainty threshold, the mean of the decision distribution is used; When epistemic uncertainty is greater than the epistemic uncertainty threshold, request more information; When the risk score is greater than the preset risk threshold, a safe default decision is adopted; In other cases, the mean of the decision distribution with warning signs is used.
8. A system for collecting decision data by intelligent agents driven by a multimodal large model, characterized in that: include: at least one processor; as well as, a memory communicatively connected to at least one of the processors; wherein, The memory stores instructions that can be executed by the processor, and the instructions are used to be executed by the processor to implement the method for collecting decision data by an intelligent agent driven by a multimodal large model as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Automatic operation and maintenance method and system for power distribution network based on artificial intelligence
CN119090490A
Graph neural network-based power grid dispatching decision-making method and large model
CN119294872A