A Multi-Agent Decision-Programming-Communication-Control System and Cooperative Method Based on Meta-Reinforcement Learning and Uncertainty-Driven Approach

CN122569550APending Publication Date: 2026-08-14CHONGQING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-27
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

传统运动规划通常基于单步或短视距的确定性预测,例如假设障碍物按恒定速度运动,无法量化预测结果的不可靠度

Benefits of technology

(1)通过深度融合协同感知信息与不确定性量化机制,系统能够突破传统短视距规划的局限,在隐空间内推演未来时空演化趋势并显式量化预测不可靠度。这种预测与风险的深度耦合,使系统能够提前识别潜在碰撞风险,在常规工况下执行兼顾效率与安全的最优规划,在极端未知环境下则自主降级为保守策略,从根本上避免了因盲目信任错误预测而导致的决策失误。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122569550A_ABST
    Figure CN122569550A_ABST
Patent Text Reader

Abstract

This invention relates to a multi-agent decision-planning-communication-control system and cooperative method based on meta-reinforcement learning and uncertainty-driven approaches, belonging to the field of multi-agent cooperative control technology. Addressing the technical problems of existing multi-agent cooperative systems, such as prediction uncertainty blind spots, wasted communication bandwidth, rigid security boundaries, and difficulties in cold start, this invention proposes a four-in-one cooperative scheme. The system includes a prediction-risk coupled decision-planning module for cooperative perception and uncertainty quantification, a reinforcement learning self-evolutionary communication and information compression cooperative module, an uncertainty-adjusted control barrier function safety filtering control module, and a meta-reinforcement learning cooperative adaptive optimization module. By explicitly quantifying prediction uncertainty and driving dynamic adjustment of communication and security boundaries, combined with meta-learning to achieve rapid generalization across scenarios, this invention improves the forward-looking decision-making safety, communication resource utilization, and underlying control flexibility of multi-agent systems in complex dynamic environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of multi-agent cooperative control technology, and relates to a multi-agent decision-planning-communication-control system and cooperative method based on meta-reinforcement learning and uncertainty-driven approach. Background Technology

[0002] In the practical deployment of multi-agent systems, the high dynamism of the environment, the limitations of sensor observation, and the constraints of communication bandwidth pose significant challenges to the safe and efficient collaboration of the system. Existing multi-agent cooperative control technologies have obvious shortcomings in several key aspects, which restricts their application effectiveness in complex dynamic environments.

[0003] First, there's the issue of uncertainty blind spots in the prediction and planning stages. Traditional motion planning typically relies on deterministic predictions based on single steps or short lines of sight, such as assuming obstacles move at constant speeds, making it impossible to quantify the unreliability of the predictions. When faced with sudden changes in neighboring intentions or extremely complex environments, the system often blindly trusts erroneous predicted trajectories, leading to a surge in collision risk.

[0004] Secondly, there is a contradiction between wasted bandwidth and lack of fidelity at the communication layer. Existing collaborative systems mostly use fixed-period communication mechanisms or static trigger thresholds, which not only wastes valuable communication bandwidth but also easily causes channel congestion under dangerous conditions. In addition, information compression often uses a fixed compression rate, which cannot dynamically adjust the fidelity of information according to the degree of danger of the current environment, resulting in the loss of critical information or redundant transmission.

[0005] Third, the safety control layer lacks dynamic adaptive boundaries. Traditional control barrier functions (CBFs) typically preset fixed safety distance boundaries. When there are delays in upper-level planning instructions or significant errors in predicting the trajectory of dynamic obstacles, fixed lower-level safety boundaries often fail to provide sufficient buffer space, leading to the breach of physical safety constraints.

[0006] Fourth, multi-module integrated systems face challenges in cold start and collaborative optimization. Decision-making, planning, communication, and control modules are typically designed and calibrated independently. When an agent is deployed to a completely new dynamic environment, traditional reinforcement learning algorithms require extremely long trial-and-error times to reconverge, and are highly susceptible to security incidents during this period—hence the difficulty of cold start. Simultaneously, the lack of a global optimization mechanism to coordinate the balance between planning efficiency, communication overhead, and control security makes it difficult for modules to work together to achieve optimal performance.

[0007] In summary, existing technologies still have significant shortcomings in areas such as environmental prediction and uncertainty quantification, adaptive communication compression, uncertainty-driven security constraints, and multi-module collaborative adaptive optimization. These limitations make it difficult to meet the comprehensive application requirements of multi-agent decision-making and planning in complex dynamic environments, which demand for foresight, security, communication efficiency, and rapid adaptability. Therefore, there is an urgent need for a full-link collaborative optimization scheme that integrates perception, decision-making, communication, and control to address these issues. Summary of the Invention

[0008] In view of this, the purpose of this invention is to provide a multi-agent decision-planning-communication-control system and cooperative method based on meta-reinforcement learning and uncertainty-driven approaches.

[0009] To achieve the above objectives, the present invention provides the following technical solution: A multi-agent decision-planning-communication-control system based on meta-reinforcement learning and uncertainty-driven, comprising a prediction-risk coupled decision-planning and planning module with collaborative perception and uncertainty quantification, a reinforcement learning self-evolutionary communication and information compression collaborative module, a control barrier function safety filtering control module with uncertainty adjustment, and a meta-reinforcement learning collaborative adaptive optimization module. The collaborative perception and uncertainty quantification prediction-risk coupled decision and planning module is used to generate future prediction hidden states, quantify prediction uncertainty and calculate comprehensive path risk, and output the optimal reference trajectory. The reinforcement learning self-evolutionary communication and information compression collaboration module is used to drive dynamic decision-making based on uncertainty and adaptively adjust the information compression rate to achieve low-bandwidth high-fidelity collaboration. The uncertainty-adjusted control barrier function safety filter control module is used to perform minimally invasive safety correction on the optimal reference trajectory and output safety control commands. The meta-reinforcement learning collaborative adaptive optimization module is used to construct a global multi-objective collaborative reward function, and achieves collaborative optimization of decision, communication and control parameters through offline meta-training and online fast adaptive updating.

[0010] Furthermore, the collaborative perception and uncertainty quantification prediction-risk coupled decision-making and planning module specifically includes: The collaborative sensing coding unit is used to fuse local observation data with the intent information of neighboring agents to generate a fused hidden state vector. ,in, For the i-th agent in t The fused hidden state vector at time step 1. For encoder network parameters, For implicit space dimension; Latent space multi-step dynamic prediction unit, used for fusion of latent state vectors With candidate spatiotemporal trajectory clusters Generate the future H The predicted hidden state sequence at each time step; Deeply integrated uncertainty quantization unit, used to... K Given two dynamically transitioning subnetworks with identical structures but initialized parameters, calculate the mean of the predicted states. With uncertainty covariance matrix ,in, , For the first k The predicted state of each subnetwork K The total number of sub-networks; The spatiotemporal evolution risk quantification unit is used to calculate the comprehensive path risk function of candidate spatiotemporal trajectories based on the predicted mean and the uncertainty covariance matrix. ,in, , , These are the weighting coefficients. The time decay factor, For the first k Predicting the timing agent i Collision object j The instantaneous collision probability; The decision and planning unit is used to output the optimal candidate trajectory under normal operating conditions. Output a preset safety trajectory under extremely uncertain operating conditions. The final output is the optimal reference trajectory. .

[0011] Furthermore, the reinforcement learning self-evolutionary communication and information compression collaborative module specifically includes: The communication-triggered decision unit is used to calculate the marginal communication value gain based on the communication strategy value network. And combined with uncertainty-driven dynamic trigger threshold Determine communication decision variables ,in, , , For communication strategy value network, This is the lower limit of the minimum communication threshold. Based on the trigger threshold, This is the uncertainty sensitivity coefficient; An adaptive information compression unit for uncertainty-driven adaptive compression fidelity. The information to be sent is compressed and encoded, where, , This is the lower limit of the minimum compression ratio. This is the sensitivity coefficient; The receiving information fusion unit is used to calculate confidence weights based on the uncertainty of neighbor information. Generate a fused state representation by integrating its own state with neighbor information. ,in, , The uncertainty attenuation coefficient is... To prevent small constants from being divided by zero, Gather for neighbors; Joint optimization unit, used for communication reward function End-to-end collaborative optimization of communication triggering strategy and information compression strategy is performed, wherein, , , , These are the weighting coefficients. For the cost of communication, This refers to the information reconstruction error.

[0012] Furthermore, the uncertainty-adjusted control barrier function security filtering control module specifically includes: Adaptive safety function building block for location-based prediction uncertainty scalar Build a dynamic adaptive safe distance ,in, , , This is a submatrix representing the corresponding positional components in the covariance matrix. Based on the safe distance, This is the adjustment coefficient for uncertainty regarding the safety distance boundary; The control barrier function constraint generation unit is used to substitute the adaptive safety function into the control barrier function constraint conditions to generate linear constraints on the control input, wherein... , The gain coefficient used to control the decay rate of the safety function; The quadratic programming safety filtering unit is used to output safety control commands by solving a quadratic programming problem. The objective function is: , U For the physically feasible set of control inputs for the actuator, For nominal reference control instructions.

[0013] Furthermore, the meta-reinforcement learning collaborative adaptive optimization module specifically includes: A global multi-objective collaborative reward function construction unit is used to construct a global multi-objective collaborative reward function that comprehensively considers planning efficiency, communication cost, and underlying security performance. ,in, , , , These are the weighting coefficients. Rewards for navigation and planning, Rewards for communication efficiency Rewards for safety interventions; Meta-learning task distribution building unit, used to encapsulate environmental tasks into task distributions. Define a global meta-initialization parameter set. ,in, , For spatiotemporal prediction and planning parameter set, The decision parameter set is used to trigger communication. For adaptive information compression and fusion parameter set, This is a set of underlying control security parameters; The offline meta-training unit is used for meta-training via a two-layer gradient update architecture, with the inner update based on the support set. Fine-tune task-specific parameters Outer meta update based on query set Optimize meta-initialization parameters ; Online collaborative fast adaptive unit, used in new deployment scenarios, based on a global multi-objective collaborative reward function. Global meta-initialization parameter set Perform online collaborative gradient updates.

[0014] A multi-agent decision-planning-communication-control collaborative method based on meta-reinforcement learning and uncertainty-driven approaches includes the following steps: S1: Through collaborative perception and uncertainty quantification, prediction-risk coupled decision-making and planning steps, generate future prediction latent states, quantify prediction uncertainty and calculate comprehensive path risk, and output the optimal reference trajectory. S2: By using reinforcement learning to self-evolutionize communication and information compression coordination steps, the communication triggering timing is driven by uncertainty, and the information compression rate is adaptively adjusted to achieve low-bandwidth high-fidelity coordination. S3: Through the safety filtering control step of the control barrier function adjusted for uncertainty, the optimal reference trajectory is modified with minimal intrusion, and a safety control command is output. S4: Through meta-reinforcement learning collaborative adaptive optimization steps, a global multi-objective collaborative reward function is constructed. Through offline meta-training and online fast adaptive updates, collaborative optimization of decision, communication, and control parameters is achieved.

[0015] Furthermore, S1 specifically includes: S11: Fuse local observation data with the intent information of neighboring agents to generate a fused hidden state vector. ,in, For the firsti An intelligent agent in t Local observation data at time, For the neighboring intelligent agent's intent information, For multimodal encoder neural networks, For encoder network parameters; S12: Based on fusion of hidden state vectors With candidate spatiotemporal trajectory clusters Through dynamic transfer network Generate the predicted hidden state for the next time step. And through the decoder neural network Restore to the physical prediction state ,in, To dynamically transfer network parameters, These are the decoder network parameters; S13: Calculate the mean of the predicted state using a deep ensemble method. With uncertainty covariance matrix ; S14: Calculate the comprehensive path risk function of candidate spatiotemporal trajectories based on the predicted mean and the uncertainty covariance matrix. ; S15: Output the optimal candidate trajectory under normal operating conditions Output a preset safety trajectory under extremely uncertain operating conditions. The final output is the optimal reference trajectory. ,in, This represents the feasible candidate trajectory space. For efficiency weighting coefficients, This is the efficiency evaluation function.

[0016] Furthermore, S2 specifically includes: S21: Value Network Based on Communication Strategy Calculate the value gain of marginal communication And combined with uncertainty-driven dynamic trigger threshold Determine communication decision variables ,when hour, ,otherwise ; S22: When At that time, adaptive compression fidelity based on uncertainty-driven The information to be sent is compressed and encoded to generate compressed information. ,in, For learnable compression encoder neural networks, For encoder network parameters, Original information to be sent; S23: The receiver uses a decoder. Recovery Information ,in, These are the decoder network parameters; S24: Calculating confidence weights based on the uncertainty of neighbor information Generate a fused state representation by integrating its own state with neighbor information. ,in, To converge networks, To integrate network parameters, The current perception state; S25: Based on the communication reward function End-to-end collaborative optimization of communication triggering strategy and information compression strategy.

[0017] Furthermore, S3 specifically includes: S31: Scalar based on location prediction uncertainty Build a dynamic adaptive safe distance ; S32: Based on dynamic adaptive safety distance Constructing an adaptive security function And substitute it into the control barrier function constraints. Generate linear constraints on the control input. ,in, and These are the Lie derivatives of the adaptive safety function along the system drift dynamics and the control input direction, respectively. The real-time rate of change of location prediction uncertainty; S33: Solving a quadratic programming problem Output safety control commands .

[0018] Furthermore, S4 specifically includes: S41: Construct a global multi-objective collaborative reward function ,in, , , , For security conflict penalties, To control the violation of the indicator function by the barrier function constraint; S42: Encapsulate environmental tasks as task distributions Define the global meta-initialization parameter set. ; S43: During the offline simulation training phase, meta-training is performed using a two-layer gradient update architecture, with the inner layer update based on the support set. Fine-tune task-specific parameters Outer meta update based on query set Optimize meta-initialization parameters ,in, For the inner layer adaptive learning rate, The learning rate of the outermost element; S44: In the new deployment scenario, load the global meta-initialization parameter set. Based on a global multi-objective collaborative reward function Online collaborative gradient updates enable rapid adaptation of decision-making, communication, and control parameters.

[0019] The beneficial effects of this invention are as follows: (1) By deeply integrating collaborative perception information and uncertainty quantification mechanisms, the system can overcome the limitations of traditional short-line planning, extrapolate future spatiotemporal evolution trends in the implicit space, and explicitly quantify the unreliability of predictions. This deep coupling of prediction and risk enables the system to identify potential collision risks in advance, execute optimal planning that balances efficiency and safety under normal operating conditions, and autonomously downgrade to a conservative strategy in extremely unknown environments, fundamentally avoiding decision-making errors caused by blindly trusting erroneous predictions.

[0020] (2) By introducing uncertainty as the core driving force, the system breaks the rigid pattern of fixed-period communication and static compression rate. The communication triggering mechanism can dynamically and adaptively adjust according to the environment, actively communicating when uncertainty is high and remaining silent when the environment is simple. At the same time, the information compression rate changes dynamically with the risk level, ensuring high-fidelity transmission of key game information under dangerous conditions and maximizing bandwidth savings under stable conditions. This self-evolving communication mechanism effectively solves the channel congestion problem in large-scale intelligent agent clusters.

[0021] (3) By explicitly introducing prediction uncertainty into the control barrier function, the system constructs a safety boundary that can dynamically expand or contract with the environment. Compared with traditional fixed boundary safety constraints, this invention can provide more ample buffer space when the environment is complex and prediction is unreliable, while allowing for more compact passage when the environment is certain. Combined with a minimally invasive safety filtering architecture, the system can retain the decision-making intent of the upper-level planning to the greatest extent while strictly adhering to the safety bottom line, thus alleviating the problem of excessive conservatism in traditional safety control methods.

[0022] (4) The collaborative optimization mechanism based on meta-reinforcement learning enables the system to learn quickly from cross-scenario general knowledge extracted offline. When deployed to a new environment, only a very small number of online interaction samples are needed to complete the collaborative evolution of decision-making, communication and control parameters. This four-in-one adaptive optimization greatly reduces the trial-and-error cost and deployment difficulty of the system in new scenarios, ensuring the feasibility of industrial implementation.

[0023] (5) This invention breaks down the design barriers between decision planning, communication and control systems, which are independent of each other, and tightly couples the four modules through a unified global multi-objective reward function. The functional boundaries of each module are clear and the interfaces are well-defined, so that they can evolve independently and operate collaboratively, forming a positive loop of perception guiding communication, communication assisting planning, planning controlled by safety filtering, and safety feedback optimizing perception, which comprehensively improves the overall collaborative efficiency of the multi-agent system in complex dynamic environments.

[0024] Other advantages, objectives, and features of the invention will be set forth in part in the description which follows, and in part will be apparent to those skilled in the art from the following examination, or may be learned from practice of the invention. The objectives and other advantages of the invention can be realized and obtained through the following description. Attached Figure Description

[0025] To make the objectives, technical solutions, and advantages of the present invention clearer, the preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings, wherein: Figure 1 The flowchart is for a prediction-risk coupled decision-making and planning algorithm based on collaborative perception and uncertainty quantification. Figure 2 Flowchart of a self-evolving communication and information compression collaborative mechanism based on reinforcement learning; Figure 3 Here is a flowchart of a safety filtering control method based on an uncertainty-adjusted control barrier function; Figure 4 Flowchart of a decision-planning-communication-control collaborative adaptive optimization mechanism based on meta-reinforcement learning; Figure 5 A flowchart of a multi-agent collaborative method integrating decision-making, planning, communication, control, and adaptation. Detailed Implementation

[0026] The following specific examples illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of the present invention. Unless otherwise specified, the following embodiments and features can be combined with each other.

[0027] The accompanying drawings are for illustrative purposes only and are schematic diagrams, not actual pictures. They should not be construed as limiting the invention. To better illustrate the embodiments of the invention, some parts in the drawings may be omitted, enlarged, or reduced, and do not represent the actual product dimensions. It is understandable to those skilled in the art that some well-known structures and their descriptions may be omitted in the drawings.

[0028] In the accompanying drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components. In the description of the present invention, it should be understood that if terms such as "upper," "lower," "left," "right," "front," and "rear" indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, they are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, the terms used to describe positional relationships in the drawings are only for illustrative purposes and should not be construed as limiting the present invention. For those skilled in the art, the specific meaning of the above terms can be understood according to the specific circumstances.

[0029] Purpose: 1. To address the shortcomings of existing methods in predicting dynamic environmental changes and the short-sightedness of motion planning, a prediction-risk coupled decision-making and planning algorithm based on collaborative perception and uncertainty quantification is proposed. The algorithm explicitly quantifies the uncertainty of spatiotemporal evolution prediction through deep integration technology and deeply couples this uncertainty with the spatiotemporal collision probability to construct a high-order decision-making and motion planning system with an autonomous circuit breaker mechanism under extreme conditions.

[0030] 2. To address the problem of low collaborative efficiency in multi-agent systems under limited communication bandwidth, a self-evolving communication and information compression collaborative mechanism based on reinforcement learning is proposed. By using uncertainty as the core driving source, a self-evolving dynamic communication trigger threshold and adaptive information compression rate are constructed to achieve efficient and high-fidelity collaboration under bandwidth constraints.

[0031] 3. To address the problem that traditional security constraint methods cannot dynamically adapt to environmental uncertainties, a security filtering control method based on uncertainty adjustment is proposed. Uncertainty is introduced into the underlying control barrier function to construct a dynamically expanding security boundary for multiple scenarios, thereby achieving minimal intrusive security filtering of upper-level nominal commands.

[0032] 4. To address the challenges of joint optimization of parameters in multi-module systems and poor adaptability to new scenarios, a collaborative adaptive optimization mechanism for decision planning, communication, and control based on meta-reinforcement learning is proposed. This mechanism introduces a meta-reinforcement learning framework, constructs a global multi-objective collaborative reward function, and enables online decoupling and updating of four independent parameter sets (decision, communication, and control) and rapid cold-start deployment with minimal samples across scenarios.

[0033] I. Prediction-Risk Coupled Decision and Planning Algorithm Based on Collaborative Perception and Uncertainty Quantification 1. Method Overview and Module Structure This method is implemented collaboratively by a collaborative sensing coding module, a latent space multi-step dynamic prediction module, a state decoding and reconstruction module, a deep integration uncertainty quantification module, a spatiotemporal evolution risk quantification module, and a prediction-risk coupled decision and planning module. The data flow relationship between the modules is as follows: the collaborative sensing coding module integrates its own sensor observations and communication information with neighboring intelligent agents to generate a fused latent state; the latent space multi-step dynamic prediction module generates a future predicted latent state based on the fused latent state and candidate spatiotemporal trajectory clusters in an autoregressive manner; the state decoding and reconstruction module restores the latent state to the task-related physical prediction state; the uncertainty quantification module outputs the prediction mean and covariance matrix; the spatiotemporal evolution risk quantification module calculates the comprehensive risk value of the candidate trajectories; and finally, the decision and planning module outputs the optimal reference trajectory that balances efficiency and safety.

[0034] 2. Cooperative sensing encoding and latent space mapping In multi-agent systems, the first At time t, the intelligent agent not only collects local observation data through vehicle-mounted sensors. Simultaneously, it receives the intent information of neighboring intelligent agents transmitted from the communication network. The perceptual coding module incorporates a multimodal encoder neural network. The high-dimensional local observation data and neighbor information are fused at the feature level and compressed into a low-dimensional latent space to obtain the fused latent state vector. : (1) in, For encoder network parameters, The hidden state vector. This is the hidden space dimension.

[0035] 3. Multi-step dynamic state prediction based on latent space To assess the future evolution trends of different planning schemes, the latent space multi-step dynamic prediction module incorporates a dynamic transition network. Given the current fusion hidden state A cluster of candidate spatiotemporal trajectories generated by sampling within the control action space The dynamic transition network outputs the predicted hidden state for the next time step: (2) in, To dynamically transfer network parameters, These represent the path points of the candidate trajectory at the next time step. The state decoding and reconstruction module incorporates a decoder neural network. This reduces the predicted hidden state to a definite physical state: (3) in, These are the decoder network parameters.

[0036] Through continuous By taking an autoregressive input of candidate trajectory points, the system can generate future trajectories within the latent space. The predicted state sequence at each time step provides forward-looking information for subsequent spatiotemporal risk assessment.

[0037] 4. Quantitative assessment of forecast uncertainty Due to the randomness of environmental dynamics and the unreliability of multi-agent intentions, the uncertainty quantification module employs a deep ensemble approach for evaluation. Parallel execution is also implemented. A dynamically transitioned subnetwork with identical structure but initialized parameters is obtained. Group prediction distribution. The mean of the predicted state is calculated based on unbiased estimation. With uncertainty covariance matrix : (4) (5) in, It is the spatiotemporal uncertainty covariance matrix that characterizes the degree of divergence in the evolution of dynamic obstacle trajectories and the ambiguity of the intentions of neighboring intelligent agents. The diagonal elements represent the prediction variance of each state component, and the off-diagonal elements represent the prediction covariance between different state components. The prediction confidence index is defined as... The smaller the prediction variance, the more confident the system is in the future evolution of the current candidate trajectory.

[0038] 5. Quantitative Calculation of Risk in Multidimensional Spatiotemporal Evolution For each sequence in the candidate trajectory cluster The spatiotemporal evolution risk quantification module combines the predicted mean and covariance matrix to calculate the cumulative collision risk in the prediction time domain. The instantaneous collision risk assessment covers static obstacles, non-communicating dynamic obstacles, and communicating neighboring agents. For any potential collision object... Based on the Gaussian distribution assumption, the first The instantaneous collision probability at the predicted time is analytically calculated using the standard normal cumulative distribution function: (6) in, To predict the mean relative distance, The standard deviation of the predicted distance is obtained by projecting the covariance matrix onto the direction of the line connecting the agent and the obstacle. This is the safety threshold.

[0039] Fusion time decay factor In conjunction with the uncertainty penalty term, define the comprehensive path risk function for candidate spatiotemporal trajectories: (7) in, and These are the weighting coefficients.

[0040] 6. Prediction-Risk Coupled Decision Making and Planning The decision-making and planning module, as the core of this method, is responsible for comprehensively weighing the spatiotemporal evolution risk value and navigation efficiency index, and executing the strategy switching mechanism according to the uncertainty state.

[0041] First, under normal operating conditions, a multi-objective trajectory optimization problem is constructed to find the optimal candidate trajectory in the current time domain: (8) in, This provides a feasible candidate trajectory space for vehicle kinematics and dynamics. For efficiency weighting coefficients, An efficiency evaluation function for measuring trajectory smoothness and proximity to the target point.

[0042] To address environments of extreme uncertainty, this method establishes an autonomous decision-making mechanism for strategy switching driven by both risk and uncertainty. The optimal reference trajectory ultimately sent to the control layer is defined as... Its decision logic satisfies the following piecewise function: (9) in, The maximum tolerable comprehensive risk threshold preset for the system. The maximum tolerable uncertainty trace threshold. For a pre-set spatiotemporal trajectory (such as an emergency braking trajectory generated based on the maximum deceleration, or a trajectory that retraces along the original path of a historical safe zone).

[0043] The physical meaning of the above decision-making mechanism lies in: when the optimal trajectory is obtained through conventional multi-objective optimization... If the residual risk still exceeds the threshold, or if the overall unreliability of the model prediction (trace of uncertainty covariance) exceeds the limit due to drastic environmental dynamics, the system will proactively abandon forward planning, force a downgrade, and output a preset conservative safety trajectory. The final output reference trajectory The results of the upper-level planning will be sent down to the lower-level motion controller and control barrier function safety filter to derive the final lower-level control commands.

[0044] The core innovation of this algorithm lies in: breaking the perception silos of a single vehicle by performing latent space feature and pre-fusion of local observations and multi-agent communication intentions; introducing a deep integration mechanism for multi-step autoregressive prediction, which, while extrapolating future spatiotemporal trajectories, achieves explicit quantification of prediction unreliability; deeply coupling prediction uncertainty with spatiotemporal collision probability, and constructing a threshold-driven switching mechanism, enabling the system to identify potential collision risks in advance, execute optimal forward planning that balances efficiency and safety under normal operating conditions, and autonomously downgrade to a conservative strategy in extreme unknown environments, fundamentally improving the forward-looking decision-making ability and safety assurance of multi-agent systems in complex dynamic environments.

[0045] Figure 1 This is a flowchart of a prediction-risk coupled decision-making and planning algorithm based on collaborative perception and uncertainty quantification.

[0046] II. A cooperative mechanism for self-evolutionary communication and information compression based on reinforcement learning In multi-agent systems, limited communication bandwidth and highly dynamic environmental changes are the main bottlenecks in collaborative decision-making. This paper proposes a reinforcement learning-based self-evolving communication and information compression collaborative mechanism, consisting of four sub-modules: communication-triggered decision-making, adaptive information compression, received information fusion, and self-evolving joint update. This mechanism leverages spatiotemporal prediction uncertainty to drive communication decisions and feeds back the fused collaborative intent to the multimodal perception coding module, achieving a deep closed loop of perception, communication, and planning.

[0047] 1. Uncertainty-driven communication-triggered decision module The communication-triggered decision module is used to determine whether the agent needs to send intent information to neighboring agents at the current moment to assist in spatiotemporal trajectory planning. Define the first... An intelligent agent in The communication decision variable at time t is .when When, it indicates an intelligent agent Trigger a communication action, encode its state information, and send it to neighboring agents within the communication range; when When, it indicates an intelligent agent Remain silent and do not occupy communication channels.

[0048] Construct a reinforcement learning communication policy value network, whose inputs include the current perceived state. Historical Hidden State and multi-step prediction uncertainty quantification indicators The output is a value estimate of taking either communication or silent action: (10) in, For the value network parameters. Calculate the value gain of triggered communication compared to silent marginal communication: (11) To enable the communication mechanism to perceive environmental dynamics, an uncertainty-driven dynamic triggering threshold function is designed. : (12) in, Based on the trigger threshold, For uncertainty sensitivity coefficient, The minimum communication threshold lower limit to maintain basic system stability.

[0049] The final communication trigger condition is: (13) The physical characteristics of this mechanism are as follows: when the environment is simple and the prediction is reliable, the threshold is increased to save bandwidth; when the prediction uncertainty increases significantly, such as when encountering complex narrow road sections or unclear neighbor intentions, the dynamic threshold is adjusted. Automatic reduction makes it possible even if the marginal communication value gain... The system is small and easily meets the triggering conditions, thus allowing experiments to demonstrate the self-evolving communication behavior of "the stronger the uncertainty, the more active the communication".

[0050] 2. Adaptive Information Compression Module When communication triggers the decision module to determine At this time, the adaptive information compression module is responsible for compressing and encoding the information to be sent in order to further reduce the communication bandwidth usage.

[0051] Define the compression mapping function as ,in For the original state space, Let this be the compressed message space. For intelligent agents At any moment The generated raw information to be sent can be specified as a combination of any one or more of the following: current local state, predicted trajectory sequence, action intention, hidden state representation, or uncertainty information.

[0052] Define a learnable compression encoder: (14) Among them, Learnable compression encoder neural network, Its network parameters, These are compression ratio control parameters. This is the compressed communication information.

[0053] The receiving agent uses a decoder Recovery information: (15) in, For decoder network parameters, To reconstruct information.

[0054] Unlike traditional fixed compression ratio methods, this invention introduces uncertainty-driven adaptive compression fidelity: (16) in, This is the lower limit of the minimum compression ratio. This is the sensitivity coefficient. When the uncertainty is high, When the value approaches 1, the system uses a low compression ratio to ensure that key game information is not distorted; when the uncertainty is low, the system uses a high compression ratio. Approaching This reduces the amount of communication data and minimizes bandwidth usage.

[0055] 3. Receiver confidence-weighted information fusion module The receiving information fusion module is used to integrate the agent's own state with the received neighbor communication information to generate a fused state representation for decision-making.

[0056] Let intelligent agents At any moment Received from neighbor set Decoding information The fusion process is represented as: (17) in, To converge networks, For network parameters, This is the confidence weight for neighbor information. This weight is determined by the uncertainty of the neighbor information. Decide: (18) in, The uncertainty attenuation coefficient is... To prevent division by zero by small constants, this weighting mechanism ensures that the system assigns lower weights to information from high-uncertainty sources and higher weights to information from low-uncertainty sources, thereby reducing the interference of noise propagation on collaborative decision-making.

[0057] The fusion status representation generated by the receiving information fusion module This will be used as a high-order global perception feature and directly input into the multimodal perception coding module as the basic input data for generating multi-step predicted states and trajectory evolution.

[0058] 4. Multi-objective joint reward and communication-compression strategy cooperative self-evolution To achieve the optimal balance between communication efficiency and collaborative performance, this invention constructs a joint reward function to perform end-to-end collaborative optimization of communication triggering strategy, information compression strategy, and action strategy.

[0059] Define the communication reward function as follows: (19) in, For reference trajectory The corresponding spatiotemporal evolution comprehensive risk value; For the cost of communication, This refers to the amount of data transmitted in this communication. For information reconstruction error; , , For the weighting coefficients, satisfying It is used to adjust the relative importance of trajectory safety, communication overhead, and compression quality.

[0060] Based on this joint reward signal, the agent uses a reinforcement learning temporal difference algorithm to update the gradient of the value network online. Let the communication decision state set be... Calculate the time-difference objective With timing difference error : (20) (twenty one) in, As a discount factor, The target value network parameters.

[0061] Construct a loss function that minimizes the squared temporal difference error, and update the communication value network using gradient descent: (twenty two) in, The learning rate of the value network.

[0062] This mechanism enables the agent to autonomously correct its long-term value assessment of communication behavior in unknown dynamic scenarios without manual recalibration of parameters. Through real-time collection of interaction samples and online gradient iteration of Equation (22), it eventually evolves into an optimal game strategy that balances high-fidelity collaborative planning and extremely low communication redundancy, thus realizing the self-evolution of the self-evolving value network.

[0063] As the value network evolves, the encoder and decoder parameters of the adaptive information compression module ( , This is also incorporated into the joint optimization framework. The system aims to minimize the penalty term (i.e., trajectory risk and reconstruction distortion) in the joint reward, and constructs the joint loss function for the compressed network: (twenty three) The gradient of the joint loss with respect to the encoder and decoder parameters is calculated using the backpropagation algorithm, and then updated synchronously end-to-end. (twenty four) in, To compress the learning rate of the network.

[0064] The aforementioned mechanism ensures that the communication triggering strategy and the information compression strategy are no longer trained in isolation. Both are trained under the same multi-objective joint reward. Driven by this, synchronous self-evolutionary updates are achieved by utilizing the temporal difference error of equation (22) and the policy gradient of equation (24). This deep joint optimization ensures that the system can autonomously explore the optimal cooperative game equilibrium point of "high-fidelity communication in high-risk situations and high compression or silence in low-risk situations" in a bandwidth-constrained dynamic environment.

[0065] Figure 2 This is a flowchart of a self-evolving communication and information compression collaborative mechanism based on reinforcement learning.

[0066] III. Security Filtering Control Method Based on Uncertainty Adjustment Control Barrier Function 1. Methodology Overview and Control Architecture In a multi-agent cooperative control system, this method serves as the bottom-level safety control layer, constructing a three-layer control closed loop of "planning-tracking-safety". It aims to perform real-time safety calibration of the reference trajectory generated by the upper-level motion planning through a safety filtering algorithm based on quadratic programming.

[0067] Specifically, the closed loop first receives the final reference trajectory generated by the prediction-risk coupled decision and planning module. Then, with the help of the underlying trajectory tracking controller, it is converted into nominal reference control commands. Based on this, the safety filtering module reads its own prediction uncertainty in real time. and the uncertainty of neighbor in communication reception Based on this, the safety boundary is dynamically adjusted, and minimally invasive modifications are applied to the nominal control command, ultimately outputting an absolute safety control command that satisfies the safety constraints. .

[0068] 2. Uncertainty-driven adaptive safety function To address the issues of fixed safety boundaries and insufficient safety margins in high-uncertainty scenarios associated with traditional control barrier functions, this method constructs an adaptive safety function based on prediction confidence. .

[0069] First, from the covariance matrix Extract the location-related submatrix and calculate the uncertainty scalar of the location prediction. : (25) in, Covariance matrix The submatrix corresponding to the position component.

[0070] Based on location uncertainty Define intelligent agents Dynamic adaptive safety distance: (26) in, The basic safety distance is determined by physical constraints such as agent size and braking distance. This is the adjustment coefficient for the uncertainty on the safety distance boundary, controlling the degree to which uncertainty amplifies the safety distance.

[0071] Adaptive security function for complex dynamic environments Specifically, it can be further divided into the following three scenarios and sub-functions: (1) Static obstacle safety function : For static obstacles in the environment Its positional uncertainty is zero, and the safety function is adjusted solely by the agent's own prediction uncertainty: (27) in, and These are the current position vectors of the agent and the static obstacle, respectively.

[0072] (2) Dynamic obstacle safety function : For dynamic obstacles that cannot communicate This method utilizes a spatiotemporal dynamic prediction module to extrapolate the trajectory of the target object in the future time domain. The moment when the predicted distance between the two objects is minimized within the prediction time domain is defined as the most dangerous moment. The safety function not only considers its own safety distance but also needs to compensate for the uncertainty in predicting dynamic obstacles. (28) in, The predicted location of the most dangerous moment. The uncertainty in the model's prediction of the dynamic obstacle at this moment is... This is the magnification factor for external dynamic objects.

[0073] (3) Cooperative security function between agents : (29) in, It is based on the received neighbor uncertainty The calculated dynamic safe distance between neighbors.

[0074] The technical meaning of the adaptive safety function is: when environmental complexity leads to increased prediction uncertainty, the adaptive safety distance... As the safety boundary increases, it automatically expands outward, forcing the system to take more conservative obstacle avoidance actions; when uncertainty is low, the safety distance approaches the baseline value. The system can allow for more compact passage behavior while ensuring safety.

[0075] 3. Control barrier function constraints for uncertainty adjustment Substituting the adaptive safety function into the control barrier function constraints, we obtain the control barrier function constraints for uncertainty adjustment: (30) in, The gain coefficient, which controls the decay rate of the safety function, determines the smoothness of the buffer as the agent approaches the safety boundary.

[0076] Combining the system's physical dynamics model with the real-time rate of change of prediction uncertainty According to the Lie derivative expansion rule, equation (30) is transformed into a function of the underlying control input. Direct linear constraints: (31) in, and These are the Lie derivatives of the adaptive safety function along the system drift mechanics and the control input direction, respectively.

[0077] The left side of this constraint is the control input. The linear function is given by a dynamic, time-varying threshold established on the right side. From a physical perspective, when the prediction uncertainty is rapidly increasing (i.e.,...), the threshold is... When this occurs, the threshold on the right side of the constraint increases significantly, thus affecting the control input. Imposing stricter restrictions prompts the system to take absolutely safe actions such as braking or evasive maneuvers in advance.

[0078] 4. Algorithm for solving security filtering based on quadratic programming This module achieves minimally invasive modification of the nominal instructions at the higher level by solving a quadratic programming problem. The objective function and constraints of the quadratic programming are defined as follows: (32) in, The set of physically feasible control inputs for the actuator; This is the optimal control input after safety correction.

[0079] The objective function of this quadratic programming problem is to find a control input that satisfies safety constraints under the constraint of minimizing modifications to the reference control command. If the nominal command derived from the trajectory planned by the upper layer... If the product is already within the adaptive safety boundary, then the optimization result... The safety filter is in a dormant state and does not intervene in the instructions. If there is a collision risk with the upper-level instructions, or if the original safety boundary expands and covers the current state due to a surge in environmental prediction uncertainty, the quadratic programming solver will search for the optimal solution within the safe feasible region that has the closest Euclidean distance to the nominal instructions. This minimally invasive modification ensures that the system retains the decision-making intent of the motion planning module of the prediction-risk coupled decision and planning module to the greatest extent possible, while maintaining absolute safety. Figure 3 This is a flowchart of a safety filtering control method based on uncertainty-adjusted control barrier function.

[0080] IV. Decision-Programming-Communication-Control Cooperative Adaptive Optimization Mechanism Based on Meta-Reinforcement Learning In multi-agent collaborative decision-making, planning, communication, and control systems, the optimal parameter configurations of the spatiotemporal evolution prediction module, the self-evolving communication module, and the control barrier function security constraint module are highly dependent on the dynamic characteristics of the current environment, such as obstacle density and communication topology jump frequency. Traditional reinforcement learning methods suffer from a severe "cold start" problem when facing entirely new deployment scenarios, resulting in extremely high trial-and-error costs and a high risk of security failure. To address this, this invention proposes a two-layer collaborative adaptive optimization mechanism based on meta-reinforcement chemistry. By extracting general meta-prior knowledge across scenarios offline, it enables the system to rapidly generalize and co-evolve parameters with minimal samples in unknown online environments.

[0081] 1. Construction of a global collaborative multi-objective reward function To drive cross-module updates in meta-learning, a global multi-objective collaborative reward function is first constructed that comprehensively considers efficiency, communication costs, and underlying security performance. : (33) in, , , These are the weighting coefficients for each sub-objective. The reward for each sub-item is defined as follows: Navigation and Planning Awards : (34) The final reference trajectory generated based on the prediction-risk coupled decision and planning module , Used to evaluate the spatiotemporal evolution risk and traffic efficiency of the planning module, aiming to encourage the system to generate trajectories with low spatiotemporal evolution risk and high efficiency.

[0082] Communication efficiency reward Using formula (19):

[0083] Used to punish high communication data volume costs With information compression and reconstruction distortion .

[0084] Safety Intervention Rewards : (35) in, This represents the cost of the security filter's intrusive modification of the original schematic diagram. This constitutes a significant penalty in the event of extreme security conflict. To control the violation of the indicator function by the barrier function constraint. Used to penalize intrusive safety corrections and absolute collision violations by the underlying controller.

[0085] 2. Construction of Meta-learning Task Distribution and Definition of Global Parameters Unlike traditional reinforcement learning, which trains in a single environment, meta-reinforcement learning focuses on sampling and learning the task distribution. This method encapsulates the obstacle kinematics, communication channel fading model, and agent density of a specific environment into an environmental task. .

[0086] Define a set of global meta-initialization parameters covering the entire link decision planning, communication and control modules. : (36) in, It is a set of parameters for spatiotemporal prediction and planning; It is a set of communication-triggered decision parameters; It is an adaptive information compression and fusion parameter set; It is the set of underlying control security parameters.

[0087] 3. Offline Meta-Reinforcement Learning Two-Layer Optimized Meta-Training Mechanism During the offline simulation training phase, the system employs model-independent meta-learning to construct a two-layer gradient update architecture, aiming to optimize a set of meta-initialization parameters that are highly sensitive to environmental changes. For each meta-training iteration, from the distribution Sampling a batch of tasks For any specific task Divide its interaction data into support sets and query set .

[0088] Inner layer update: In the mission Internally, the intelligent agent utilizes support sets Perform exploratory testing, calculate gradients on local tasks, and perform rapid parameter fine-tuning in one or more steps: (37) in, For the inner layer adaptive learning rate, For the current environment only The task-specific parameters have been fine-tuned. This step allows the parameters of each module to be initially adapted to the local scenario.

[0089] Outer layer update The true purpose of meta-learning is not optimization. Instead, it utilizes the finely tuned parameters. In query set Tests were conducted to evaluate its fast generalization ability, and then the results were calculated relative to the original meta-parameters. The meta-gradient performs a global update across tasks: (38) in, Let be the learning rate of the outermost element. Through the second derivative information transmission in equation (38), the system forces... It converges to a "parameter-sensitive region", meaning that starting from this state, it can reach the optimal solution with only a very small number of gradient steps when facing any new environment.

[0090] 4. Zero-latency cold start and rapid adaptive deployment via online collaboration Based on high-quality meta-prior knowledge extracted in the offline phase This system is designed with an efficient online multi-module collaborative generalization mechanism.

[0091] When an intelligent agent is deployed in a completely new real-world scenario with unknown physical dynamics and communication interference, it is directly loaded... Initialize each submodule. Based on this "meta-prior" foundation, the system only needs to collect a very small number of online prior sample sequences (typically 10-20 decision interaction cycles) to directly trigger the online collaborative gradient update law: Planning perception layer update: (39) in, The learning rate is the parameter set for spatiotemporal prediction and planning.

[0092] Update on communication decision-making level: (40) in, The learning rate for triggering the decision parameter set for communication.

[0093] Compressed fusion layer update: (41) in, The learning rate is used for the adaptive information compression and fusion parameter set.

[0094] Security boundary layer update: (42) in, The learning rate for the underlying control security parameter set.

[0095] Through the decoupled online update design that combines meta-learning mechanisms, the decision-making, communication, and low-level control modules can achieve the same global multi-objective reward. Driven by this, it quickly transitioned from the "factory-standard mode" to the "current environment-optimal customized mode," breaking through the technical bottleneck of traditional reinforcement learning's extremely slow online convergence and high risk of catastrophic collisions.

[0096] Figure 4 This is a flowchart of a decision-planning-communication-control collaborative adaptive optimization mechanism based on meta-reinforcement learning.

[0097] Advantages and innovations of this invention A prediction-risk coupled decision-making and planning algorithm based on collaborative perception and uncertainty quantification is proposed. By predicting the future environmental evolution through multi-step recursion in the latent space and quantifying the prediction credibility, the prediction-risk coupled decision-making and strategy switching protection are combined, enabling the system to identify potential collision risks in advance and improve the foresight and safety of motion planning in complex dynamic environments.

[0098] We propose a self-evolving communication and information compression collaboration mechanism based on reinforcement learning. This mechanism uses a communication policy network to adaptively decide on communication timing and leverages uncertainty-driven compression rate control to dynamically adjust information fidelity. This significantly reduces communication bandwidth usage in large-scale agent clusters, achieving an optimal balance between communication efficiency and collaboration performance.

[0099] A safety filtering control method based on uncertainty adjustment is proposed. The prediction uncertainty is explicitly introduced into the definition of the safety function and safety boundary, an adaptive safety distance is constructed, and a two-layer control of planning and safety is achieved through a quadratic programming safety filtering architecture. This alleviates the over-conservative problem of traditional control barrier function methods and improves the passage efficiency in narrow scenarios.

[0100] We propose a decision-planning-communication-control collaborative adaptive optimization mechanism based on meta-reinforcement learning. By extracting general meta-prior knowledge across scenarios offline, it can drive the relevant parameters of decision-making, communication, compression and control with only a small number of samples in newly deployed scenarios. Based on global multi-objective rewards, it achieves synchronous evolution and rapid convergence, which greatly improves the generalization ability and deployment efficiency of multi-agent systems in industrial applications.

[0101] A four-in-one multi-agent collaborative method of "decision planning-communication-control-adaptation" is proposed. Each module has clear functional boundaries, well-defined interfaces, and collaborative operation. It can be deployed on heterogeneous computing platforms and achieves full-link self-evolution of the system while ensuring real-time performance.

[0102] Figure 5 A flowchart of a multi-agent collaborative method integrating decision-making, planning, communication, control, and adaptation.

[0103] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A multi-agent decision-planning-communication-control system based on meta-reinforcement learning and uncertainty-driven approaches, characterized in that: It includes a prediction-risk coupled decision-making and planning module with collaborative perception and uncertainty quantification, a reinforcement learning self-evolutionary communication and information compression collaborative module, a control barrier function security filtering control module with uncertainty adjustment, and a meta-reinforcement learning collaborative adaptive optimization module. The collaborative perception and uncertainty quantification prediction-risk coupled decision and planning module is used to generate future prediction hidden states, quantify prediction uncertainty and calculate comprehensive path risk, and output the optimal reference trajectory. The reinforcement learning self-evolutionary communication and information compression collaboration module is used to drive dynamic decision-making based on uncertainty and adaptively adjust the information compression rate to achieve low-bandwidth high-fidelity collaboration. The uncertainty-adjusted control barrier function safety filter control module is used to perform minimally invasive safety correction on the optimal reference trajectory and output safety control commands. The meta-reinforcement learning collaborative adaptive optimization module is used to construct a global multi-objective collaborative reward function, and achieves collaborative optimization of decision, communication and control parameters through offline meta-training and online fast adaptive updating.

2. The multi-agent decision-planning-communication-control system based on meta-reinforcement learning and uncertainty-driven approach according to claim 1, characterized in that: The collaborative perception and uncertainty quantification prediction-risk coupled decision-making and planning module specifically includes: The collaborative sensing coding unit is used to fuse local observation data with the intent information of neighboring agents to generate a fused hidden state vector. ,in, For the i-th agent in t The fused hidden state vector at time step 1. For encoder network parameters, For implicit space dimension; Latent space multi-step dynamic prediction unit, used for fusion of latent state vectors With candidate spatiotemporal trajectory clusters Generate the future H The predicted hidden state sequence at each time step; Deeply integrated uncertainty quantization unit, used to... K Given two dynamically transitioning subnetworks with identical structures but initialized parameters, calculate the mean of the predicted states. With uncertainty covariance matrix ,in, , For the first k The predicted state of each subnetwork K The total number of sub-networks; The spatiotemporal evolution risk quantification unit is used to calculate the comprehensive path risk function of candidate spatiotemporal trajectories based on the predicted mean and the uncertainty covariance matrix. ,in, , , These are the weighting coefficients. The time decay factor, For the first k Predicting the timing agent i Collision object j The instantaneous collision probability; The decision and planning unit is used to output the optimal candidate trajectory under normal operating conditions. Output a preset safety trajectory under extremely uncertain operating conditions. The final output is the optimal reference trajectory. .

3. The multi-agent decision-planning-communication-control system based on meta-reinforcement learning and uncertainty-driven approach according to claim 1, characterized in that: The reinforcement learning self-evolutionary communication and information compression collaborative module specifically includes: The communication-triggered decision unit is used to calculate the marginal communication value gain based on the communication strategy value network. And combined with uncertainty-driven dynamic trigger threshold Determine communication decision variables ,in, , , For communication strategy value network, This is the lower limit of the minimum communication threshold. Based on the trigger threshold, This is the uncertainty sensitivity coefficient; An adaptive information compression unit for uncertainty-driven adaptive compression fidelity. The information to be sent is compressed and encoded, where, , This is the lower limit of the minimum compression ratio. This is the sensitivity coefficient; The receiving information fusion unit is used to calculate confidence weights based on the uncertainty of neighbor information. Generate a fused state representation by integrating its own state with neighbor information. ,in, , The uncertainty attenuation coefficient is... To prevent small constants from being divided by zero, Gather for neighbors; Joint optimization unit, used for communication reward function End-to-end collaborative optimization of communication triggering strategy and information compression strategy is performed, wherein, , , , These are the weighting coefficients. For the cost of communication, This refers to the information reconstruction error.

4. The multi-agent decision-planning-communication-control system based on meta-reinforcement learning and uncertainty-driven approach according to claim 1, characterized in that: The uncertainty-adjusted control barrier function security filtering control module specifically includes: Adaptive safety function building block for location-based prediction uncertainty scalar Build a dynamic adaptive safe distance ,in, , , This is a submatrix representing the corresponding positional components in the covariance matrix. Based on the safe distance, This is the adjustment coefficient for uncertainty regarding the safety distance boundary; The control barrier function constraint generation unit is used to substitute the adaptive safety function into the control barrier function constraint conditions to generate linear constraints on the control input, wherein... , The gain coefficient used to control the decay rate of the safety function; The quadratic programming safety filtering unit is used to output safety control commands by solving a quadratic programming problem. The objective function is: , U For the physically feasible set of control inputs for the actuator, For nominal reference control instructions.

5. The multi-agent decision-planning-communication-control system based on meta-reinforcement learning and uncertainty-driven approach according to claim 1, characterized in that: The meta-reinforcement learning collaborative adaptive optimization module specifically includes: A global multi-objective collaborative reward function construction unit is used to construct a global multi-objective collaborative reward function that comprehensively considers planning efficiency, communication cost, and underlying security performance. ,in, , , , These are the weighting coefficients. Rewards for navigation and planning, Rewards for communication efficiency Rewards for safety interventions; Meta-learning task distribution building unit, used to encapsulate environmental tasks into task distributions. Define a global meta-initialization parameter set. ,in, , For spatiotemporal prediction and planning parameter set, The set of decision parameters is used to trigger communication. For adaptive information compression and fusion parameter set, This is a set of underlying control security parameters; The offline meta-training unit is used for meta-training via a two-layer gradient update architecture, with the inner update based on the support set. Fine-tune task-specific parameters Outer meta update based on query set Optimize meta-initialization parameters ; Online collaborative fast adaptive unit, used in new deployment scenarios, based on a global multi-objective collaborative reward function. Global meta-initialization parameter set Perform online collaborative gradient updates.

6. A multi-agent decision-planning-communication-control collaborative method based on meta-reinforcement learning and uncertainty-driven approaches, characterized in that: Includes the following steps: S1: Through collaborative perception and uncertainty quantification, prediction-risk coupled decision-making and planning steps, generate future prediction latent states, quantify prediction uncertainty and calculate comprehensive path risk, and output the optimal reference trajectory; S2: By using reinforcement learning to coordinate self-evolutionary communication and information compression steps, the communication triggering timing is driven by uncertainty, and the information compression rate is adaptively adjusted to achieve low-bandwidth high-fidelity coordination. S3: Through the safety filtering control step of the control barrier function adjusted for uncertainty, the optimal reference trajectory is modified with minimal intrusion, and a safety control command is output. S4: Through meta-reinforcement learning collaborative adaptive optimization steps, a global multi-objective collaborative reward function is constructed. Through offline meta-training and online fast adaptive updates, collaborative optimization of decision, communication, and control parameters is achieved.

7. The multi-agent decision-planning-communication-control cooperative method based on meta-reinforcement learning and uncertainty-driven approach according to claim 6, characterized in that: S1 specifically includes: S11: Fuse local observation data with the intent information of neighboring agents to generate a fused hidden state vector. ,in, For the first i An intelligent agent in t Local observation data at time, For the neighboring intelligent agent's intent information, For multimodal encoder neural networks, For encoder network parameters; S12: Based on fusion of hidden state vectors With candidate spatiotemporal trajectory clusters Through dynamic transfer network Generate the predicted hidden state for the next time step. And through the decoder neural network Restore to the physical prediction state ,in, To dynamically transfer network parameters, These are the decoder network parameters; S13: Calculate the mean of the predicted state using a deep ensemble method. With uncertainty covariance matrix ; S14: Calculate the comprehensive path risk function of candidate spatiotemporal trajectories based on the predicted mean and the uncertainty covariance matrix. ; S15: Output the optimal candidate trajectory under normal operating conditions Output a preset safety trajectory under extremely uncertain operating conditions. The final output is the optimal reference trajectory. ,in, This represents the feasible candidate trajectory space. For efficiency weighting coefficients, This is the efficiency evaluation function.

8. The multi-agent decision-planning-communication-control cooperative method based on meta-reinforcement learning and uncertainty-driven approach according to claim 6, characterized in that: S2 specifically includes: S21: Value Network Based on Communication Strategy Calculate the value gain of marginal communication And combined with uncertainty-driven dynamic trigger threshold Determine communication decision variables ,when hour, ,otherwise ; S22: When At that time, adaptive compression fidelity based on uncertainty-driven The information to be sent is compressed and encoded to generate compressed information. ,in, For learnable compression encoder neural networks, For encoder network parameters, Original information to be sent; S23: The receiver uses a decoder. Recovery Information ,in, These are the decoder network parameters; S24: Calculating confidence weights based on the uncertainty of neighbor information Generate a fused state representation by integrating its own state with neighbor information. ,in, To converge networks, To integrate network parameters, The current perception state; S25: Based on the communication reward function End-to-end collaborative optimization of communication triggering strategy and information compression strategy.

9. The multi-agent decision-planning-communication-control cooperative method based on meta-reinforcement learning and uncertainty-driven approach according to claim 6, characterized in that: S3 specifically includes: S31: Scalar based on location prediction uncertainty Build a dynamic adaptive safe distance ; S32: Based on dynamic adaptive safety distance Constructing an adaptive security function And substitute it into the control barrier function constraints. Generate linear constraints on the control input. ,in, and These are the Lie derivatives of the adaptive safety function along the system drift dynamics and the control input direction, respectively. The real-time rate of change of location prediction uncertainty; S33: Solving a quadratic programming problem Output safety control commands .

10. The multi-agent decision-planning-communication-control cooperative method based on meta-reinforcement learning and uncertainty-driven approach according to claim 6, characterized in that: S4 specifically includes: S41: Construct a global multi-objective collaborative reward function ,in, , , , For security conflict penalties, To control the violation of the indicator function by the barrier function constraint; S42: Encapsulate environmental tasks as task distributions Define the global meta-initialization parameter set. ; S43: During the offline simulation training phase, meta-training is performed using a two-layer gradient update architecture, with the inner layer update based on the support set. Fine-tune task-specific parameters Outer meta update based on query set Optimize meta-initialization parameters ,in, The inner layer adaptive learning rate, The outermost learning rate; S44: In the new deployment scenario, load the global meta-initialization parameter set. Based on a global multi-objective collaborative reward function Online collaborative gradient updates enable rapid adaptation of decision-making, communication, and control parameters.