Multi-agent vehicle-road-cloud integrated collaborative decision-making and control architecture system and method based on federated reinforcement learning
By adopting a multi-agent vehicle-road-cloud integrated collaborative decision-making architecture based on federated reinforcement learning, the problems caused by information redundancy and privacy awareness in intelligent transportation systems are solved, realizing autonomous driving with deep collaborative decision-making between vehicles and traffic, and improving decision-making efficiency and safety.
Patent Information
- Application Number
- PCT/CN2024/092572
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-23
- Filing Date
- 2024-05-11
- Publication Date
- 2025-10-30
AI Technical Summary
Existing intelligent transportation systems in the field of autonomous driving suffer from problems such as traffic information redundancy, difficulty in extracting key information, and huge communication overhead, resulting in poor vehicle-road cooperative decision-making efficiency and control effect. Furthermore, information asymmetry between vehicles and roads caused by privacy concerns limits the development of integrated cooperative technologies.
A multi-agent vehicle-road-cloud integrated collaborative decision-making architecture based on federated reinforcement learning is adopted. By embedding vehicle dynamics characteristics, the semantic matrix generated by the roadside is used as the input for vehicle-side reinforcement learning to construct a fusion reward function, realizing a comprehensive consideration of vehicle-side safety and comfort. The neural network parameters are transmitted through V2I communication to solve the information asymmetry problem caused by privacy concerns.
It has enabled autonomous driving with deep decision-making and control coordination between vehicles and traffic in intelligent transportation systems, improving decision-making efficiency and safety, solving the problem of information asymmetry, and achieving the selection of local optimal strategies and the balancing of globally shared models.
Smart Images

Figure CN2024092572_30102025_PF_FP_ABST
Abstract
Description
A Multi-Agent Vehicle-Road-Cloud Integrated Collaborative Decision-Making Architecture and Method Based on Federated Reinforcement Learning Technical Field
[0001] This invention belongs to the field of transportation and relates to a multi-agent vehicle-road-cloud integrated collaborative decision-making architecture system and method based on federated reinforcement learning. Background Technology
[0002] Due to limitations in perception range and computing power, single-vehicle intelligence struggles to cope with complex traffic conditions. Intelligent transportation systems based on vehicle-road cooperative technology, through information exchange between vehicles and road infrastructure, can achieve vehicle-road cooperative perception and vehicle-side computing load transfer, providing a new solution for autonomous driving in complex traffic situations.
[0003] Intelligent transportation systems based on vehicle-to-infrastructure (V2X) technology rely on V2X communication, including V2V (vehicle-to-vehicle) communication and V2I (vehicle-to-infrastructure) communication. Through V2X communication, intelligent transportation systems can acquire comprehensive traffic information, including information on road participants and obstacles. V2X technology offers several advantages. First, it expands the driver's field of vision from the driver's perspective to a higher and wider BEV (Battery on the Road) bird's-eye view. Second, V2I communication enables the deployment of some autonomous driving algorithms on the infrastructure, reducing the computational burden on intelligent vehicles. Third, the fixed sensing systems of roadside units (RSUs) can access historical road information, making it easier to detect random obstacles on the road. Based on V2X technology, intelligent transportation systems fully utilize the advantages of the roadside infrastructure, thereby solving challenges such as the limited perception range of a single vehicle, the limited computing power of a single vehicle, and the difficulty in detecting random obstacles.
[0004] Existing intelligent transportation system architectures can provide drivers with additional traffic information and driving assistance. However, in the field of autonomous driving, challenges such as traffic information redundancy, difficulties in extracting key information, and huge communication overhead directly affect the decision-making efficiency and control effectiveness of vehicle-to-infrastructure (V2I) cooperation. This results in the immaturity of research on integrated collaborative architectures serving autonomous driving and a lack of practical application technologies. Furthermore, current V2I technologies typically assume that connected vehicles are connected to the cooperative network, and the information asymmetry between vehicles and roads caused by privacy concerns further restricts the development of integrated collaborative technologies.
[0005] Summary of the Invention
[0006] To address the aforementioned technical challenges, this invention provides a multi-agent vehicle-road-cloud integrated collaborative decision-making architecture based on federated reinforcement learning. Employing a multi-agent federated reinforcement learning framework with embedded vehicle dynamics characteristics, it solves the problem of deep information fusion between Intelligent Transportation Systems (ITS) and Intelligent Vehicles (IV), achieving autonomous driving based on deep collaborative decision-making between vehicles and traffic. A semantic matrix generated from the roadside perspective is used as input for vehicle-side reinforcement learning, constructing roadside-guided global and local trajectory planning for the vehicle. Utilizing the advantages of the roadside, a driving safety field model is implemented, constructing a fusion reward function used in vehicle-side reinforcement learning, achieving a comprehensive consideration of vehicle safety and comfort. Based on the roadside federated learning architecture, vehicle-side neural network parameters are transmitted via V2I communication, solving the problem of information asymmetry between vehicles and roads caused by privacy concerns. For different environmental sample distributions, a locally optimal strategy for the current environment is selected based on a neural network screening process, and a globally shared model benefiting from different environments is synthesized, achieving a balance between sample efficiency and model robustness.
[0007] This invention provides a technical solution for a multi-agent vehicle-road-cloud integrated collaborative decision-making architecture system based on federated reinforcement learning, comprising three main components: cloud, road, and vehicle.
[0008] The cloud platform is primarily used for information transmission and delivery, enabling cloud control applications at different levels. It consists of a cloud control platform and related support platforms, achieving efficient information exchange through V2N vehicle-to-cloud communication combined with optical network communication.
[0009] The cloud control platform can be categorized into three application levels: connected vehicle empowerment, traffic management and control, and traffic data empowerment. These correspond to three cloud control levels within the platform: edge cloud, regional cloud, and central cloud.
[0010] The supporting platform includes, but is not limited to, a traffic management platform that provides a complete road network and real-time traffic information, a logistics platform that provides logistics supervision information, a map platform that provides high-precision maps, a positioning platform that provides high-precision positioning and navigation, and a meteorological platform that provides real-time weather information.
[0011] Based on real-time traffic dynamics collected by the edge cloud, the cloud control platform provides enabling services to improve the driving safety of connected vehicles and enhance traffic efficiency. By combining information from the edge cloud with that from the support platform, the regional cloud provides services such as traffic situation awareness and assessment, traffic planning, traffic order management, and transportation management, achieving benefits such as reducing traffic accidents, alleviating road congestion, and improving road network efficiency. Furthermore, by combining information from the regional cloud with that from the support platform, the central cloud, empowered by traffic big data, provides big data analysis services and business support to universities, research institutions, mobility service providers, and automobile manufacturers.
[0012] For roadside units, the main function is to coordinate vehicle groups based on different levels of cloud-based control applications. A roadside unit consists of multiple roadside units, each containing an information processing module and a group coordination and decision-making module.
[0013] The information processing module processes the bird's-eye view image input into a semantic bird's-eye view map, and generates a Gaussian safety field using roadside advantages. The semantic bird's-eye view map serves as the input to the vehicle-side perception processing module, while the Gaussian safety field participates in generating the fusion reward function in the reinforcement learning module. The Gaussian safety field is established through the following equation: φ=b x / c y =l v / w v
[0014] Among them, S sta C represents the static safety field strength. a The static safety field strength coefficient is represented by x0 and y0, which represent the coordinates of the static risk center O(x0, y0), and ε represents the safety field shape coefficient. x and c y This represents the appearance coefficient of the intelligent connected vehicle, where φ is the aspect ratio of the intelligent connected vehicle, and l v Indicates the length of the vehicle, w v Indicates the width of the vehicle.
[0015] When the intelligent connected vehicle moves, the risk center O(x0,y0) of the Gaussian safety field will shift to a new risk center O′(x′0,y′0) as the vehicle moves:
[0016] Where, k v Indicates the moving adjustment factor, and The sign is related to the direction of motion; β represents the transfer vector of the connected vehicle. The angle between the coordinate axes and the coordinate system in the Cartesian coordinate system. This represents the velocity vector of the connected vehicle. A virtual vehicle of length l′ is formed in the dynamic safety field under the influence of risk center transfer. v Width is w′ v S dyn The dynamic safety field strength is represented by the new aspect ratio φ′=b′. x / c′ y =l′ v / w′ v .
[0017] The group coordination and decision-making module acquires the neural network parameters of intelligent connected vehicles in the vehicle-to-everything (V2I) group through vehicle-to-infrastructure communication. Based on federated learning, it filters and aggregates local vehicle-to-everything (V2I) group neural network parameters and shared neural network parameters to achieve multi-agent group optimization. The group coordination and decision-making process only transmits network parameters, not training samples, ensuring agent privacy while reducing communication and computational overhead. The neural network used for parameter uploading and downloading is a reinforcement learning-based actor-critic network, where the actor network outputs the control policy and the critic network evaluates the policy. For different environmental sample distributions, a locally optimal policy for the current environment is selected based on the neural network filtering process, and a globally shared model that benefits from different environments is synthesized, achieving a balance between sample efficiency and model robustness.
[0018] In the neural network selection process described in this invention, driving safety and driving aggression based on Gaussian safety fields are used as network selection indicators. A rule-based method is used as the basis for selection, achieving rule-driven integration with data, and the interpretability of the algorithm framework is enhanced through interpretable rules. The actor network employs the following aggregation process:
[0019] in, This indicates that the parameters θ of the intelligent connected vehicle itself are... μ,i And another intelligent connected vehicle parameter θ μ,i′ New network parameters obtained through aggregation. Safety The safety-related reward function for intelligent connected vehicles is calculated from both driving safety and driving aggression aspects by the roadside Gaussian safety field. r Safety =r Risk +r Agg
[0020] Among them, R i,j (t) represents the driving risk posed by intelligent connected vehicle j to intelligent connected vehicle i. Let k represent the field strength of intelligent connected vehicle j with respect to intelligent connected vehicle i. c Indicates the risk perception coefficient. Let θ represent the speed of the intelligent connected vehicle j at time t. i,j (t) represents the angle between intelligent connected vehicle i and intelligent connected vehicle j at time t, r Risk Let f represent the reward function related to driving risk. Risk (ξ) represents the driving risk score, Rthr τ represents the risk threshold. rc R represents the duration exceeding the risk threshold. j,i (t′) represents the driving risk posed by intelligent connected vehicle i to intelligent connected vehicle j. Let i represent the field strength of intelligent connected vehicle i for intelligent connected vehicle j. This represents the field strength between intelligent connected vehicle j and intelligent connected vehicle i. Let θ represent the speed of the intelligent connected vehicle i at time t′. j,i (t′) represents the angle between intelligent connected vehicle j and intelligent connected vehicle i at time t′, r Agg Let f represent the reward function related to driving aggression. Agg (ξ) represents the integral of driving aggression.
[0021] The critic network employs the following aggregation process:
[0022] in, This indicates that the parameters θ of the intelligent connected vehicle itself are... l,i And another intelligent connected vehicle parameter θ l,i′ The new network parameters obtained through aggregation This represents the value of the Critic network's output for the state-action pair (s, a).
[0023] For the vehicle side, it is mainly used to output planning and control quantities based on roadside input and sensor information from intelligent connected vehicles. A single roadside unit contains multiple vehicle groups, and each vehicle group consists of N intelligent connected vehicles, which share information through V2V vehicle-to-vehicle communication. The intelligent connected vehicle consists of a perception processing module, a trajectory prediction module, a reinforcement learning module, and an intelligent chassis coupling system.
[0024] The perception processing module matches and segments the semantic bird's-eye view provided by the roadside based on the vehicle's positioning. It then fuses the vehicle's dynamic information (sensor data) and the roadside static information into a state variable matrix, which serves as the input to the reinforcement learning module's neural network. The dynamic information includes the vehicle's speed and position, while the roadside static information includes road, desired path, and lane information. The state variable matrix i... RL ∈[0,1] W×H×C Where W×h=56 represents the matrix size, actually covering a physical area of approximately 14m×14m, and C=4 represents the number of channels.
[0025] The trajectory prediction module uses the semantic bird's-eye view provided by the roadside to predict trajectories at multiple time steps. The prediction results will be used as part of the input to the reinforcement learning module.
[0026] The reinforcement learning module, through inputs from the perception processing module and the trajectory prediction module, combined with a fusion reward function generated by the roadside safety field, interactively generates planning control outputs in the CARLA simulator. This interactive process involves the reinforcement learning actor network generating the planning control output 'a' based on the fusion state variables. t1 With a t2 The output is mapped to the range [-1, 1] using an activation function, where a t1 Indicates the amount of steering wheel control, a t2 It is divided into two parts: [-1, 0] and [0, 1], which represent the braking and throttle control values, respectively.
[0027] The intelligent chassis coupling system is primarily used by a single intelligent connected vehicle to output secondary planning control quantities based on reinforcement learning control output and the desired path. Through a distributed controller, corresponding intelligent chassis subsystems are controlled. These subsystems include, but are not limited to, a steering subsystem, a DYC subsystem, a suspension subsystem, and a drive subsystem. State coupling and cost coupling exist between these subsystems. Specifically, the steering and DYC subsystems calculate the lateral control quantities of the intelligent connected vehicle, the suspension subsystem calculates the vertical control quantities, and the drive subsystem calculates the longitudinal control quantities. The steering and DYC subsystems calculate the yaw moment of the intelligent connected vehicle based on the steering wheel control quantity output by the reinforcement learning module and the pre-aiming error between the actual path and the desired path. The drive subsystem obtains the desired target speed and calculates the four-wheel drive / braking torque of the intelligent connected vehicle based on the throttle control quantity and the current speed of the intelligent connected vehicle. First, the driving force equation of the intelligent connected vehicle is constructed based on PID control.
[0028] Among them, F x (t) represents the driving force of the intelligent connected vehicle at time t, K p Let K represent the proportional gain, e(t) represent the velocity error at time t, and K represent the velocity error at time t. i K represents the integral gain. d This represents the differential gain.
[0029] Then consider the following optimization problem:
[0030] Where F xij F represents the longitudinal force on the four wheels of an intelligent connected vehicle. yij F represents the lateral force on the four wheels of an intelligent connected vehicle. zfl F zfr F zrl F zrrThis represents the vertical force on all four wheels of the intelligent connected vehicle, where fl represents the left front wheel, fr represents the right front wheel, rl represents the left rear wheel, rr represents the right rear wheel, and μ represents the adjustment factor. Since the steering subsystem and DYC subsystem have already calculated the four-wheel steering angles, i.e., F... yij Since ,ij=fl,fr,rl,rr are constants, the above equation simplifies to:
[0031] subject to
[0032] Where m represents the mass of the intelligent connected vehicle, a x d represents the driving acceleration, and M represents the wheelbase of the intelligent connected vehicle. z R represents the additional yaw moment. w T represents the wheel radius of a smart connected vehicle. min T represents the minimum driving torque or braking torque. max This represents the maximum driving torque. The three formulas above respectively indicate that the total driving force satisfies the constraints of the drive subsystem, the additional yaw moment satisfies the constraints of the steering subsystem and DYC subsystem, and the driving force satisfies the actuator constraints.
[0033] Finally, by solving the optimization problem, the four-wheel drive or braking torque is obtained: T ij =F xij R w ,i=fl,fr,rl,rr
[0034] After the intelligent chassis coupling system outputs the secondary programming control quantity, it controls the intelligent connected vehicle in the CARLA simulator to interact and generate experience samples. Based on the experience samples, it trains the neural network parameters of the reinforcement learning module. Then, the roadside group coordination decision module obtains the vehicle-side neural network parameters of the intelligent connected vehicles in the vehicle group through V2I vehicle-to-infrastructure communication. Based on federated learning, it filters and aggregates the local vehicle-side group neural network parameters and shared neural network parameters to achieve multi-agent group optimization.
[0035] The technical solution of the multi-agent vehicle-road-cloud integrated collaborative decision-making method based on federated reinforcement learning of this invention includes the following steps:
[0036] Step 1: Build a cloud platform, providing three cloud control applications—connected vehicle empowerment, traffic management and control, and traffic data empowerment—through three levels of cloud control. Obtain comprehensive road network and real-time traffic information, logistics monitoring information, high-precision map information, high-precision positioning and navigation information, and real-time weather information through relevant cloud support platforms. The cloud control platform, in conjunction with relevant support platforms, provides high-level information such as traffic scheduling and management to roadside operations.
[0037] Step 2: Build a roadside platform. The roadside information processing module processes the bird's-eye view image input into a semantic bird's-eye view map, and leverages the advantages of the roadside to generate a Gaussian safety field. The semantic bird's-eye view map serves as input to the vehicle-side perception processing module, while the Gaussian safety field participates in generating the fusion reward function in the reinforcement learning module. The roadside group coordination and decision-making module obtains the vehicle-side neural network parameters from the intelligent connected vehicle group. Based on federated learning, it filters and aggregates local vehicle-side group neural network parameters and shared neural network parameters to build a multi-agent group optimization architecture.
[0038] Step 3: Build the vehicle-side platform. A single roadside unit contains multiple vehicle-side groups, each consisting of N intelligent connected vehicles, sharing information through V2V vehicle-to-vehicle communication. The intelligent connected vehicle comprises a perception processing module, a trajectory prediction module, a reinforcement learning module, and an intelligent chassis coupling system. The perception processing module matches and segments the semantic bird's-eye view provided by the roadside based on the vehicle's location, fusing the vehicle's dynamic information (sensor data) and the roadside static information into a state matrix, which serves as input to the reinforcement learning module. The trajectory prediction module performs multi-time-step trajectory prediction based on the semantic bird's-eye view provided by the roadside, and the prediction results serve as input to the reinforcement learning module. The reinforcement learning module models the planning and control process as a Markov decision process, using the output of the perception processing module as reinforcement learning input, and combining it with the fusion reward function generated by the roadside safety field. The resulting planning and control output is interactively generated in the CARLA simulator. Then, based on the reinforcement learning planning and control output and the desired path, the intelligent chassis coupling system outputs the secondary planning control quantity for each intelligent connected vehicle. Finally, a distributed controller controls the corresponding intelligent chassis subsystem.
[0039] Step 4: After the intelligent chassis coupling system of a single intelligent connected vehicle outputs the secondary planning control quantity, it interacts with the CARLA simulator to generate experience samples, and trains the neural network parameters of the reinforcement learning module based on the experience samples.
[0040] Step 5: The roadside group coordination and decision-making module obtains the vehicle-side neural network parameters of the intelligent connected vehicles in the vehicle-side group through V2I vehicle-road communication, and realizes multi-agent group optimization by filtering, aggregating local vehicle-side group neural network parameters and shared neural network parameters based on federated learning.
[0041] Preferably, in step 1, the three cloud control levels are edge cloud, regional cloud, and central cloud. Based on the real-time traffic dynamics collected by the edge cloud, the cloud control platform provides enabling services to meet the driving safety requirements of connected vehicles and improve traffic efficiency. Through the information provided by the edge cloud and the support platform, the regional cloud provides services such as traffic situation awareness and assessment, traffic planning, order management, and transportation management, achieving benefits such as reducing traffic accidents, alleviating road congestion, and improving road network efficiency. Furthermore, through the information provided by the regional cloud and the support platform, the central cloud, empowered by traffic big data, provides big data analysis services and business support to universities, research institutions, mobility service providers, and automobile manufacturers.
[0042] Preferably, in step 2, the Gaussian safety field is established by the following equation: φ=b x / c y =l v / w v
[0043] Among them, S sta C represents the static safety field strength. a The static safety field strength coefficient is represented by x0 and y0, which represent the coordinates of the static risk center O(x0, y0). x and c y This represents the appearance coefficient of the intelligent connected vehicle, where φ is the aspect ratio of the intelligent connected vehicle, and l v Indicates the length of the vehicle, w v Indicates the width of the vehicle.
[0044] When the intelligent connected vehicle moves, the risk center O(x0,y0) of the Gaussian safety field will shift to a new risk center O′(x′0,y′0) as the vehicle moves:
[0045] Where, k v Indicates the moving adjustment factor, and The sign is related to the direction of motion; β represents the transfer vector of the connected vehicle. The angle between the vehicle and the coordinate axes in the Cartesian coordinate system. A virtual vehicle of length l′ is formed in the dynamic safety field under the influence of risk center transfer. v Width is w′ v S dyn The dynamic safety field strength is represented by the new aspect ratio φ′=b′. x / c′ y =l′ v / w′ v .
[0046] Preferably, in step 2, the neural network screening process uses driving safety and driving aggression based on Gaussian safety fields as network screening indicators. A rule-based method is used as the screening basis, achieving rule-driven integration with data, and the interpretability of the algorithm framework is enhanced through interpretable rules.
[0047] The actor network employs the following aggregation process:
[0048] in, This indicates that the parameters θ of the intelligent connected vehicle itself are... μ,i And another intelligent connected vehicle parameter θ μ,i′ New network parameters obtained through aggregation. Safety The reward function related to the safety of intelligent connected vehicles is calculated by the roadside driving safety field from two aspects: driving safety and driving aggression. r Safety =r Risk +r Agg
[0049] Among them, R i,j (t) represents the driving risk posed by intelligent connected vehicle j to intelligent connected vehicle i. Let k represent the field strength of intelligent connected vehicle j with respect to intelligent connected vehicle i. c Indicates the risk perception coefficient. Let θ represent the speed of the intelligent connected vehicle j at time t. i,j (t) represents the angle between intelligent connected vehicle i and intelligent connected vehicle j at time t, r Risk Let f represent the reward function related to driving risk. Risk (ξ) represents the driving risk score, R thr τ represents the risk threshold. rc R represents the duration exceeding the risk threshold. j,i (t′) represents the driving risk posed by intelligent connected vehicle i to intelligent connected vehicle j. Let i represent the field strength of intelligent connected vehicle i for intelligent connected vehicle j. Let θ represent the speed of the intelligent connected vehicle i at time t′. j,i (t′) represents the angle between intelligent connected vehicle j and intelligent connected vehicle i at time t′, r Agg Let f represent the reward function related to driving aggression. Agg (ξ) represents the integral of driving aggression.
[0050] The critic network employs the following aggregation process:
[0051] in, This indicates that the parameters θ of the intelligent connected vehicle itself are... l,i And another intelligent connected vehicle parameter θ l,i′ The new network parameters obtained through aggregation The value output of the critic network for a state-action pair (s, a).
[0052] Preferably, in step 3, the intelligent subsystem includes, but is not limited to, a steering subsystem and a DYC subsystem for calculating the lateral control of the intelligent connected vehicle, a suspension subsystem for calculating the vertical control of the intelligent connected vehicle, and a drive subsystem for calculating the longitudinal control of the intelligent connected vehicle. The steering subsystem and DYC subsystem calculate the yaw moment of the intelligent connected vehicle based on the steering wheel control output by the reinforcement learning module and the pre-aiming error between the actual path and the desired path. The drive subsystem obtains the desired target speed and calculates the four-wheel drive / braking torque of the intelligent connected vehicle based on the throttle control and the current speed of the intelligent connected vehicle.
[0053] Preferably, in step 4, the experience sample consists of tuples (s) t ,a t ,r t ,s t+1 ) describes, where s t Corresponding state variable matrix i RL ∈[0,1] W×H×C Where W×H=56 represents the matrix size, actually covering a physical area of approximately 14m×14m, and C=4 represents the number of channels. t The corresponding reinforcement learning actor network generates a control output 'a' based on the state variables. t1 With a t2 The output is mapped to the range [-1, 1] using an activation function, where a t1 Indicates the amount of steering wheel control, a t2 It is divided into two parts: [-1, 0] and [0, 1], representing the brake and accelerator control values, respectively. t Corresponding to the expected path-related reward function and the security-related reward function, where s t+1 This represents the state variable matrix for the next frame.
[0054] Preferably, in step 5, the vehicle-to-infrastructure communication only transmits network parameters, not training samples, thus ensuring agent privacy while reducing communication and computational overhead. The neural network used for parameter uploading and downloading is a reinforcement learning actor-critic network, where the actor network outputs the regulatory policy and the critic network evaluates the policy.
[0055] The beneficial effects of this invention are:
[0056] (1) This invention provides a multi-agent vehicle-road-cloud integrated collaborative decision-making architecture based on federated reinforcement learning. It adopts a multi-agent federated reinforcement learning decision-making framework with embedded vehicle dynamics characteristics, which solves the problem of deep information fusion between intelligent transportation system (ITS) and intelligent vehicle (IV), and realizes autonomous driving based on deep decision-making collaboration between vehicles and traffic.
[0057] (2) Based on the roadside perspective, a semantic matrix is generated as the input for vehicle-side reinforcement learning, and a global and local trajectory planning for the vehicle under the guidance of the roadside is constructed. The advantages of the roadside are used to realize the modeling of the driving safety field, and a fusion reward function used by vehicle-side reinforcement learning is constructed to realize the comprehensive consideration of vehicle-side safety and comfort. Based on the roadside federated learning architecture, the vehicle-side neural network parameters are transmitted through V2I communication, which solves the problem of information asymmetry between vehicles and roads caused by privacy awareness.
[0058] (3) For different environmental sample distributions, a local optimal strategy for the current environment is selected based on the neural network screening process, and a globally shared model that benefits from different environments is synthesized to achieve a balance between sample efficiency and model robustness. Attached Figure Description
[0059] Figure 1. Schematic diagram of the vehicle-road-cloud integrated collaborative decision-making architecture proposed in this invention;
[0060] Figure 2. Schematic diagram of vehicle-road-cloud communication proposed in this invention;
[0061] Figure 3. Schematic diagram of the group coordination decision-making proposed in this invention;
[0062] Figure 4. Schematic diagram of the neural network screening used in this invention; Detailed Implementation
[0063] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings, but the content of the present invention is not limited thereto.
[0064] This invention provides a multi-agent vehicle-road-cloud integrated collaborative decision-making and control architecture based on federated reinforcement learning, as shown in Figure 1. It can realize multi-agent autonomous driving in complex driving scenarios based on integrated vehicle-road-cloud collaborative decision-making and control. A schematic diagram of vehicle-road-cloud communication is shown in Figure 2. The specific steps include:
[0065] (1) Build a cloud platform to provide three cloud control applications: connected vehicle empowerment, traffic management and control, and traffic data empowerment, through three cloud control levels of the cloud platform. Obtain complete traffic network and real-time traffic information, logistics supervision information, high-precision map information, high-precision positioning and navigation information, real-time weather information, etc. through relevant cloud support platforms. Provide high-level information such as traffic scheduling and management to the roadside through the cloud control platform in combination with relevant support platforms.
[0066] (2) A roadside platform is built. The roadside information processing module processes the bird's-eye view image input into a semantic bird's-eye view map, and utilizes the advantages of the roadside to generate a Gaussian safety field. The semantic bird's-eye view map serves as the input to the vehicle-side perception processing module, while the Gaussian safety field participates in generating the fusion reward function in the vehicle-side reinforcement learning module. The roadside group coordination decision module obtains the vehicle-side neural network parameters of the intelligent connected vehicles in the vehicle group, as shown in Figure 3. Based on federated learning, the local vehicle-side group neural network parameters and shared neural network parameters are selected and aggregated to build a multi-agent group optimization architecture. The Gaussian safety field is established through the following equation: φ=b x / c y =l v / w v
[0067] Among them, S sta C represents the static safety field strength. a The static safety field strength coefficient is represented by x0 and y0, which represent the coordinates of the static risk center O(x0, y0). x and c y This represents the appearance coefficient of the intelligent connected vehicle, where φ is the aspect ratio of the intelligent connected vehicle, and l v Indicates the length of the vehicle, w v Indicates the width of the vehicle.
[0068] When the intelligent connected vehicle moves, the risk center O(x0,y0) of the Gaussian safety field will shift to a new risk center O′(x′0,y′0) as the vehicle moves:
[0069] Where, k v Indicates the moving adjustment factor, and The sign is related to the direction of motion; β represents the transfer vector of the connected vehicle. The angle between the vehicle and the coordinate axes in the Cartesian coordinate system. A virtual vehicle of length l′ is formed in the dynamic safety field under the influence of risk center transfer. v Width is w′ v S dyn The dynamic safety field strength is represented by the new aspect ratio φ′=b′. x / c′ y =l′ v / w′ v .
[0070] (3) A vehicle-side platform is established, with multiple vehicle-side groups under a single roadside unit. Each vehicle-side group consists of N intelligent connected vehicles, and information sharing is achieved through V2V vehicle-to-vehicle communication. The intelligent connected vehicle consists of a perception processing module, a trajectory prediction module, a reinforcement learning module, and an intelligent chassis coupling system. The perception processing module matches and segments the semantic bird's-eye view provided by the roadside based on the vehicle's positioning, and fuses the vehicle's dynamic information (sensor data) and the roadside static information into a state matrix, which is used as part of the reinforcement learning input. The trajectory prediction module performs multi-time-step trajectory prediction based on the semantic bird's-eye view provided by the roadside, and the prediction result is used as part of the reinforcement learning module input. The reinforcement learning module models the planning and control process as a Markov decision process, uses the output of the perception processing module as the reinforcement learning input, and combines it with the roadside Gaussian safety field to generate a fusion reward function, which is interactively generated in the CARLA simulator to produce the planning and control output. Then, based on the reinforcement learning planning and control output and the desired path, the intelligent chassis coupling system outputs the secondary planning control quantity of a single intelligent connected vehicle. Finally, the corresponding intelligent chassis subsystem is controlled through a distributed controller. The intelligent subsystem includes, but is not limited to, the steering subsystem, the DYC subsystem, the suspension subsystem, and the drive subsystem. These intelligent subsystems are coupled in both state and cost. The steering and DYC subsystems calculate the lateral control variables of the intelligent connected vehicle, the suspension subsystem calculates the vertical control variables, and the drive subsystem calculates the longitudinal control variables. The steering and DYC subsystems calculate the yaw moment of the intelligent connected vehicle based on the steering wheel control variables output by the reinforcement learning module and the aiming error between the actual and desired paths. The drive subsystem obtains the desired target speed and calculates the four-wheel drive / braking torque of the intelligent connected vehicle based on the throttle control variables and the current speed of the intelligent connected vehicle. First, the driving force equation of the intelligent connected vehicle is constructed based on PID control.
[0071] Among them, F x (t) represents the driving force of the intelligent connected vehicle at time t, K p Let K represent the proportional gain, e(t) represent the velocity error at time t, and K represent the velocity error at time t. i K represents the integral gain. dLet represent the differential gain. Then consider the following optimization problem:
[0072] Where F xij F represents the longitudinal force on the four wheels of an intelligent connected vehicle. yij F represents the lateral force on the four wheels of an intelligent connected vehicle. zfl F zfr F zfl F zrr This represents the vertical force on all four wheels of the intelligent connected vehicle, where fl represents the left front wheel, fr represents the right front wheel, rl represents the left rear wheel, rr represents the right rear wheel, and μ represents the adjustment factor. Since the steering subsystem and DYC subsystem have already calculated the four-wheel steering angles, i.e., F... yij Since ,ij=fl,fr,rl,rr are constants, the above equation simplifies to:
[0073] subject to
[0074] Where m represents the mass of the intelligent connected vehicle, a x d represents the driving acceleration, and M represents the wheelbase of the intelligent connected vehicle. z R represents the additional yaw moment. w T represents the wheel radius of a smart connected vehicle. min T represents the minimum driving torque / braking torque. max This represents the maximum driving torque. The above three equations respectively indicate that the total driving force satisfies the constraints of the drive subsystem, the additional yaw moment satisfies the constraints of the steering subsystem and DYC subsystem, and the driving force satisfies the actuator constraints. Finally, by solving the optimization problem, the four-wheel drive / braking torque is obtained: T ij =F xij R w ,i=fl,fr,rl,rr
[0075] (4) After the intelligent chassis coupling system of a single intelligent connected vehicle outputs the secondary programming control quantity, it interacts with the CARLA simulator to generate experience samples, and trains the neural network parameters of the reinforcement learning module based on the experience samples. The experience samples consist of tuples (s t ,a t ,r t ,s t+1 ) describes, where s t Corresponding state variable matrix i RL ∈[0,1] W×H×C Where W×H=56 represents the matrix size, actually covering a physical area of approximately 14m×14m, and X=4 represents the number of channels. tThe corresponding reinforcement learning actor network generates a control output 'a' based on the state variables. t1 With a t2 The output is mapped to the range [-1, 1] using an activation function, where a t1 Indicates the amount of steering wheel control, a t2 It is divided into two parts: [-1, 0] and [0, 1], representing the brake and accelerator control values, respectively. t Corresponding to the expected path-related reward function and the security-related reward function, where s t+1 This represents the state variable matrix for the next frame.
[0076] (5) The roadside group coordination decision-making module obtains the vehicle-side neural network parameters of intelligent connected vehicles in the vehicle-side group through V2I vehicle-road communication, and filters and aggregates the local vehicle-side group neural network parameters and shared neural network parameters based on federated learning. The neural network filtering process uses driving safety and driving aggression based on Gaussian risk fields as network filtering indicators, as shown in Figure 4. A rule-based method is used as the basis for filtering to achieve the fusion of rules and data, and the interpretability of the algorithm framework is enhanced through interpretable rules. The actor network adopts the following aggregation process:
[0077] in, This indicates that the parameters θ of the intelligent connected vehicle itself are... μ,i And another intelligent connected vehicle parameter θ μ,i′ New network parameters obtained through aggregation. Safety The safety-related reward function for intelligent connected vehicles is calculated from both driving safety and driving aggression aspects by the roadside Gaussian safety field. r Safety =r Risk +r Agg
[0078] Among them, R i,j (t) represents the driving risk posed by intelligent connected vehicle j to intelligent connected vehicle i. Let k represent the field strength of intelligent connected vehicle j with respect to intelligent connected vehicle i. c Indicates the risk perception coefficient. Let θ represent the speed of the intelligent connected vehicle j at time t. i,j (t) represents the angle between intelligent connected vehicle i and intelligent connected vehicle j at time t, r Risk Let f represent the reward function related to driving risk. Risk (ν) represents the driving risk score, R thrτ represents the risk threshold. rc R represents the duration exceeding the risk threshold. j,i (t′) represents the driving risk posed by intelligent connected vehicle i to intelligent connected vehicle j. Let i represent the field strength of intelligent connected vehicle i for intelligent connected vehicle j. Let θ represent the speed of the intelligent connected vehicle i at time t′. j,i (t′) represents the angle between intelligent connected vehicle j and intelligent connected vehicle i at time t′, r Agg Let f represent the reward function related to driving aggression. Agg (ξ) represents the integral of driving aggression. The critic network uses the following aggregation process:
[0079] in, This indicates that the parameters θ of the intelligent connected vehicle itself are... l,i And another intelligent connected vehicle parameter θ l,i′ The new network parameters obtained through aggregation The value output of the critic network for a state-action pair (s, a).
[0080] In summary, this invention proposes a multi-agent vehicle-road-cloud integrated collaborative decision-making architecture based on federated reinforcement learning. Employing a multi-agent federated reinforcement learning framework with embedded vehicle dynamics characteristics, it solves the problem of deep information fusion between Intelligent Transportation Systems (ITS) and Intelligent Vehicles (IV), achieving autonomous driving based on deep collaborative decision-making between vehicles and traffic. A semantic matrix generated from the roadside perspective is used as input for vehicle-side reinforcement learning, constructing roadside-guided global and local trajectory planning for the vehicle. Utilizing the advantages of the roadside, driving safety field modeling is achieved, and a fusion reward function used by vehicle-side reinforcement learning is constructed, realizing a comprehensive consideration of vehicle safety and comfort. Based on the roadside federated learning architecture, vehicle-side neural network parameters are transmitted via V2I communication, solving the problem of information asymmetry between vehicles and roads caused by privacy concerns. For different environmental sample distributions, a locally optimal strategy for the current environment is selected based on a neural network screening process, and a globally shared model benefiting from different environments is synthesized, achieving a balance between sample efficiency and model robustness.
[0081] The detailed descriptions listed above are merely specific descriptions of feasible embodiments of the present invention, and are not intended to limit the scope of protection of the present invention. All equivalent methods or modifications that do not depart from the technology of the present invention should be included within the scope of protection of the present invention.
Claims
1. A multi-agent vehicle-road-cloud integrated collaborative decision-making and control architecture system based on federated reinforcement learning, characterized in that: include: Cloud, road, and vehicle; The cloud platform includes a cloud control platform and a support platform. The cloud control platform includes three cloud control applications: connected vehicle empowerment, traffic management and control, and traffic data empowerment. The support platform includes a traffic management platform that provides a complete road network and real-time traffic information, a logistics platform that provides logistics supervision information, a map platform that provides high-precision maps, a positioning platform that provides high-precision positioning and navigation, and a meteorological platform that provides real-time weather information. The roadside unit is used to coordinate a group of vehicles based on different cloud control applications. The roadside unit includes multiple roadside units, each containing an information processing module and a group coordination decision-making module. The information processing module processes the bird's-eye view image input into a semantic bird's-eye view map and generates a Gaussian safe field. The semantic bird's-eye view map serves as the input to the vehicle-side perception processing module, and the Gaussian safe field is used to generate the fusion reward function in the vehicle-side reinforcement learning module. The group coordination decision-making module obtains the neural network parameters of the intelligent connected vehicle reinforcement learning module within the vehicle group through V2I vehicle-road communication, and achieves multi-agent group optimization by filtering, aggregating local vehicle-side group neural network parameters, and sharing neural network parameters based on federated learning. The vehicle-side unit is used to output planning control quantities based on roadside input and intelligent connected vehicle sensor information. A single roadside unit contains multiple vehicle-side groups, each consisting of N intelligent connected vehicles, sharing information through V2V vehicle-to-vehicle communication. The vehicle-side unit includes a perception processing module, a trajectory prediction module, a reinforcement learning module, and an intelligent chassis coupling system. The perception processing module matches and segments the semantic bird's-eye view provided by the roadside unit based on the vehicle's location, fusing the vehicle's dynamic information and the roadside static information into a state matrix, which serves as input to the reinforcement learning module's neural network. The trajectory prediction module uses the semantic bird's-eye view provided by the roadside unit to predict trajectories at multiple time steps; the prediction result is used as part of the reinforcement learning module's input. The reinforcement learning module generates a fusion reward function by combining the inputs from the perception processing module, the trajectory prediction module, and the roadside Gaussian safety field, interactively producing the planning control quantity output. The intelligent chassis coupling system is used by a single intelligent connected vehicle to output a secondary planning control quantity based on the planning control quantity output by the reinforcement learning module and the desired path, controlling the corresponding intelligent chassis subsystem through a distributed controller.
2. The multi-agent vehicle-road-cloud integrated collaborative decision-making and control architecture system based on federated reinforcement learning according to claim 1, characterized in that, The cloud control platform includes an edge cloud, a regional cloud, and a central cloud. The edge cloud collects real-time traffic data and provides services that meet the driving safety requirements of connected vehicles and improve traffic efficiency. The regional cloud provides services such as traffic situation awareness and assessment, traffic planning, order management, and transportation management. The central cloud is empowered by traffic big data and provides big data analysis services and business support.
3. The multi-agent vehicle-road-cloud integrated collaborative decision-making architecture system based on federated reinforcement learning according to claim 1, characterized in that, The Gaussian safety field is established by the following equation: φ=b x / c y =l v / w v Among them, S sta C represents the static safety field strength. a The static safety field strength coefficient is represented by x0 and y0, which represent the coordinates of the static risk center O(x0, y0). x and c y This represents the appearance coefficient of the intelligent connected vehicle, where φ is the aspect ratio of the intelligent connected vehicle, and l v Indicates the length of the vehicle, w v Indicates vehicle width; When the intelligent connected vehicle moves, the risk center O(x0,y0) of the Gaussian safety field will shift to a new risk center O′(x′0,y0′) as the vehicle moves: Where, k v Indicates the moving adjustment factor, and The sign is related to the direction of motion; β represents the transfer vector of the connected vehicle. The angle between the vehicle and the coordinate axes in the Cartesian coordinate system forms a virtual vehicle of length l′ under the influence of the risk center transfer in the dynamic safety field. v Width is w′ v S dyn The dynamic safety field strength is represented by the new aspect ratio φ′=b′. x / c′ y =l′ v / w′ v .
4. The multi-agent vehicle-road-cloud integrated collaborative decision-making architecture system based on federated reinforcement learning according to claim 1, characterized in that... The neural network at the vehicle end is a reinforcement learning actor-critic network, in which the actor network outputs the control policy and the critic network evaluates the policy. The Actor network employs the following aggregation process: in, This indicates that the parameters θ of the intelligent connected vehicle itself are... μ,i And another intelligent connected vehicle parameter θ μ,i′ The newly generated network parameters, r Safety The safety-related reward function for intelligent connected vehicles is calculated from both driving safety and driving aggression aspects by the roadside Gaussian safety field. r Safety =r Risk +r Agg Among them, R i,j (t) represents the driving risk posed by intelligent connected vehicle j to intelligent connected vehicle i. Let k represent the field strength of intelligent connected vehicle j with respect to intelligent connected vehicle i. c Indicates the risk perception coefficient. Let θ represent the speed of the intelligent connected vehicle j at time t. i,j (t) represents the angle between intelligent connected vehicle i and intelligent connected vehicle j at time t, r Risk Let f represent the reward function related to driving risk. Risk (ξ) represents the driving risk score, R thr τ represents the risk threshold. rc R represents the duration exceeding the risk threshold. j,i (t′) represents the driving risk posed by intelligent connected vehicle i to intelligent connected vehicle j. Let i represent the field strength of intelligent connected vehicle i for intelligent connected vehicle j. Let θ represent the speed of the intelligent connected vehicle i at time t′. j,i (t′) represents the angle between intelligent connected vehicle j and intelligent connected vehicle i at time t′, r Agg Let f represent the reward function related to driving aggression. Agg (ξ) represents the integral of driving aggression; The critic network employs the following aggregation process: in, This indicates that the parameters θ of the intelligent connected vehicle itself are... l,i And another intelligent connected vehicle parameter θ l,i′ The new network parameters obtained through aggregation This represents the value of the Critic network's output for the state-action pair (s, a).
5. The multi-agent vehicle-road-cloud integrated collaborative decision-making and control architecture system based on federated reinforcement learning according to claim 1, characterized in that, The dynamic information includes vehicle speed and location information, the roadside static information includes road, desired path, and lane information, and the state variable matrix i RL ∈[0,1] W×H×C Where W×H=56 represents the matrix size, actually covering a physical area of approximately 14m×14m, and C=4 represents the number of channels.
6. The multi-agent vehicle-road-cloud integrated collaborative decision-making and control architecture system based on federated reinforcement learning according to claim 1, characterized in that, The interaction of the reinforcement learning module generates the planning control output, which is produced by the reinforcement learning actor network based on the fused state matrix. t1 With a t2 The output is mapped to the range [-1, 1] using an activation function, where a t1 Indicates the amount of steering wheel control, a t2 It is divided into two parts: [-1, 0] and [0, 1], which represent the brake and accelerator control values, respectively.
7. The multi-agent vehicle-road-cloud integrated collaborative decision-making architecture system based on federated reinforcement learning according to claim 1, characterized in that, The intelligent subsystem includes a steering subsystem, a DYC subsystem, a suspension subsystem, and a drive subsystem. The steering subsystem and DYC subsystem calculate the lateral control parameters of the intelligent connected vehicle, the suspension subsystem calculates the vertical control parameters of the intelligent connected vehicle, and the drive subsystem calculates the longitudinal control parameters of the intelligent connected vehicle. The steering subsystem and DYC subsystem calculate the yaw moment of the intelligent connected vehicle based on the steering wheel control parameters output by the reinforcement learning module and the pre-aiming error between the actual path and the desired path. The drive subsystem obtains the desired target speed and calculates the four-wheel drive or braking torque of the intelligent connected vehicle based on the throttle control quantity and the current speed of the intelligent connected vehicle.
8. The multi-agent vehicle-road-cloud integrated collaborative decision-making architecture system based on federated reinforcement learning according to claim 7, characterized in that, The specific control of the intelligent subsystem is as follows: First, construct the driving force equation for intelligent connected vehicles based on PID: Among them, F x (t) represents the driving force of the intelligent connected vehicle at time t, K p Let e(t) represent the proportional gain, e(t) represent the velocity error at time t, and k represent the velocity error at time t. i K represents the integral gain. d Represents differential gain. Then consider the following optimization problem: Where F xij F represents the longitudinal force on the four wheels of an intelligent connected vehicle. yij F represents the lateral force on the four wheels of an intelligent connected vehicle. zfl This represents the vertical force on all four wheels of the intelligent connected vehicle, where fl represents the left front wheel, fr represents the right front wheel, rl represents the left rear wheel, rr represents the right rear wheel, and μ represents the adjustment factor. Since the steering subsystem and DYC subsystem have already calculated the four-wheel steering angles, i.e., F... yij Since ,ij=fl,fr,rl,rr are constants, the above equation simplifies to: subject to The three formulas above respectively indicate that the total driving force satisfies the constraints of the drive subsystem, the additional yaw moment satisfies the constraints of the steering subsystem and DYC subsystem, and the driving force satisfies the actuator constraints; where m represents the mass of the intelligent connected vehicle, a x d represents the driving acceleration, and M represents the wheelbase of the intelligent connected vehicle. z R represents the additional yaw moment. w T represents the wheel radius of a smart connected vehicle. min T represents the minimum driving torque or braking torque. max Indicates the maximum driving torque; Finally, by solving the optimization problem, the four-wheel drive or braking torque is obtained: T ij =F xij R w ,i=fl,fr,rl,rr。 9. The multi-agent vehicle-road-cloud integrated collaborative decision-making architecture system based on federated reinforcement learning according to claim 1 or 8, characterized in that, After the intelligent chassis coupling system outputs the secondary programming control quantity, it also includes using the secondary programming control quantity to generate experience samples, and training the neural network parameters of the reinforcement learning module based on the experience samples; then the roadside group coordination decision module obtains the intelligent connected vehicle neural network parameters in the vehicle group through V2I vehicle-road communication, and achieves multi-agent group optimization based on federated learning to filter, aggregate local vehicle group neural network parameters and shared neural network parameters.
10. A multi-agent vehicle-road-cloud integrated collaborative decision-making and control method based on federated reinforcement learning, characterized in that... Including the following: Step 1: Build a cloud platform. Provide three cloud control applications through the three cloud control levels of the cloud control platform: connected vehicle empowerment, traffic management and control, and traffic data empowerment. Obtain complete traffic network and real-time traffic information, logistics supervision information, high-precision map information, high-precision positioning and navigation information, and real-time weather information through the cloud support platform. Provide traffic scheduling and management information to the roadside through the cloud control platform in combination with the support platform. Step 2: Build a roadside platform. The roadside information processing module processes the bird's-eye view image input into a semantic bird's-eye view map and generates a Gaussian safety field. The semantic bird's-eye view map serves as the input to the vehicle-side perception processing module, while the Gaussian safety field participates in generating the fusion reward function in the reinforcement learning module. Step 3: Build a vehicle-side platform. A single roadside unit contains multiple vehicle-side groups, each consisting of N intelligent connected vehicles. Information sharing is achieved through V2V vehicle-to-vehicle communication. Each intelligent connected vehicle includes a perception processing module, a trajectory prediction module, a reinforcement learning module, and an intelligent chassis coupling system. The perception processing module matches and segments the semantic bird's-eye view provided by the roadside based on the vehicle's location, fusing the vehicle's dynamic information and the roadside static information into a state matrix, which serves as part of the input to the reinforcement learning module. The trajectory prediction module predicts multi-time-step trajectories based on the semantic bird's-eye view provided by the roadside. The prediction results are then used as part of the input to the reinforcement learning module. The reinforcement learning module models the planning and control process as a Markov decision process, and the output of the perception processing module is used as the reinforcement learning input. The input is combined with the Gaussian safety field at the road end to generate a fusion reward function. The control output is generated interactively in the CARLA simulator. Based on the reinforcement learning control output and the desired path, the secondary planning control quantity of a single intelligent connected vehicle is output through the intelligent chassis coupling system. The corresponding intelligent chassis subsystem is controlled through the distributed controller. Step 4: After the intelligent chassis coupling system of a single intelligent connected vehicle outputs the secondary planning control quantity, it interacts with the CARLA simulator to generate experience samples, and trains the neural network parameters of the reinforcement learning module based on the experience samples. Step 5: The roadside group coordination decision module obtains the vehicle-side neural network parameters of the intelligent connected vehicles in the vehicle group through V2I vehicle-road communication, and achieves multi-agent group optimization based on federated learning to filter, aggregate local vehicle-side group neural network parameters and shared neural network parameters; In step 1, the three cloud control levels correspond to edge cloud, regional cloud, and central cloud, respectively. Based on the real-time traffic dynamics collected by the edge cloud, the cloud control platform provides services that enable connected vehicles to improve driving safety and traffic efficiency. Through the edge cloud combined with the information provided by the support platform, the regional cloud provides traffic situation awareness and assessment, traffic planning, order management, and transportation management services. Through the regional cloud combined with the information provided by the support platform, the central cloud provides big data analysis services and business support based on traffic big data. In step 2, the Gaussian safety field is established using the following equation: φ=b x / c y =l v / w v Among them, S sta C represents the static safety field strength. a The static safety field strength coefficient is represented by x0 and y0, which represent the coordinates of the static risk center O(x0, y0). x and c y This represents the appearance coefficient of the intelligent connected vehicle, where φ is the aspect ratio of the intelligent connected vehicle, and l v Indicates the length of the vehicle, w v Indicates vehicle width; When the intelligent connected vehicle moves, the risk center O(x0,y0) of the Gaussian safety field will shift to a new risk center O′(x′0,y′0) as the vehicle moves: Where, k v Indicates the moving adjustment factor, and The sign is related to the direction of motion; β represents... Connected vehicle transfer vector The angle between the vehicle and the coordinate axes in the Cartesian coordinate system forms a virtual vehicle of length l′ under the influence of the risk center transfer in the dynamic safety field. v Width is w′ v S dyn The dynamic safety field strength is represented by the new aspect ratio φ′=b′. x / c′ y =l′ v / w′ v ; In step 2, the neural network screening process uses driving safety and driving aggression based on Gaussian safety fields as network screening indicators, and uses rule-based methods as the screening basis to achieve the fusion of rules and data, and enhances the interpretability of the algorithm framework through interpretable rules. The actor network employs the following aggregation process: in, This indicates that the parameters θ of the intelligent connected vehicle itself are... μ,i And another intelligent connected vehicle parameter θ μ,i′ The newly generated network parameters, r Safety The reward function related to the safety of intelligent connected vehicles is calculated by the roadside driving safety field from two aspects: driving safety and driving aggression. r Safety =r Risk +r Agg Among them, R i,j (t) represents the driving risk posed by intelligent connected vehicle j to intelligent connected vehicle i. Let k represent the field strength of intelligent connected vehicle j with respect to intelligent connected vehicle i. c Indicates the risk perception coefficient. Let θ represent the speed of the intelligent connected vehicle j at time t. i,j (t) represents the angle between intelligent connected vehicle i and intelligent connected vehicle j at time t, r Risk Let f represent the reward function related to driving risk. Risk (ξ) represents the driving risk score, R thr τ represents the risk threshold. rc R represents the duration exceeding the risk threshold. j,i (t′) represents the driving risk posed by intelligent connected vehicle i to intelligent connected vehicle j. Let i represent the field strength of intelligent connected vehicle i for intelligent connected vehicle j. Indicating intelligent connected vehicles i The velocity at time t′, θ j,i (t′) represents the angle between intelligent connected vehicle j and intelligent connected vehicle i at time t′, r Agg Let f represent the reward function related to driving aggression. Agg (ξ) represents the integral of driving aggression; The critic network employs the following aggregation process: in, This indicates that the parameters θ of the intelligent connected vehicle itself are... l,i And another intelligent connected vehicle parameter θ l,i′ The new network parameters obtained through aggregation The value output of the critic network for a state-action pair (s, a); In step 3, the intelligent subsystem includes a steering subsystem and a DYC subsystem for calculating the lateral control of the intelligent connected vehicle, a suspension subsystem for calculating the vertical control of the intelligent connected vehicle, and a drive subsystem for calculating the longitudinal control of the intelligent connected vehicle. The steering subsystem and the DYC subsystem calculate the yaw moment of the intelligent connected vehicle based on the steering wheel control output by the reinforcement learning module and the pre-aiming error between the actual path and the desired path. The drive subsystem obtains the desired target speed and calculates the four-wheel drive or braking torque of the intelligent connected vehicle based on the throttle control and the current speed of the intelligent connected vehicle. In step 4, the empirical samples consist of tuples (s) t ,a t ,r t ,s t+1 ) describes, where s t Corresponding state variable matrix i RL ∈[0,1] W×H×C Where W×H=56 represents the matrix size, actually covering a physical area of approximately 14m×14m, C=4 represents the number of channels, and a t The corresponding reinforcement learning actor network generates a control output 'a' based on the state variables. t1 With a t2 The output is mapped to the range [-1, 1] using an activation function, where a t1 Indicates the amount of steering wheel control, a t2 It is divided into two parts, [-1, 0] and [0, 1], representing the brake and throttle control values respectively, r t Corresponding to the expected path-related reward function and the security-related reward function, where s t+1 This represents the state variable matrix for the next frame; In step 5, the vehicle-to-infrastructure communication only transmits neural network parameters. The neural network for uploading and downloading parameters is an actor-critic network, which is a reinforcement learning network. The actor network outputs the control policy and the critic network evaluates the policy.
Citation Information
Patent Citations
Vehicle-road cooperation system, analog simulation method, vehicle-mounted equipment and roadside equipment
CN113256976A
End-edge-cloud vehicle road collaborative fusion sensing architecture and construction method thereof
CN113743479A
Vehicle-road collaborative automatic driving decision-making method based on hierarchical reinforcement learning
CN115100866A
All-domain collaborative awareness and decision-making method and device based on vehicle and road cloud interface
CN115116216A
Multi-Level Collaborative Control System With Dual Neural Network Planning For Autonomous Vehicle Control In A Noisy Environment
US20200174471A1
Cited By
Double-four-foot series combination control method, system, medium and equipment
CN121043154A
Multi-agent cooperative scheduling method based on federal reinforcement learning and digital twinning
CN121543628A
Comprehensive energy system regulation and control method and system considering traffic logistics facilities
CN121563156A
Rollover early warning device for dangerous chemical transport vehicle, medium and electronic equipment
CN121617215A
Internet-of-vehicles-oriented communication, inductance and calculation fusion resource collaborative optimization method and system
CN121764662A