New energy automatic driving automobile coordination control system and method based on reinforcement learning

Through the reinforcement learning-based coordinated control system of new energy autonomous driving vehicles, the coordinated control of the suspension and drive systems is adjusted in real time, solving the problem of independent control of the suspension and drive systems, achieving a dynamic balance between energy saving, safety and comfort, and improving energy recovery efficiency and vehicle stability.

CN120646011APending Publication Date: 2025-09-16JIANGSU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510737761.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

In existing technologies, the independent control of the suspension and drive systems cannot be adjusted in real time, resulting in energy recovery efficiency that does not fully match the vehicle driving conditions, and energy saving and safety cannot be taken into account at the same time. The autonomous driving visual recognition technology does not fully identify road defects, resulting in energy waste.

Method used

A coordinated control system for new energy autonomous driving vehicles based on reinforcement learning is adopted, including a vehicle real-time status monitoring module, a visual recognition module, an intelligent ranging module, an intelligent suspension control module, and an intelligent drive control module. The system dynamically adjusts the coordinated control strategy of the suspension and drive systems through a deep reinforcement learning algorithm, combines visual recognition and intelligent ranging to detect road slope changes and disease characteristics in real time, and dynamically adjusts the suspension damping coefficient, multi-cavity aerodynamic spring stiffness, and the front and rear torque distribution of the drive system.

Benefits of technology

It has improved the comprehensive endurance under complex road conditions by 12% and the vibration energy recovery efficiency by 15%, achieving a dynamic balance between energy saving, safety and comfort under the premise of safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120646011A_ABST
    Figure CN120646011A_ABST
Patent Text Reader

Abstract

According to the reinforcement learning-based new energy automatic driving vehicle coordination control system and method provided by the invention, road slope change and disease characteristics (such as cracks, pavement subsidence and upheaval) are detected in real time through a visual identification and intelligent distance measurement module; the damping coefficient (500-5000Ns / m) of the suspension, the rigidity and the height of the multi-cavity aerodynamic spring, the front and rear torque distribution ratio (front-drive: rear-drive = 0: 100-100: 0) of the driving system and the energy recovery intensity (5-50kW) of the driving system are dynamically adjusted by combining a deep reinforcement learning model, and cooperative control (delay lt; 1 ms). The system can improve the comprehensive endurance by more than or equal to 12%, the vibration energy recovery efficiency is more than or equal to 15%, and the system is suitable for energy conservation and safety balance under complex road conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a new energy autonomous driving vehicle coordinated control system and method based on reinforcement learning, belonging to the technical field of new energy vehicles. Background Art

[0002] With the popularization of new energy vehicle technology and the development of autonomous driving technology, energy management of autonomous vehicles in complex road conditions has become a key challenge. Existing technologies have the following shortcomings:

[0003] The suspension and drive systems are independently controlled: they cannot be adjusted in real time according to road changes, resulting in energy recovery efficiency that does not fully match the vehicle driving conditions.

[0004] Energy saving and safety cannot be taken into account at the same time: Traditional suspension control methods often cannot take into account energy saving, safety and comfort at the same time, especially in scenarios with sudden obstacles, where it is difficult to take into account both braking energy recovery and vehicle stability.

[0005] Autonomous driving visual recognition technology does not fully recognize road conditions: Current autonomous driving visual recognition technology lacks specific countermeasures for the specific conditions of road defects and generally treats them as ordinary obstacles. As a result, the car cannot select appropriate active control strategies based on the defect conditions in advance, resulting in energy waste.

[0006] The present invention aims to solve the above problems and ensure that new energy autonomous driving vehicles achieve a dynamic balance between energy saving, safety and comfort under the premise of safety through multi-module collaborative control and deep reinforcement learning algorithms. Summary of the Invention

[0007] Purpose of the invention: In response to the deficiencies in the prior art, the present invention provides a new energy autonomous driving vehicle coordinated control system and method based on reinforcement learning to solve the problems mentioned in the above background technology.

[0008] Technical solution: A coordinated control system for new energy autonomous vehicles based on reinforcement learning, including a real-time vehicle status monitoring module, a visual recognition module, an intelligent ranging module, an intelligent suspension control module, an intelligent drive control module, a data transmission module, and a vehicle control module.

[0009] The vehicle real-time status monitoring module, visual recognition module and intelligent ranging module are respectively connected to the data transmission module signal, and the collected vehicle speed, suspension pitch angle, vehicle sideslip angle, sideslip angle, road disease degree, road surface roughness coefficient and distance between the vehicle and the obstacle are transmitted to the data transmission module;

[0010] The data transmission module is connected to the vehicle control module by signal, and the vehicle control module controls the intelligent suspension control module and the intelligent drive control module to perform actions in real time according to the collected data.

[0011] The vehicle real-time status monitoring module includes an inertial measurement unit (IMU) installed in the upper center of the vehicle chassis. It integrates a MEMS sensor, a three-axis gyroscope, a three-axis accelerometer, and a three-axis magnetometer. The Kalman filter algorithm is used to calculate the vehicle body pitch angle θ in real time. pitch 、Vehicle sideslip angle θ slide and the vehicle vertical vibration acceleration a vibration , data is transmitted to the vehicle control main processor via TSN Ethernet;

[0012] The visual recognition module includes a global shutter camera installed on the top of the front windshield, integrated with an FPGA image and processing unit and a GPU image depth processing unit, and realizes information transmission through the TSN Ethernet data transmission network;

[0013] The intelligent ranging module includes a laser radar and millimeter-wave radar installed in the center of the vehicle's front bumper, as well as a wheel speed sensor installed at the wheel bearing. Through the distance measurement main control processor, it integrates the laser radar point cloud, millimeter-wave radar data and wheel speed data in real time, and uses the Kalman filter algorithm to compensate for the sensor data offset caused by vehicle movement, thereby generating the relative positions of obstacles and road defects and the real-time speed of the vehicle.

[0014] The vehicle control module is equipped with a vehicle control main processor, runs a deep reinforcement learning model based on a hierarchical actor-critic network, and dynamically generates a coordinated control strategy for the suspension and drive system.

[0015] The intelligent suspension control module includes an electromagnetic active suspension actuator and a multi-cavity air spring; the electromagnetic active suspension actuator adjusts the damping fluid viscosity and flow rate in real time to achieve real-time adjustment of the suspension damping coefficient; the multi-cavity air spring adjusts the stiffness in real time;

[0016] The intelligent drive control module includes a permanent magnet synchronous motor, a silicon carbide inverter, a bidirectional DC-DC converter, and a supercapacitor module for transient energy buffering.

[0017] The method for a coordinated control system of a new energy autonomous driving vehicle based on reinforcement learning includes the following steps:

[0018] Step 1: Algorithm preset, including multi-objective optimization function construction, construction of multi-mode control strategy and weight base value, and state space construction;

[0019] Step 2: Collect visual information and vehicle information during driving;

[0020] Step 3: Execute a safety mode, an energy-saving mode, or an adaptive mode according to the visual information and vehicle information collected in step 2.

[0021] The step 1 is specifically as follows:

[0022] Step 1.1, multi-objective optimization function construction;

[0023] The multi-objective optimization function is set as:

[0024] J=α*E SAVE +β*C comfort +γ*S safety

[0025] Among them, J is the immediate reward, α, β, γ are weight coefficients and α+β+γ=1, which is adaptively adjusted by the deep reinforcement learning model according to the road slope and road damage conditions. SAVE E is the comprehensive energy efficiency score under actual road conditions, which takes into account the energy recovery efficiency of the suspension system and the energy recovery efficiency of the drive system, and is obtained by fitting the experimental data. SAVE The calculation formula is:

[0026] E save =0.1*η suspension +0.9*η drive

[0027] Among them, η suspension is the energy recovery efficiency of the suspension, which is calculated by the ratio of the linear generator output power of the electromagnetic actuator to the vibration energy input, η drive The energy recovery efficiency of the drive system is the ratio of the energy recovered during braking / coasting to the kinetic energy reduced by the vehicle during braking;

[0028] Establish a road surface status recognition mechanism and set I rough The road roughness index is 0-1 and is calculated using the lidar point cloud data. The formula is: The preset threshold is 0.1; D severity The severity of the damage is scored on a scale of 0-1. The visually recognized damage parameters include crack depth, rut width, and bulge slope and height. ΔSlope is the slope change rate, calculated by fusing wheel speed sensor and inertial measurement unit (IMU) data.

[0029] Definition D severity Normalization calculation formula

[0030]

[0031] Among them, ε i is the influence coefficient of the disease type i: 0.3 for crack disease, 0.5 for road subsidence disease, and 0.2 for bumps; T i,moderateis the severity classification threshold of the disease type corresponding to the national standard for disease i, n is the number of diseases within the observation range, and n≤3. If i>3, the calculation is restarted to prevent the occurrence of too many diseases within the observation distance, which may lead to unreasonable adjustment strategies;

[0032] c i,grade is the severity of the disease i, with 0.1 for mild and 0.2 for severe; P i The measured parameters for disease i include crack width and rutting depth;

[0033] C comfort is the comfort score, which is calculated as follows:

[0034]

[0035] Among them, a virbration Vertical vibration acceleration (unit: m / S 2 ), measured by the vehicle body acceleration sensor; a vmax The maximum vertical vibration acceleration that a human body in a vehicle can normally accept; θ pitch The vehicle body pitch angle is fed back in real time by the inertial measurement unit (IMU); θ vmax is the maximum pitch angle of the vehicle body;

[0036] S safety The safety score is calculated as follows:

[0037]

[0038] Among them, d obsacle is the real-time distance between the vehicle and the obstacle in front, θ slide is the vehicle sideslip angle, reflecting the degree of lateral slip of the vehicle body, θ slide,max The maximum permissible sideslip angle is set according to the vehicle stability standard. Expresses the brake margin. The larger the value, the longer the safety buffer time. Its upper limit is set to 1. vehicle Vehicle real-time speed m / s, measured by vehicle status sensor

[0039]

[0040] a brake is the vehicle's real-time deceleration m / s 2 .

[0041] The step 1 is specifically as follows:

[0042] Step 1.2: Construct a multi-mode control strategy and weight base value;

[0043] Multi-mode control strategy includes energy-saving mode, safety mode and adaptive mode;

[0044] The energy-saving mode is specifically:

[0045] The road is smooth rough ≤0.1, no significant slope change (ΔSlope≤2%), light road damage (D) severity ≤0.1, no obstacles ahead, i.e. d obstacle ≥50m and the vehicle is in a medium or low speed state, that is, v vehicle When the speed is ≤80km / h, the energy-saving mode is triggered; in the energy-saving mode, α is ≥0.6, and the deep reinforcement learning algorithm continuously optimizes the reward value based on this;

[0046] In energy-saving mode, the rear-wheel drive priority is fixed, that is, the front-wheel drive: rear-wheel drive = 0:100, the energy recovery intensity is maximized and the low-damping suspension configuration is configured, that is, 500Ns / m to improve energy recovery efficiency. The slope change ΔSlope>2% or the disease is more serious D severity When it is >0.1, the α weight is gradually reduced and the adaptive mode is switched. When an obstacle suddenly appears, d obstacle ≤20m, immediately switch to safe mode, i.e. γ≥0.9;

[0047] The safety mode is specifically as follows: the safety mode is the highest priority mode, and when the safety mode is triggered, other modes are immediately overwritten. Under sudden road conditions: sudden obstacles d obstacle ≤20m, wet and slippery road μ friction <0.3, or when the vehicle body is in abnormal condition: the sideslip angle increases θ slide >5° or pitch angle θ pitch >3°, safety mode is triggered;

[0048] Fixed front-wheel drive priority (FWD:RWD = 100:0), minimized energy recovery intensity to prevent braking from affecting vehicle balance, increased low-damping suspension damping to 5000Ns / m, air spring stiffness to 2000N / mm, suspension damping and drive torque distribution settings adjusted within 1ms to ensure overall vehicle structural stability and suppress vehicle vibration. Once the danger is eliminated, the obstacle distance is restored to d obstacle ≥30m, gradually reduce the γ weight and switch to adaptive mode;

[0049] The adaptive mode is specifically:

[0050] In non-energy-saving mode and safety mode, the vehicle state is adaptively controlled based on road conditions and deep reinforcement learning algorithms, gradually transitioning to energy-saving mode in non-emergency situations;

[0051] Weight base value: α base ,β base ,γ baseis the initial weight value of the multi-objective optimization function when the car starts. The new energy vehicle starts in the initial energy-saving mode. base =0.6,β base =0.05,γ base =0.35.

[0052] The step 1 is specifically as follows:

[0053] Step 1.3: State space construction; specifically:

[0054] Input state vector:

[0055] S t =[ΔSlope,D severity ,I rough ,v vehicle ,a virbration ,d obstacle ,θ pitch ,η suspension ,η drive ]

[0056] Normalize: Each parameter is normalized to [0,1], and we get:

[0057]

[0058] The step 1 is specifically as follows:

[0059] It also includes the generation of weight corrections and adjustment commands based on a hierarchical actor-critic network;

[0060] The hierarchical actor-critic network includes an upper-level decision network that outputs real-time α, β, and γ, a middle-level mapping network that outputs suspension damping, multi-cavity air spring stiffness and height, and the dynamic distribution ratio of front and rear torque of the drive system, and a lower-level actuator control network that outputs damping fluid flow, actuator excitation current, multi-cavity air spring pressure, and multi-cavity air spring volume;

[0061] The upper decision network includes:

[0062] Input layer: 9 neurons corresponding to the state vector S t,origin The nine dimensions of

[0063] Middle layer: 2 layers of fully connected neurons, 256 neurons in each layer, activation function RELU;

[0064] Output layer: 3 neurons, corresponding to the output vector [Δα raw ,Δβ raw ,Δγ raw ], activation function RELU, Δα raw ,Δβraw ,Δγ raw The value range is [-1,1];

[0065] Scaling: Δα=0.2*Δα raw ,Δβ=0.2*Δβ raw ,Δγ=0.2*Δγ raw , that is, limiting the correction range of each operation to ±0.2;

[0066] Use the softmax function for normalization: The same applies to β and γ;

[0067] And calculate the reward value J based on real-time α, β, γ

[0068] When the car starts:

[0069] The middle-level mapping network includes a suspension damping sub-network, a multi-cavity aerodynamic spring stiffness sub-network, a multi-cavity aerodynamic spring height sub-network, and a front-rear torque dynamic distribution sub-network in a parallel relationship;

[0070] The suspension damping sub-network includes:

[0071] Input layer: receives the weights α, β, γ and 9-dimensional state vector S output by the upper layer t , a total of 12-dimensional vector input;

[0072] Middle layer: two fully connected layers, using ReLU activation function, with 128 neurons in each layer;

[0073] Output layer: uses sigmoid activation function, 1 neuron, and outputs the normalized suspension damping coefficient value σ damp ∈[0,1]; Construct the suspension damping coefficient denormalization formula C damp =500+4500σ damp , get the real-time suspension damping coefficient C damp , ranging from 500-5000Ns / m;

[0074] The multi-cavity aerodynamic spring stiffness sub-network:

[0075] Input layer: receives the weights α, β, γ and 9-dimensional state vector S output by the upper layer t , a total of 12-dimensional vector input;

[0076] Middle layer: two fully connected layers, using ReLU activation function, with 128 neurons in each layer;

[0077] Output layer: uses sigmoid activation function, 1 neuron, and outputs the normalized multi-cavity aerodynamic spring stiffness coefficient value σspring ∈[0,1]; Construct the spring stiffness denormalization formula K spring =500+1500σ spring , get the real-time multi-cavity aerodynamic spring stiffness value K spring , ranging from 500-2000N / mm;

[0078] The multi-cavity air spring height sub-network:

[0079] Input layer: receives the weights α, β, γ and 9-dimensional state vector S output by the upper layer t , a total of 12-dimensional vector input;

[0080] Middle layer: two fully connected layers, using ReLU activation function, with 128 neurons in each layer;

[0081] Output layer: uses sigmoid activation function, 1 neuron, and outputs the normalized spring height adjustment coefficient value σ height ∈[-1,1]; Construct the spring height denormalization formula H spring =200+50σ height , get the real-time spring height value H spring , ranging from 150-250mm;

[0082] The front and rear torque dynamic distribution sub-network:

[0083] Input layer: receives the weights α, β, γ and 9-dimensional state vector S output by the upper layer t , a total of 12-dimensional vector input;

[0084] Middle layer: two fully connected layers, using ReLU activation function, with 128 neurons in each layer;

[0085] Output layer: uses sigmoid activation function, 1 neuron, and outputs the normalized front and rear torque dynamic distribution ratio value σ torque ∈[0,1], retain three decimal places, and get the real-time dynamic torque distribution ratio T time =100*σ torque (%);

[0086] The multi-objective optimization function is the instant reward value of the deep reinforcement learning algorithm:

[0087] J=α*E SAVE +β*C comfort +γ*S safety

[0088] The dynamic adjustment formula of the suspension damping coefficient is:

[0089] C damp =500+4500σdamp

[0090] C damp is the suspension damping coefficient;

[0091] The height adjustment formula of the multi-chamber air spring is:

[0092] K spring =500+1500σ spring

[0093] K spring is the multi-cavity aerodynamic spring stiffness;

[0094] The multi-cavity air spring stiffness adjustment formula is:

[0095] H spring =200+50σ height

[0096] H spring is the height of the multi-chamber air spring;

[0097] The formula for distributing the front and rear driving torque of the drive system is:

[0098] T time =100*σ torque

[0099] H spring Distribute the driving torque between the front and rear axles;

[0100] The damping fluid flow subnetwork:

[0101] Input layer: synthetic vector And enter;

[0102] Middle layer: two fully connected layers, using ReLU activation function, with 64 neurons in each layer;

[0103] Output layer: uses sigmoid activation function, 1 neuron, outputs the real-time flow rate of magnetorheological fluid;

[0104] The actuator excitation current sub-network:

[0105] Input layer: synthetic vector And enter;

[0106] Middle layer: two fully connected layers, using ReLU activation function, with 64 neurons in each layer;

[0107] Output layer: uses sigmoid activation function, 1 neuron, outputs the excitation current of the electromagnetic active suspension actuator;

[0108] The multi-cavity air spring pneumatic sub-network:

[0109] Input layer: synthetic vector

[0110] Middle layer: two fully connected layers, using ReLU activation function, with 64 neurons in each layer;

[0111] Output layer: uses sigmoid activation function, 1 neuron, outputs the real-time air pressure of the multi-cavity air spring;

[0112] The multi-cavity air spring volume subnetwork:

[0113] Input layer: synthetic vector

[0114] Middle layer: two fully connected layers, using ReLU activation function, with 64 neurons in each layer;

[0115] Output layer: uses sigmoid activation function, 1 neuron, and outputs the real-time volume of the multi-cavity air spring.

[0116] The step 2 is specifically as follows:

[0117] This includes image acquisition and processing, using a global shutter camera to capture road images at 60 FPS, covering a horizontal viewing angle of 120° and a vertical viewing angle of 60°; and performing FPGA-based image preprocessing on the captured images.

[0118] It includes noise reduction, which eliminates high-frequency noise based on FPGA algorithms; color correction, which performs white balance processing and gamma correction γ=2.2 on real-time images to ensure lighting consistency; edge enhancement, which performs Canny edge detection on real-time images to enhance crack and rutting features; and ROI extraction, which dynamically defines the 70% area in the center of the real-time image as a high-priority detection area; further, it performs image depth processing based on the on-board GPU; and feature recognition: an improved lightweight convolutional neural network model, which is a backbone network model based on efficient visual feature extraction, for real-time detection of road slope ΔSlope, disease type, and severity score D severity ; and output parameters: ΔSlope, D severity , take 0-1, road roughness index I rough ; and data fusion: LiDAR point cloud is aligned with visual recognition results, and false detections are eliminated through the ICP algorithm;

[0119] It also includes intelligent distance measurement, specifically, generating a three-dimensional point cloud of the road through the laser radar and calculating the obstacle distance d obstacle Millimeter-wave radar compensates for visual failure in rainy and foggy weather, with a detection distance of 200m; wheel speed sensor provides real-time feedback on vehicle speed v vehicle , used for Kalman filter motion compensation;

[0120] It also includes data fusion, specifically, fusing radar and visual data through the Kalman filter algorithm to output the corrected obstacle distance d obstacle .

[0121] Beneficial Effects: Vision recognition and intelligent ranging modules detect road slope changes and road damage characteristics (such as cracks, road subsidence, and bumps) in real time. Combined with a deep reinforcement learning model, the system dynamically adjusts the suspension damping coefficient (500-5000Ns / m), the stiffness and height of the multi-cavity aerodynamic springs, the front-to-rear torque distribution ratio of the drive system (front-wheel drive: rear-wheel drive = 0:100 to 100:0), and the energy recovery intensity (5-50kW). Furthermore, coordinated control of the suspension and drive systems is achieved via TSN Ethernet (latency <1ms). The system can improve overall range by ≥12% and vibration energy recovery efficiency by ≥15%, making it suitable for balancing energy conservation and safety in complex road conditions. BRIEF DESCRIPTION OF THE DRAWINGS

[0122] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.

[0123] Figure 1 A schematic diagram of the system composition.

[0124] Figure 2 Dynamically adjust logic graphs for multi-objective optimization function weights.

[0125] Figure 3 This is the system control logic diagram. DETAILED DESCRIPTION

[0126] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0127] In the description of the present invention, it should be understood that the terms "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", etc., indicating directions or positional relationships, are based on the directions or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific direction, be constructed and operated in a specific direction, and therefore should not be understood as limiting the present invention.

[0128] In the present invention, unless otherwise expressly specified or limited, a first feature being "above" or "below" a second feature may include the first and second features being in direct contact, or may include the first and second features being in contact not directly but through another feature between them. Furthermore, a first feature being "above," "above," and "above" a second feature may include the first feature being directly above or obliquely above the second feature, or may simply mean that the first feature is higher in level than the second feature. A first feature being "below," "below," and "below" a second feature may include the first feature being directly below or obliquely below the second feature, or may simply mean that the first feature is lower in level than the second feature.

[0129] like Figure 1 As shown in the figure, the coordinated control system of new energy autonomous driving vehicles based on reinforcement learning includes a vehicle real-time status monitoring module, a visual recognition module, an intelligent ranging module, an intelligent suspension control module, an intelligent drive control module, a data transmission module and a vehicle control module;

[0130] The vehicle real-time status monitoring module, visual recognition module and intelligent ranging module are respectively connected to the data transmission module signal, and the collected vehicle speed, suspension pitch angle, vehicle sideslip angle, sideslip angle, road disease degree, road surface roughness coefficient and distance between the vehicle and the obstacle are transmitted to the data transmission module;

[0131] The data transmission module is connected to the vehicle control module by signal, and the vehicle control module controls the intelligent suspension control module and the intelligent drive control module to perform actions in real time according to the collected data.

[0132] The vehicle real-time status monitoring module includes an inertial measurement unit (IMU) installed in the upper center of the vehicle chassis. It integrates a MEMS sensor, a three-axis gyroscope, a three-axis accelerometer, and a three-axis magnetometer. The Kalman filter algorithm is used to calculate the vehicle body pitch angle θ in real time. pitch 、Vehicle sideslip angle θ slide and the vehicle vertical vibration acceleration a vibration , data is transmitted to the vehicle control main processor via TSN Ethernet;

[0133] The visual recognition module includes a global shutter camera installed on the top of the front windshield, integrated with an FPGA image and processing unit and a GPU image depth processing unit, and realizes information transmission through the TSN Ethernet data transmission network;

[0134] The intelligent ranging module includes a laser radar and millimeter-wave radar installed in the center of the vehicle's front bumper, as well as a wheel speed sensor installed at the wheel bearing. Through the distance measurement main control processor, it integrates the laser radar point cloud, millimeter-wave radar data and wheel speed data in real time, and uses the Kalman filter algorithm to compensate for the sensor data offset caused by vehicle movement, thereby generating the relative positions of obstacles and road defects and the real-time speed of the vehicle.

[0135] The vehicle control module is equipped with a vehicle control main processor, runs a deep reinforcement learning model based on a hierarchical actor-critic network, and dynamically generates a coordinated control strategy for the suspension and drive system.

[0136] The intelligent suspension control module includes an electromagnetic active suspension actuator and a multi-cavity air spring; the electromagnetic active suspension actuator adjusts the damping fluid viscosity in real time to achieve real-time adjustment of the suspension damping coefficient; the multi-cavity air spring adjusts the stiffness in real time;

[0137] The intelligent drive control module includes a permanent magnet synchronous motor, a silicon carbide inverter, a bidirectional DC-DC converter, and a supercapacitor module for transient energy buffering.

[0138] like Figure 2 and 3 As shown, the method for coordinated control system of new energy autonomous driving vehicle based on reinforcement learning includes the following steps:

[0139] Step 1: Algorithm preset, including multi-objective optimization function construction, construction of multi-mode control strategy and weight base value, and state space construction;

[0140] Step 2: Collect visual information and vehicle information during driving;

[0141] Step 3: Execute a safety mode, an energy-saving mode, or an adaptive mode according to the visual information and vehicle information collected in step 2.

[0142] The step 1 is specifically as follows:

[0143] Step 1.1, multi-objective optimization function construction;

[0144] The multi-objective optimization function is set as:

[0145] J=α*E SAVE +β*C comfort +γ*S safety

[0146] Among them, J is the immediate reward, α, β, γ are weight coefficients and α+β+γ=1, which is adaptively adjusted by the deep reinforcement learning model according to the road slope and road damage conditions. SAVE E is the comprehensive energy efficiency score under actual road conditions, which takes into account the energy recovery efficiency of the suspension system and the energy recovery efficiency of the drive system, and is obtained by fitting the experimental data. SAVE The calculation formula is:

[0147] E save =0.,1*η suspension +0.9*η drive

[0148] Among them, η suspension is the energy recovery efficiency of the suspension, which is calculated by the ratio of the linear generator output power of the electromagnetic actuator to the vibration energy input, η drive The energy recovery efficiency of the drive system is the ratio of the energy recovered during braking / coasting to the kinetic energy reduced by the vehicle during braking;

[0149] Establish a road surface status recognition mechanism and set I rough The road roughness index is 0-1 and is calculated using the lidar point cloud data. The formula is: The preset threshold is 0.1; D severity The severity of the damage is scored on a scale of 0-1. The visually recognized damage parameters include crack depth, rut width, and bulge slope and height. ΔSlope is the slope change rate, calculated by fusing wheel speed sensor and inertial measurement unit (IMU) data.

[0150] Definition D severity Normalization calculation formula

[0151]

[0152] Among them, ε i is the influence coefficient of the disease type i: 0.3 for crack disease, 0.5 for road subsidence disease, and 0.2 for bumps; T i,moderate is the national standard disease severity classification threshold corresponding to disease i, see Table 1 for details; n is the number of diseases within the observation range, and n≤3. If i>3, the calculation is restarted to prevent the occurrence of too many diseases within the observation distance, which may lead to unreasonable adjustment strategies; c i,grade is the severity of the disease i, with 0.1 for mild and 0.2 for severe; P i The measured parameters for disease i include crack width and rutting depth;

[0153] Table 1 Road damage severity classification table

[0154]

[0155] C comfort is the comfort score, which is calculated as follows:

[0156]

[0157] Among them, a virbration Vertical vibration acceleration (unit: m / S 2 ), measured by the vehicle body acceleration sensor; a vmax The maximum vertical vibration acceleration that a human body in a vehicle can normally accept; θ pitch The vehicle body pitch angle is fed back in real time by the inertial measurement unit (IMU); θ vmax is the maximum pitch angle of the vehicle body;

[0158] S safety The safety score is calculated as follows:

[0159]

[0160] Among them, μ friction is the road friction coefficient, which is determined by the tire driving conditions and road conditions, d obstacle is the real-time distance between the vehicle and the obstacle in front, θ slide is the vehicle sideslip angle, reflecting the degree of lateral slip of the vehicle body, θ slide,max The maximum permissible sideslip angle is set according to the vehicle stability standard. Expresses the brake margin. The larger the value, the longer the safety buffer time of the safety brake. Its upper limit is set to 1. vehicle Vehicle real-time speed m / s, measured by vehicle status sensor

[0161]

[0162] a brake is the vehicle's real-time deceleration m / s 2 .

[0163] The step 1 is specifically as follows:

[0164] Step 1.2: Construct a multi-mode control strategy and weight base value;

[0165] Multi-mode control strategy includes energy-saving mode, safety mode and adaptive mode;

[0166] The energy-saving mode is specifically:

[0167] The road is smooth rough ≤0.1, no significant slope change (ΔSlope≤2%), light road damage (D) severity ≤0.1, no obstacles ahead, i.e. dobstacle ≥50m and the vehicle is in a medium or low speed state, that is, v vehicle When the speed is ≤80km / h, the energy-saving mode is triggered; in the energy-saving mode, α is ≥0.6, and the deep reinforcement learning algorithm continuously optimizes the reward value based on this;

[0168] In energy-saving mode, the rear-wheel drive priority is fixed, that is, the front-wheel drive: rear-wheel drive = 0:100, the energy recovery intensity is maximized and the low-damping suspension configuration is configured, that is, 500Ns / m to improve energy recovery efficiency. The slope change ΔSlope>2% or the disease is more serious D severity When it is >0.1, the α weight is gradually reduced and the adaptive mode is switched. When an obstacle suddenly appears, d obstacle ≤20m, immediately switch to safe mode, i.e. γ≥0.9;

[0169] The safety mode is specifically as follows: the safety mode is the highest priority mode, and when the safety mode is triggered, other modes are immediately overwritten. Under sudden road conditions: sudden obstacles d obstacle ≤20m, wet and slippery road μ friction <0.3, or when the vehicle body is in abnormal condition: the sideslip angle increases θ slide >5° or pitch angle θ pitch >3°, safety mode is triggered;

[0170] Fixed front-wheel drive priority (FWD:RWD = 100:0), minimized energy recovery intensity to prevent braking from affecting vehicle balance, increased low-damping suspension damping to 5000Ns / m, air spring stiffness to 2000N / mm, suspension damping and drive torque distribution settings adjusted within 1ms to ensure overall vehicle structural stability and suppress vehicle vibration. Once the danger is eliminated, the obstacle distance is restored to d obstacle ≥30m, gradually reduce the γ weight and switch to adaptive mode;

[0171] The adaptive mode is specifically:

[0172] In non-energy-saving mode and safety mode, the vehicle state is adaptively controlled based on road conditions and deep reinforcement learning algorithms, gradually transitioning to energy-saving mode in non-emergency situations;

[0173] Weight base value: α base ,β base ,γ base is the initial weight value of the multi-objective optimization function when the car starts. The new energy vehicle starts in the initial energy-saving mode. base =0.6,β base =0.05,γ base =0.35.

[0174] The step 1 is specifically as follows:

[0175] Step 1.3: State space construction; specifically:

[0176] Input state vector:

[0177] S t =[ΔSlope,D severity ,I rough ,v vehicle ,a virbration ,d obstacle ,θ pitch ,η suspension ,η drive ]

[0178] Normalize: Each parameter is normalized to [0,1], and we get:

[0179]

[0180] The step 1 is specifically as follows:

[0181] It also includes the generation of weight corrections and adjustment commands based on a hierarchical actor-critic network;

[0182] The hierarchical actor-critic network includes an upper-level decision network that outputs real-time α, β, and γ, a middle-level mapping network that outputs suspension damping, multi-cavity air spring stiffness and height, and the dynamic distribution ratio of front and rear torque of the drive system, and a lower-level actuator control network that outputs damping fluid flow, actuator excitation current, multi-cavity air spring pressure, and multi-cavity air spring volume;

[0183] The Critic network uses the instant reward value obtained by the Actor network evaluation. Its characteristic is that the adaptive coordinated control system, based on meeting safety requirements, takes energy saving as the core goal, adjusts the (α, β, γ) values ​​in real time through multi-objective optimization functions and deep reinforcement learning models, and generates control commands for coordinating the intelligent suspension system and the intelligent drive system. During the calculation process, β is limited to <0.15 to ensure that the adaptive control system focuses on optimizing energy efficiency and safety while taking into account comfort requirements, and optimizes the parameters of the deep reinforcement learning model through the back propagation algorithm. In actual operation, the Critic network will compare the actual reward J real and the predicted value J pred , J pred Calculated by the weight before adjustment, J real The difference between the two (i.e. TD error | J) is calculated by adjusting the weights. real -J pred|) is used as the basis to determine the adjustment range, and through back propagation, the Critic network transmits the error information to the Actor network and the entire deep reinforcement learning model, prompting the model to adjust parameters to improve the accuracy of prediction and the effectiveness of control strategy. When the vehicle encounters new road conditions (such as sudden obstacles) during driving, the actual reward J real The Critic network calculates the TD error and performs backpropagation through the Actor network structure, so that the Actor network can adjust the output control instructions, such as increasing suspension damping, adjusting torque distribution, etc., to better adapt to new road conditions.

[0184] The upper decision network includes:

[0185] Input layer: 9 neurons corresponding to the state vector S t,origin The nine dimensions of

[0186] Middle layer: 2 layers of fully connected neurons, 256 neurons in each layer, activation function RELU;

[0187] Output layer: 3 neurons, corresponding to the output vector [Δα raw ,Δβ raw ,Δγ raw ], activation function RELU, Δα raw ,Δβ raw ,Δγ raw The value range is [-1,1];

[0188] Scaling: Δα=0.2*Δα raw ,Δβ=0.2*Δβ raw ,Δγ=0.2*Δγ raw , that is, limiting the correction range of each operation to ±0.2;

[0189] Use the softmax function for normalization: The same applies to β and γ;

[0190] And calculate the reward value J based on real-time α, β, γ

[0191] When the car starts:

[0192] The middle-level mapping network includes a suspension damping sub-network, a multi-cavity aerodynamic spring stiffness sub-network, a multi-cavity aerodynamic spring height sub-network, and a front-rear torque dynamic distribution sub-network in a parallel relationship;

[0193] The suspension damping sub-network includes:

[0194] Input layer: receives the weights α, β, γ and 9-dimensional state vector S output by the upper layer t, a total of 12-dimensional vector input;

[0195] Middle layer: two fully connected layers, using ReLU activation function, with 128 neurons in each layer;

[0196] Output layer: uses sigmoid activation function, 1 neuron, and outputs the normalized suspension damping coefficient value σ damp ∈[0,1]; Construct the suspension damping coefficient denormalization formula C damp =500+4500σ damp , get the real-time suspension damping coefficient C damp , ranging from 500-5000Ns / m;

[0197] The multi-cavity aerodynamic spring stiffness sub-network:

[0198] Input layer: receives the weights α, β, γ and 9-dimensional state vector S output by the upper layer t , a total of 12-dimensional vector input;

[0199] Middle layer: two fully connected layers, using ReLU activation function, with 128 neurons in each layer;

[0200] Output layer: uses sigmoid activation function, 1 neuron, and outputs the normalized multi-cavity aerodynamic spring stiffness coefficient value σ spring ∈[0,1]; Construct the spring stiffness denormalization formula K spring =500+1500σ spring , get the real-time multi-cavity aerodynamic spring stiffness value K spring , ranging from 500-2000N / mm;

[0201] The multi-cavity air spring height sub-network:

[0202] Input layer: receives the weights α, β, γ and 9-dimensional state vector S output by the upper layer t , a total of 12-dimensional vector input;

[0203] Middle layer: two fully connected layers, using ReLU activation function, with 128 neurons in each layer;

[0204] Output layer: uses sigmoid activation function, 1 neuron, and outputs the normalized spring height adjustment coefficient value σ height ∈[-1,1]; Construct the spring height denormalization formula H spring =200+50σ height , get the real-time spring height value H spring , ranging from 150-250mm;

[0205] The front and rear torque dynamic distribution sub-network:

[0206] Input layer: receives the weights α, β, γ and 9-dimensional state vector S output by the upper layer t , a total of 12-dimensional vector input;

[0207] Middle layer: two fully connected layers, using ReLU activation function, with 128 neurons in each layer;

[0208] Output layer: uses sigmoid activation function, 1 neuron, and outputs the normalized front and rear torque dynamic distribution ratio value σ torque ∈[0,1], retain three decimal places, and get the real-time dynamic torque distribution ratio T time =100*σ torque (%);

[0209] The multi-objective optimization function is the instant reward value of the deep reinforcement learning algorithm:

[0210] J=α*E SAVE +β*C comfort +γ*S safety

[0211] The dynamic adjustment formula of the suspension damping coefficient is:

[0212] C damp =500+4500σ damp

[0213] C damp is the suspension damping coefficient;

[0214] The height adjustment formula of the multi-chamber air spring is:

[0215] K spring =500+1500σ spring

[0216] K spring is the multi-cavity aerodynamic spring stiffness;

[0217] The multi-cavity air spring stiffness adjustment formula is:

[0218] H spring =200+50σ height

[0219] H spring is the height of the multi-chamber air spring;

[0220] The formula for distributing the front and rear driving torque of the drive system is:

[0221] T time =100*σ torque

[0222] H springDistribute the driving torque between the front and rear axles;

[0223] The damping fluid flow subnetwork:

[0224] Input layer: synthetic vector And enter;

[0225] Middle layer: two fully connected layers, using ReLU activation function, with 64 neurons in each layer;

[0226] Output layer: uses sigmoid activation function, 1 neuron, outputs the real-time flow rate of magnetorheological fluid;

[0227] The actuator excitation current sub-network:

[0228] Input layer: synthetic vector And enter;

[0229] Middle layer: two fully connected layers, using ReLU activation function, with 64 neurons in each layer;

[0230] Output layer: uses sigmoid activation function, 1 neuron, outputs the excitation current of the electromagnetic active suspension actuator;

[0231] The multi-cavity air spring pneumatic sub-network:

[0232] Input layer: synthetic vector

[0233] Middle layer: two fully connected layers, using ReLU activation function, with 64 neurons in each layer;

[0234] Output layer: uses sigmoid activation function, 1 neuron, outputs the real-time air pressure of the multi-cavity air spring;

[0235] The multi-cavity air spring volume subnetwork:

[0236] Input layer: synthetic vector

[0237] Middle layer: two fully connected layers, using ReLU activation function, with 64 neurons in each layer;

[0238] Output layer: uses sigmoid activation function, 1 neuron, and outputs the real-time volume of the multi-cavity air spring.

[0239] The step 2 is specifically as follows:

[0240] This includes image acquisition and processing, using a global shutter camera to capture road images at 60 FPS, covering a horizontal viewing angle of 120° and a vertical viewing angle of 60°; and performing FPGA-based image preprocessing on the captured images.

[0241] It includes noise reduction, which eliminates high-frequency noise based on FPGA algorithms; color correction, which performs white balance processing and gamma correction γ=2.2 on real-time images to ensure lighting consistency; edge enhancement, which performs Canny edge detection on real-time images to enhance crack and rutting features; and ROI extraction, which dynamically defines the 70% area in the center of the real-time image as a high-priority detection area; further, it performs image depth processing based on the on-board GPU; and feature recognition: an improved lightweight convolutional neural network model, which is a backbone network model based on efficient visual feature extraction, for real-time detection of road slope ΔSlope, disease type, and severity score D severity ; and output parameters: ΔSlope, D severity , take 0-1, road roughness index I rough ; and data fusion: LiDAR point cloud is aligned with visual recognition results, and false detections are eliminated through the ICP algorithm;

[0242] It also includes intelligent distance measurement, specifically, generating a three-dimensional point cloud of the road through the laser radar and calculating the obstacle distance d obstacle Millimeter-wave radar compensates for visual failure in rainy and foggy weather, with a detection distance of 200m; wheel speed sensor provides real-time feedback on vehicle speed v vehicle , used for Kalman filter motion compensation;

[0243] It also includes data fusion, specifically, fusing radar and visual data through the Kalman filter algorithm to output the corrected obstacle distance d obstacle .

[0244] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Reference can be made to the common and similar parts between the various embodiments. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method description.

[0245] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not limited to the embodiments shown herein but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A coordinated control system for new energy autonomous driving vehicles based on reinforcement learning, characterized by: It includes vehicle real-time status monitoring module, visual recognition module, intelligent ranging module, intelligent suspension control module, intelligent drive control module, data transmission module and vehicle control module; The vehicle real-time status monitoring module, visual recognition module and intelligent ranging module are respectively connected to the data transmission module signal, and the collected vehicle speed, suspension pitch angle, vehicle sideslip angle, sideslip angle, road disease degree, road surface roughness coefficient and distance between the vehicle and the obstacle are transmitted to the data transmission module; The data transmission module is connected to the vehicle control module by signal, and the vehicle control module controls the intelligent suspension control module and the intelligent drive control module to perform actions in real time according to the collected data.

2. The coordinated control system for new energy autonomous driving vehicles based on reinforcement learning according to claim 1, characterized in that: The vehicle real-time status monitoring module includes an inertial measurement unit (IMU) installed in the upper center of the vehicle chassis. It integrates a MEMS sensor, a three-axis gyroscope, a three-axis accelerometer, and a three-axis magnetometer. The Kalman filter algorithm is used to calculate the vehicle body pitch angle θ in real time. pitch 、Vehicle sideslip angle θ slide and the vehicle vertical vibration acceleration a vibration , data is transmitted to the vehicle control main processor via TSN Ethernet; The visual recognition module includes a global shutter camera installed on the top of the front windshield, integrated with an FPGA image and processing unit and a GPU image depth processing unit, and realizes information transmission through the TSN Ethernet data transmission network; The intelligent ranging module includes a laser radar and millimeter-wave radar installed in the center of the vehicle's front bumper, as well as a wheel speed sensor installed at the wheel bearing. Through the distance measurement main control processor, it integrates the laser radar point cloud, millimeter-wave radar data and wheel speed data in real time, and uses the Kalman filter algorithm to compensate for the sensor data offset caused by vehicle movement, thereby generating the relative positions of obstacles and road defects and the real-time speed of the vehicle.

3. The coordinated control system for new energy autonomous driving vehicles based on reinforcement learning according to claim 1 is characterized in that: The vehicle control module is equipped with a vehicle control main processor, runs a deep reinforcement learning model based on a hierarchical actor-critic network, and dynamically generates a coordinated control strategy for the suspension and drive system.

4. The coordinated control system for new energy autonomous driving vehicles based on reinforcement learning according to claim 1, characterized in that: The intelligent suspension control module includes an electromagnetic active suspension actuator and a multi-cavity air spring; the electromagnetic active suspension actuator adjusts the damping fluid viscosity and flow rate in real time to achieve real-time adjustment of the suspension damping coefficient; the multi-cavity air spring adjusts the stiffness in real time; The intelligent drive control module includes a permanent magnet synchronous motor, a silicon carbide inverter, a bidirectional DC-DC converter, and a supercapacitor module for transient energy buffering.

5. The coordinated control system and method for new energy autonomous driving vehicles based on reinforcement learning according to any one of claims 1 to 4, characterized in that: The following steps are involved: Step 1: Algorithm preset, including multi-objective optimization function construction, construction of multi-mode control strategy and weight base value, and state space construction; Step 2: Collect visual information and vehicle information during driving; Step 3: Execute a safety mode, an energy-saving mode, or an adaptive mode according to the visual information and vehicle information collected in step 2.

6. The coordinated control method for new energy autonomous driving vehicles based on reinforcement learning according to claim 5 is characterized in that: The step 1 is specifically as follows: Step 1.1, multi-objective optimization function construction; The multi-objective optimization function is set as: J=α*E SAVE +β*C comfort +γ*S safety Among them, J is the immediate reward, α, β, γ are weight coefficients and α+β+γ=1, which is adaptively adjusted by the deep reinforcement learning model according to the road slope and road damage conditions. SAVE E is the comprehensive energy efficiency score under actual road conditions, which takes into account the energy recovery efficiency of the suspension system and the energy recovery efficiency of the drive system, and is obtained by fitting the experimental data. SAVE The calculation formula is: E save =0.1*h suspension +0.9*h drive Among them, η suspension is the energy recovery efficiency of the suspension, which is calculated by the ratio of the linear generator output power of the electromagnetic actuator to the vibration energy input, η drive The energy recovery efficiency of the drive system when braking is taken, which is the ratio of the energy recovered by braking / coasting to the kinetic energy reduced by the vehicle during braking; Establish a road surface status recognition mechanism and set I rough The road roughness index is 0-1 and is calculated using the lidar point cloud data. The formula is: The preset threshold is 0.1; D severity The severity of the damage is scored on a scale of 0-1. The visually recognized damage parameters include crack depth, rut width, and bulge slope and height. ΔSlope is the slope change rate, calculated by fusing wheel speed sensor and inertial measurement unit (IMU) data. Definition D severity Normalization calculation formula Among them, ε i is the influence coefficient of the disease type i: 0.3 for crack disease, 0.5 for road subsidence disease, and 0.2 for bumps; T i,moderate is the severity classification threshold of the disease type corresponding to the national standard for disease i, n is the number of diseases within the observation range, and n≤3. If i>3, the calculation is restarted to prevent the occurrence of too many diseases within the observation distance, which may lead to unreasonable adjustment strategies; c i,grade is the severity of the disease i, with 0.1 for mild and 0.2 for severe; P i The measured parameters for disease i include crack width and rutting depth; C comfort is the comfort score, which is calculated as follows: Among them, a virbration Vertical vibration acceleration (unit: m / S 2 ), measured by the vehicle body acceleration sensor; a vmax The maximum vertical vibration acceleration that a human body in a vehicle can normally accept; θ pitch The vehicle body pitch angle is fed back in real time by the inertial measurement unit (IMU); θ vmax is the maximum pitch angle of the vehicle body; S safety The safety score is calculated as follows: Among them, d obstacle is the real-time distance between the vehicle and the obstacle in front, θ slide is the vehicle sideslip angle, reflecting the degree of lateral slip of the vehicle body, θ slide,max The maximum permissible sideslip angle is set according to the vehicle stability standard. Expresses the brake margin. The larger the value, the longer the safety buffer time. Its upper limit is set to 1. vehicle Vehicle real-time speed m / s, measured by vehicle status sensor a brake is the vehicle's real-time deceleration m / s 2 .

7. The coordinated control method for new energy autonomous driving vehicles based on reinforcement learning according to claim 6 is characterized in that: The step 1 is specifically as follows: Step 1.2: Construct a multi-mode control strategy and weight base value; Multi-mode control strategy includes energy-saving mode, safety mode and adaptive mode; The energy-saving mode is specifically: The road is smooth rough ≤0.1, no significant slope change (ΔSlope≤2%), light road damage (D) severity ≤0.1, no obstacles ahead, i.e. d obstacle ≥50m and the vehicle is in a medium or low speed state, that is, v vehicle When the speed is ≤80km / h, the energy-saving mode is triggered; in the energy-saving mode, α is ≥0.6, and the deep reinforcement learning algorithm continuously optimizes the reward value based on this; In energy-saving mode, the rear-wheel drive priority is fixed, that is, the front-wheel drive: rear-wheel drive = 0:100, the energy recovery intensity is maximized and the low-damping suspension configuration is configured, that is, 500Ns / m to improve energy recovery efficiency. The slope change ΔSlope>2% or the disease is more serious D severity When it is >0.1, the α weight is gradually reduced and the adaptive mode is switched. When an obstacle suddenly appears, d obstacle ≤20m, immediately switch to safe mode, i.e. γ≥0.9; The safety mode is specifically as follows: the safety mode is the highest priority mode, and when the safety mode is triggered, other modes are immediately overwritten. Under sudden road conditions: sudden obstacles d obstacle ≤20m, wet and slippery road μ friction <0.3, or when the vehicle body is in abnormal condition: the sideslip angle increases θ slide >5° or pitch angle θ pitch >3°, safety mode is triggered; Fixed front-wheel drive priority (FWD:RWD = 100:0), minimized energy recovery intensity to prevent braking from affecting vehicle balance, increased low-damping suspension damping to 5000Ns / m, air spring stiffness to 2000N / mm, suspension damping and drive torque distribution settings adjusted within 1ms to ensure overall vehicle structural stability and suppress vehicle vibration. Once the danger is eliminated, the obstacle distance is restored to d obstacle ≥30m, gradually reduce the γ weight and switch to adaptive mode; The adaptive mode is specifically: In non-energy-saving mode and safety mode, the vehicle state is adaptively controlled based on road conditions and deep reinforcement learning algorithms, gradually transitioning to energy-saving mode in non-emergency situations; Weight base value: α base ,β base ,γ base is the initial weight value of the multi-objective optimization function when the car starts. The new energy vehicle starts in the initial energy-saving mode. base =0.6,β base =0.05,γ base =0.

35.

8. The coordinated control method for new energy autonomous driving vehicles based on reinforcement learning according to claim 7 is characterized in that: The step 1 is specifically as follows: Step 1.3: State space construction; specifically: Input state vector: S t =[ΔSlope,D severity ,I rough ,v vehicle ,a virbration ,d obstacle ,i pitch ,or suspension ,or drive ] Normalize: Each parameter is normalized to [0,1], and we get:

9. The coordinated control method for new energy autonomous driving vehicles based on reinforcement learning according to claim 8, characterized in that: The step 1 is specifically as follows: It also includes the generation of weight corrections and adjustment commands based on a hierarchical actor-critic network; The hierarchical actor-critic network includes an upper-level decision network that outputs real-time α, β, and γ, a middle-level mapping network that outputs suspension damping, multi-cavity air spring stiffness and height, and the dynamic distribution ratio of front and rear torque of the drive system, and a lower-level actuator control network that outputs damping fluid flow, actuator excitation current, multi-cavity air spring pressure, and multi-cavity air spring volume; The upper decision network includes: Input layer: 9 neurons corresponding to the state vector S t,origin The nine dimensions of Middle layer: 2 layers of fully connected neurons, 256 neurons in each layer, activation function RELU; Output layer: 3 neurons, corresponding to the output vector [Δα raw ,Δβ raw ,Δγ raw ], activation function RELU, Δα raw ,Δβ raw ,Δγ raw The value range is [-1,1]; Scaling: Δα=0.2*Δα raw ,Δβ=0.2*Δβ raw ,Δγ=0.2*Δγ raw , that is, limiting the correction range of each operation to ±0.2; Use the softmax function for normalization: The same applies to β and γ; And calculate the reward value J based on real-time α, β, γ When the car starts: The middle-level mapping network includes a suspension damping sub-network, a multi-cavity aerodynamic spring stiffness sub-network, a multi-cavity aerodynamic spring height sub-network, and a front-rear torque dynamic distribution sub-network in a parallel relationship; The suspension damping sub-network includes: Input layer: receives the weights α, β, γ and 9-dimensional state vector S output by the upper layer t , a total of 12-dimensional vector input; Middle layer: two fully connected layers, using ReLU activation function, with 128 neurons in each layer; Output layer: uses sigmoid activation function, 1 neuron, and outputs the normalized suspension damping coefficient value σ damp ∈[0,1]; Construct the suspension damping coefficient denormalization formula C damp =500+4500σ damp , get the real-time suspension damping coefficient C damp , ranging from 500-5000Ns / m; The multi-cavity aerodynamic spring stiffness sub-network: Input layer: receives the weights α, β, γ and 9-dimensional state vector S output by the upper layer t , a total of 12-dimensional vector input; Middle layer: two fully connected layers, using ReLU activation function, with 128 neurons in each layer; Output layer: uses sigmoid activation function, 1 neuron, and outputs the normalized multi-cavity aerodynamic spring stiffness coefficient value σ spring ∈[0,1]; Construct the spring stiffness denormalization formula K spring =500+1500σ spring , get the real-time multi-cavity aerodynamic spring stiffness value K spring , ranging from 500-2000N / mm; The multi-cavity air spring height sub-network: Input layer: receives the weights α, β, γ and 9-dimensional state vector S output by the upper layer t , a total of 12-dimensional vector input; Middle layer: two fully connected layers, using ReLU activation function, with 128 neurons in each layer; Output layer: uses sigmoid activation function, 1 neuron, and outputs the normalized spring height adjustment coefficient value σ height ∈[-1,1]; Construct the spring height denormalization formula H spring =200+50σ height , get the real-time spring height value H spring , ranging from 150-250mm; The front and rear torque dynamic distribution sub-network: Input layer: receives the weights α, β, γ and 9-dimensional state vector S output by the upper layer t , a total of 12-dimensional vector input; Middle layer: two fully connected layers, using ReLU activation function, with 128 neurons in each layer; Output layer: uses sigmoid activation function, 1 neuron, and outputs the normalized front and rear torque dynamic distribution ratio value σ torque ∈[0,1], retain three decimal places, and get the real-time dynamic torque distribution ratio T time =100*σ torque (%); The multi-objective optimization function is the instant reward value of the deep reinforcement learning algorithm: J=α*E SAVE +β*C comfort +γ*S safety The dynamic adjustment formula of the suspension damping coefficient is: C damp =500+4500σ damp C damp is the suspension damping coefficient; The height adjustment formula of the multi-chamber air spring is: K spring =500+1500σ spring K spring is the multi-cavity aerodynamic spring stiffness; The multi-cavity air spring stiffness adjustment formula is: H spring =200+50σ height H spring is the height of the multi-chamber air spring; The formula for distributing the front and rear driving torque of the drive system is: T time =100*σ torque H spring Distribute the driving torque between the front and rear axles; The damping fluid flow subnetwork: Input layer: synthetic vector And enter; Middle layer: two fully connected layers, using ReLU activation function, with 64 neurons in each layer; Output layer: uses sigmoid activation function, 1 neuron, outputs the real-time flow rate of magnetorheological fluid; The actuator excitation current sub-network: Input layer: synthetic vector And enter; Middle layer: two fully connected layers, using ReLU activation function, with 64 neurons in each layer; Output layer: uses sigmoid activation function, 1 neuron, outputs the excitation current of the electromagnetic active suspension actuator; The multi-cavity air spring pneumatic sub-network: Input layer: synthetic vector Middle layer: two fully connected layers, using ReLU activation function, with 64 neurons in each layer; Output layer: uses sigmoid activation function, 1 neuron, outputs the real-time air pressure of the multi-cavity air spring; The multi-cavity air spring volume subnetwork: Input layer: synthetic vector Middle layer: two fully connected layers, using ReLU activation function, with 64 neurons in each layer; Output layer: uses sigmoid activation function, 1 neuron, and outputs the real-time volume of the multi-cavity air spring.

10. The coordinated control method for new energy autonomous driving vehicles based on reinforcement learning according to claim 9 is characterized in that: The step 2 is specifically as follows: This includes image acquisition and processing, using a global shutter camera to capture road images at 60 FPS, covering a horizontal viewing angle of 120° and a vertical viewing angle of 60°; and performing FPGA-based image preprocessing on the captured images. This includes noise reduction, which eliminates high-frequency noise based on an FPGA algorithm; and color correction, which performs white balance processing and gamma correction γ=2.2 on real-time images to ensure lighting consistency; and edge enhancement, which performs Canny edge detection on real-time images to enhance crack and rutting features; And ROI extraction, that is, dynamically demarcating the central 70% area of ​​the real-time image as a high-priority detection area; further, performing image depth processing based on the on-board GPU; include Feature recognition: An improved lightweight convolutional neural network model, namely a backbone network model based on efficient visual feature extraction, is used to detect road slope ΔSlope, disease type, and severity score D in real time. severity ; and output parameters: ΔSlope, D severity , take 0-1, road roughness index I rough ; and data fusion: LiDAR point cloud is aligned with visual recognition results, and false detections are eliminated through the ICP algorithm; It also includes intelligent distance measurement, specifically, generating a three-dimensional point cloud of the road through the laser radar and calculating the obstacle distance d obstacle Millimeter-wave radar compensates for visual failure in rainy and foggy weather, with a detection distance of 200m; wheel speed sensor provides real-time feedback on vehicle speed v vehicle , used for Kalman filter motion compensation; It also includes data fusion, specifically, fusing radar and visual data through the Kalman filter algorithm to output the corrected obstacle distance d obstacle .