Data-driven intelligent chassis integrated control method and system
Through the data-driven intelligent chassis integrated control method, the Actor-Critic policy network is used to integrate vehicle-environment interactive observation information, coordinate the control amount of each chassis subsystem, and solve the problems of complexity and insufficient performance of the chassis intelligent integrated control algorithm in the existing technology, achieving more efficient and stable vehicle control.
Patent Information
- Application Number
- CN202510097223.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-06-06
AI Technical Summary
The existing chassis intelligent integrated fusion control algorithm has problems such as system integration complexity, insufficient stability and robustness, insufficient multi-objective optimization, model uncertainty and software and hardware collaborative design difficulties, resulting in a lack of overall performance and adaptability.
Using data-driven intelligent chassis integrated control method, the vehicle-environment interaction observation information is integrated through the Actor-Critic policy network, and the expected control quantity signals in the XYZ direction of the actuators of each subsystem of the chassis are output to realize real-time coordinated distribution of chassis suspension, braking and steering systems.
It improves the overall performance and robustness of chassis control, enhances adaptability to the external environment and its own hardware, and achieves the improvement of vehicle handling and stability.
Smart Images

Figure CN120096594A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent chassis integrated control, and in particular to a data-driven intelligent chassis integrated control method and system. Background Art
[0002] With the rapid development of automobile intelligence, future-oriented intelligent chassis control has ushered in new demands and challenges: integrating and coordinating vehicle motion and control through cross-domain wire-controlled actuators, integrating multiple control systems or control strategies to achieve more efficient and stable vehicle performance management. Chassis fusion control involves multiple aspects of the chassis, such as suspension control, steering control, and braking control. Its main purpose is to improve the vehicle's handling, stability, and comfort under various driving conditions to achieve a safer, more comfortable, and agile driving experience. At present, both new energy vehicle companies and traditional vehicle companies have invested important R&D efforts and resource costs in chassis integration technology and chassis intelligence technology. For example, the Tuling chassis focusing on multi-domain collaborative control, the AI digital chassis dedicated to improving vehicle performance, the Tianxing chassis system equipped with a chassis controller, and the Linglizard digital chassis and Yunnian chassis system that can achieve unconventional motion control, etc.
[0003] The intelligent wire-controlled chassis with wire-controlled braking, wire-controlled steering, and active suspension as the core content can integrate the chassis control system, electronic control system, and intelligent system to achieve XYZ three-way coordinated control of the vehicle body, thereby improving the intelligence and precision of chassis control. At the same time, it can also realize the integration of the chassis with other automotive domains (such as the power domain and the body domain), realize cross-domain coordinated control, and improve the performance of the whole vehicle. The development of this technology will bring many changes to the automotive industry, including improving the safety and driving experience of automobiles, reducing energy consumption, and promoting the development of the automotive industry in the direction of intelligence and electrification, etc., and provide better control stability and comfort. For example, if you turn on an icy road or when the car is driving at high speed, it is very difficult for ordinary drivers to ensure that the vehicle does not lose control, does not push the head, and smoothly completes the turning dynamics when they turn the steering wheel sharply or make a large correction. At this time, if you can combine wire-controlled steering, wire-controlled braking, and active suspension, by collaboratively adjusting the control amount of each actuator, you can maximize the steering and stabilization capabilities of the vehicle itself to avoid accidents.
[0004] Under the coordinated control of the chassis controller, the integrated control of wire-controlled systems such as braking, steering, and suspension can achieve faster dynamic response, shorter pressure building capabilities, fine-grained precision wire-controlled braking, wire-controlled steering, and electronic suspension parameter adjustment (flexible adjustment of damping, stiffness, height, shock absorption control), etc., and can continuously and actively learn the laws of road conditions and continuously improve the chassis comfort performance, thereby creating a digital chassis cerebellum similar to the first reaction of the human nervous system, laying the foundation for more advanced advanced driver assistance systems (ADAS) and autonomous driving, and becoming an important trend in the era of intelligent driving.
[0005] At present, the research on chassis intelligent integrated fusion control algorithm is not mature yet, and there are many technical pain points and defects that need to be tackled and broken through, mainly including: (1) System integration complexity: Chassis intelligent integration requires integrating multiple subsystems (such as suspension system, braking system, steering system, etc.) into a unified control system, which increases the complexity of the system and requires a highly integrated control algorithm to coordinate the work of each subsystem. (2) Stability and robustness: Under various working conditions, the chassis control system needs to maintain stability and robustness, especially when facing external disturbances and changes in system parameters. (3) Multi-objective optimization: Chassis control systems often need to consider multiple performance indicators at the same time, such as comfort, handling, safety, etc., which requires the control algorithm to be able to perform multi-objective optimization. (4) Model uncertainty: The actual vehicle dynamics model may have uncertainties, such as tire models, nonlinear characteristics of the suspension system, etc. The control algorithm needs to be able to handle these uncertainties. (5) Co-design of software and hardware: The control algorithm needs to be co-designed with hardware devices such as actuators and sensors to achieve the best control effect. (6) Algorithm adaptability: The control algorithm needs to be able to adapt to different vehicle characteristics and driving styles and have good adaptability.
[0006] For example, the invention application with application number 202211429906.6 discloses an integrated control method and system for intelligent vehicle chassis and task load. The application scheme constructs a reconfigurable dynamic model by collecting road environment information, vehicle posture information and task target information; under the constraints of vehicle state observations and road state observations, based on the force balance optimization algorithm, the reconfigurable dynamic model is solved to obtain the acceleration parameter target value of each vehicle action actuator and the acceleration parameter target value of each task action actuator; the control command of each vehicle action actuator and each task action actuator is generated; according to the control instructions of each vehicle action actuator and the control instructions of each task action actuator, the corresponding vehicle action actuator and the corresponding task action actuator are controlled to act, and the actuators in the chassis domain and the task load domain are controlled to coordinate actions, and finally the integrated control of the chassis and the task load is realized. However, its scheme also has the following problems: complex system integration, insufficient multi-objective optimization such as comfort, handling, and safety.
[0007] Therefore, in reality, it is necessary to further optimize the integrated control of the intelligent chassis, further improve the integration, modularization and intelligence of the chassis control algorithm, accelerate technological innovation, and promote the sustainable development of the automotive industry. Summary of the invention
[0008] In view of the above problems, the purpose of the present invention is to provide a data-driven intelligent chassis integrated control method and system, realize the coordinated control of the vehicle chassis in three directions XYZ, realize the real-time coordinated allocation of the control quantities of the chassis suspension system, braking system and steering system, and improve the overall performance of chassis control. At the same time, chassis control is combined with data drive to improve the robustness and adaptability of the control algorithm to the external environment and its own hardware.
[0009] The embodiments of the present invention provide a data-driven intelligent chassis integrated control method and system.
[0010] A first aspect: A data-driven intelligent chassis integrated control method, comprising:
[0011] The vehicle-environment interaction observation information is collected as input, and the trained and verified Actor-Critic strategy network is used to output the desired control quantity signals in the XYZ directions of the actuators of each chassis subsystem to control the desired states of the actuators of each chassis subsystem.
[0012] Optionally, the collected vehicle-environment interaction observation information is used as input, and the vector representation is:
[0013]
[0014] Among them, σ steer Input for driver's steering wheel angle; are the longitudinal, lateral and vertical accelerations of the vehicle, The yaw, roll and pitch angular velocity signals of the vehicle; is the vehicle attitude angular acceleration signal; It is the vehicle wheel speed signal; It is the estimated signal of the cylinder pressure feedback of the vehicle wheel end; It is the vertical speed signal of the suspension wheel end; Estimation signal for wheel end vertical load;
[0015] Among them, the interactive observation information increases or decreases according to different tasks and application scenarios.
[0016] Optionally, the chassis subsystems include: a steering system, a braking system and a suspension system, and the desired control quantity signals in the XYZ directions of the actuators of each subsystem are represented by vectors as follows:
[0017] Y output=[Γ d ,Δp dwheel ,ΔF S ,ΔF D ]
[0018]
[0019] Among them, Γ d represents the desired steering system compensation, Δp dwheel Indicates the expected change in wheel end brake system cylinder pressure, ΔF S Denotes the desired change in suspension system spring force, ΔF D Indicates the desired change in the spring damping force of the suspension system;
[0020] Among them, the desired control quantity signals in the XYZ directions of each subsystem actuator are increased or decreased according to different tasks and application scenarios.
[0021] Optionally, the Actor-Critic strategy network includes an Actor network and a Critic network, wherein: the Actor network is responsible for generating action strategies, the Critic network uses the observation information and privileged information obtained through interaction to obtain rewards, generates action value information valuation based on the rewards and calculates the cost to guide the Actor network to perform strategy update training.
[0022] Optionally, the privileged information is directly obtained from the simulation environment, including: the road adhesion coefficient felt by the tire Vehicle speed signal v, center of mass sideslip angle signal β and vehicle desired steering angle σ dsteer .
[0023] Optionally, the rewards include:
[0024] Slippage rate bonus, the formula is:
[0025]
[0026] Among them, μ i is the tire-ground adhesion coefficient, is the wheel slip rate;
[0027] Turning to tracking rewards, the formula is expressed as:
[0028] r steering =-exp(-v·|ε steer |)
[0029] Among them, ε steer =σ steer -σ dsteer is the steering wheel steering error;
[0030] Steering-brake stability reward, the formula is expressed as:
[0031]
[0032] in, and is the excessive yaw rate when the vehicle turns, Δp i is the cylinder pressure of the wheel end brake system;
[0033] Steering-suspension stability bonus, the formula is expressed as:
[0034]
[0035] in, is too large yaw rate;
[0036] Vertical stability reward, the formula is expressed as:
[0037]
[0038] Among them, i represents the wheel code, and c represents a positive constant; Indicates a tire with large vertical force, a y z i Excessive vertical acceleration;
[0039] Horizontal macro rewards, the formula is expressed as:
[0040]
[0041] Among them, |a y |lateral acceleration, |β| sideslip angle of center of mass;
[0042] Braking reward, the formula is expressed as:
[0043] r brake =-|a x |-||v wheel ||-var(v wheel )
[0044] Among them, |a x | is the longitudinal acceleration;
[0045] Action smoothness reward, the formula is expressed as:
[0046] r action =-||Y output -Y last_output ||
[0047] Among them, ||Y output -Y last_output || represents the output amplitude of the constraint action.
[0048] Optionally, when the vehicle is braking with ABS or AEB, the slip rate bonus is expressed as:
[0049]
[0050] Among them, λ d =0.2 is the expected slip rate of effective wheel braking. When the reward slip rate is greater than 0.5, the pressure is reduced. When the reward slip rate is less than 0.05, the pressure is increased. i | is the wheel end cylinder pressure.
[0051] Optionally, the Actor-Critic strategy network training process includes:
[0052] S11. Use the current policy network to forward propagate the observation information to obtain action information, value function estimation information and entropy information;
[0053] S12, adjust the learning rate, specify the target KL divergence and calculate the KL divergence difference between the current strategy and the old strategy, and adaptively adjust the learning rate based on the result;
[0054] S13. Calculate the Actor-Critic strategy network loss, including action strategy loss, entropy loss, and estimated value loss;
[0055] S14. Use the optimizer to optimize and update the Actor-Critic strategy network parameters.
[0056] A second aspect: A data-driven intelligent chassis integrated control system, comprising:
[0057] Data acquisition module, collecting vehicle-environment interaction observation information;
[0058] The Actor-Critic strategy network module deploys the trained and verified Actor-Critic strategy network model, collects vehicle-environment interaction observation information as input, and outputs the desired control quantity signals in the XYZ directions of the actuators of each chassis subsystem;
[0059] The signal output module controls the desired state of the actuators of each chassis subsystem according to the control signal.
[0060] Optionally, the control system adopts PC-CAN loop model verification and bench-semi-physical loop model verification, wherein:
[0061] The PC-CAN ring model verification process includes:
[0062] S21. Deploy the trained Actor-Critic strategy network model on the control board and configure the CAN communication port of the control board.
[0063] S22, configuring the CarSim working environment on the PC to obtain the input and output of the vehicle running in the simulation environment;
[0064] S23, connect the control board and the PC using a USB-to-CAN adapter and a CAN-to-USB adapter;
[0065] S24, based on the input and output of the vehicle running in the simulation environment, the Actor-Critic strategy model deployed by the control board is used to verify the effect of the PC-CAN loop model;
[0066] The bench-half-physical ring model verification process includes:
[0067] S31. Deploy the trained Actor-Critic strategy network model on the control board and configure the CAN communication port of the control board.
[0068] S32, configuring the CarSim working environment on the PC to obtain the input and output of the vehicle running in the simulation environment;
[0069] S33, configure the industrial computer to integrate the vehicle signal output by CarSim and the feedback signals of each actuator on the test bench to obtain observation information, and send the observation information to the CAN bus;
[0070] S34, the control board receives the observation information, sends it to the Actor-Critic strategy network model for reasoning, converts the reasoning result into a control quantity signal, and sends it to the CAN bus;
[0071] S35. The industrial computer receives the control quantity signal, solves the control instructions of each actuator on the test bench, and updates the operation status of the current vehicle in the PC CarSim according to the execution status of the instructions on the test bench actuator side, completing a control loop cycle and verifying the effect of the test bench-semi-physical loop model.
[0072] Beneficial effects of the present invention:
[0073] 1. The intelligent chassis integrated control method and system of the present invention integrates the longitudinal, lateral and vertical functional control of the chassis control, and uses data-driven technology to integrate multi-source sensor data of the vehicle to establish a global optimization strategy model. It can simultaneously consider multiple control objectives (such as comfort, stability, energy consumption, etc.) to achieve coordinated work of the chassis system.
[0074] 2. The intelligent chassis integrated control method and system of the present invention uses data-driven fusion control to learn from massive vehicle operation data in simulation, and collaboratively handles the coupling problems of chassis subsystems without the need for complex dynamics solutions and system decoupling. At the same time, the control strategy can be automatically adjusted according to environmental changes to adapt to a variety of usage scenarios (such as slippery roads, ramps, curves, etc.).
[0075] 3. The intelligent chassis integrated control method and system of the present invention can increase or decrease observation information and specific control quantities according to different tasks and application scenarios, support the expansion of the number and type of observation information, support the expansion of chassis actuator-side control quantities, and provide an expandable vector interface for access to chassis RL training.
[0076] 4. The intelligent chassis integrated control method and system of the present invention uses integrated control to improve the handling and stability of the entire vehicle by overall coordination of the chassis subsystems. Especially during intense driving or extreme working conditions, it can break through the upper limit of hardware control through mutual compensation between actuators.
[0077] 5. The intelligent chassis integrated control method and system of the present invention can quickly generate a controller through a large amount of simulation and actual data training, reducing manual intervention and development cycle.
[0078] 6. The intelligent chassis integrated control method and system of the present invention adopts the high inference frequency of the Actor-Critic strategy network to meet the control requirements of different frequencies of the chassis. BRIEF DESCRIPTION OF THE DRAWINGS
[0079] Figure 1 The schematic diagram of the data-driven intelligent chassis integrated control method of the present invention;
[0080] Figure 2 It is a structural schematic diagram of the data-driven intelligent chassis integrated control system of the present invention;
[0081] Figure 3 This is a logic block diagram of the Actor-Critic strategy network model training module of the present invention;
[0082] Figure 4 The flowchart of the Actor-Critic strategy network model algorithm of the present invention;
[0083] Figure 5 It is the PC-CAN ring model verification flow chart of the present invention;
[0084] Figure 6 It is a flow chart of the bench-semi-physical ring model verification of the present invention;
[0085] Figure 7 It is a schematic diagram of the structure of the electronic device of the present invention. DETAILED DESCRIPTION
[0086] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar symbols throughout represent the same or similar elements or elements with the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and cannot be understood as limiting the present invention.
[0087] Existing chassis intelligent integration requires integrating multiple subsystems into a unified control system, which increases the complexity of the system and requires a highly integrated control algorithm to coordinate the work of each subsystem; the chassis control system has poor stability and robustness.
[0088] In view of the above problems, the present invention provides a data-driven intelligent chassis integrated control method. Figure 1 The schematic diagram of the data-driven intelligent chassis integrated control method provided by the present invention includes:
[0089] The vehicle-environment interaction observation information is collected as input, and the trained and verified Actor-Critic strategy network is used to output the desired control quantity signals in the XYZ directions of the actuators of each chassis subsystem to control the desired states of the actuators of each chassis subsystem.
[0090] like Figure 2 As shown, the control system proposed in the present invention collects vehicle-environment interaction observation information as input. The interaction observation information comes from external command signals and internal sensor signals, including driver intention information, intelligent driving domain information, chassis function control information and vehicle status feedback.
[0091] The driver's intention information includes: steering wheel angle, brake pedal pressure, accelerator pedal pressure, etc.; this information represents the driver's expected control behavior for the vehicle.
[0092] Intelligent driving domain information means that the intelligent driving system will provide reference motion control information based on current perception information, vehicle status, etc., including expected reference vehicle speed, reference driving trajectory, reference yaw angle, etc., to help optimize driving decisions and assist in achieving more precise control.
[0093] Chassis function control information refers to receiving chassis function switch instructions, such as anti-lock braking system (ABS), electronic stability program (ESC), automatic emergency braking (AEB), etc., and activating or disabling specific chassis functions according to the vehicle's control requirements or the driver's interactive operations.
[0094] Vehicle state feedback information means that the chassis control model also needs the vehicle's own state feedback information as reasoning input, such as its own sensor signals, chassis subsystem estimation signals, etc.
[0095] Based on the external command signal and the internal sensor signal as the observation information of the domain control, the expected control amount is distributed to each subsystem of the chassis for specific execution through the reasoning of the Actor-Critic strategy network model.
[0096] The specific interactive observation information can be quantified as the vehicle's: wheel speed, wheel cylinder pressure, vehicle yaw angular velocity, vehicle lateral / longitudinal / vertical acceleration, steering wheel angle, vehicle body pitch / sideways / displacement, etc. This information is the feedback signal that can be directly obtained during the actual vehicle control process.
[0097] The interaction observation information collected above is used as the input of the Actor-Critic strategy network model, and the vector representation is:
[0098]
[0099] Among them, σ steer Input for driver's steering wheel angle; It is the longitudinal, lateral and vertical acceleration of the vehicle, obtained by filtering the signal of the vehicle IMU; The vehicle attitude angular velocity signal, the vehicle's yaw, roll and pitch angular velocity signals; is the vehicle attitude angular acceleration signal; It is the vehicle wheel speed signal; It is the estimated signal of the cylinder pressure feedback of the vehicle wheel end; It is the vertical speed signal of the suspension wheel end; It is the signal for estimating the vertical load at the wheel end. According to the control function and application scenario, the observed signal vector retains dimensional margin and can be increased or decreased at will.
[0100] After the input observation information is processed by the Actor-Critic strategy network model, the Actor-Critic strategy network model outputs the expected control quantity signal. The expected control quantity signal is mainly oriented to the three functional subsystems of the steering system, the braking system and the suspension system. Specifically:
[0101] The steering system is mainly responsible for adjusting the steering angle of the wheels according to the driver's intention and the vehicle's driving state, thereby controlling the vehicle's driving direction and acting on the lateral control of the vehicle chassis. The main control quantities involved include: wheel angle compensation and steering torque, etc. In the four-wheel steering system, the distribution of the four-wheel steering angle is also involved.
[0102] The braking system is responsible for slowing down or stopping the movement of the vehicle, ensuring that the vehicle can cope with various working conditions and emergencies during driving, ensuring driving stability, and acting on the longitudinal control of the vehicle chassis. The main control quantity involved is the wheel end brake cylinder pressure.
[0103] The suspension system provides chassis handling stability by absorbing and alleviating the impact of road conditions on the vehicle chassis, and acts on the vertical control of the vehicle chassis. For semi-active suspension and active suspension, the main control variables involved are the suspension spring damping force and stiffness.
[0104] The desired control quantity signal in the XYZ directions of each subsystem actuator is expressed as:
[0105] Y output =[Γ d ,Δp dwheel ,ΔF S ,ΔF D ]
[0106]
[0107] Among them, Γ d represents the desired steering system compensation, Δp dwheel Indicates the expected change in wheel end brake system cylinder pressure, ΔF S represents the expected change in the spring force of the suspension system, which is convenient for simulation environment training and is directly related to the suspension stiffness, ΔF D It represents the expected change in the spring damping force of the suspension system; the expected control quantity signal in the XYZ directions of each subsystem actuator is increased or decreased according to different tasks and application scenarios.
[0108] Therefore, the Actor-Critic strategy network model can be described as: output =f(X input ), which is a complex mapping from input to output. This parameter form can be used for subsequent network model training.
[0109] The intelligent chassis integrated control system proposed in the present invention can increase or decrease the above-mentioned input and output signals according to the actual chassis function control needs. Therefore, the control system has structural adaptability and functional scalability. It can be combined with typical chassis control function application scenarios to give full play to the control characteristics of each chassis subsystem and verify the effectiveness and control performance of the chassis control system.
[0110] For example, on rugged mountain roads or extreme lane-changing driving conditions, the control system can be divided into the following: the steering system directly calculates the yaw moment intervention amount according to the driver's intention and vehicle status to prevent the vehicle from oversteering or understeering; the braking system calculates the wheel-end braking force distribution according to the wheel operating status, wheel yaw angular velocity and driver's intention to assist in vehicle steering and prevent the vehicle from skidding; the suspension system actively adjusts the suspension damping according to the vehicle's lateral and vertical acceleration, compensates for load transfer, provides lateral support for the vehicle's steering dynamics, and improves the vehicle's driving stability under poor road conditions.
[0111] In addition to the observation information collected above, the control system can also learn with the help of privileged information of the simulation environment during the training process of the Actor-Critic strategy network model. This privileged information is difficult to obtain directly in the actual environment, but can be obtained directly in the simulation environment. With the help of this privileged information, the convergence of the network model strategy can be accelerated and the adaptability of the Actor-Critic strategy model to the environment and configuration can be enhanced.
[0112] Privileged information includes: The road adhesion coefficient felt by the tire Vehicle speed signal v, center of mass side slip angle signal β, and the vehicle's desired steering angle σ sent by the upper layer perception or intelligent driving system dsteer wait.
[0113] The Actor-Critic strategy network model of the present invention includes an Actor network and a Critic network, wherein: the Actor network is mainly responsible for generating action strategies and is the main part of the integrated control strategy model; the Critic network is a value network used to guide the updating of the Actor network.
[0114] Specifically, the Actor network is built using a multi-layer perceptron (MLP) neural network, which contains four hidden layers with dimensions of 256, 128, 64, and 32. The input of the Actor network is the direct observation information generated by the interaction between the vehicle and the environment.
[0115] The Critic network uses observation information and privileged information to obtain rewards, generates action value information valuation based on the rewards and calculates the cost to guide the Actor network to perform strategy update training; the Critic network is also built using a multi-layer perceptron (MLP) neural network, which contains four hidden layers with dimensions of 256, 128, 64 and 32 respectively. In addition to the observation information received by the Actor network, the input also includes privileged information, including vehicle speed, sideslip angle of the center of mass, sideslip angular velocity of the center of mass, friction coefficient between the ground and the tire, etc.
[0116] The introduction of privileged information helps the Critic network estimate the value function to better guide the Actor network to update its strategy. The output of the Critic network is the state value, which then calculates the advantage, reward, and loss functions to update the entire Actor-Critic strategy network model.
[0117] The Actor-Critic policy network model training process includes: the Critic network first calculates the valuation of the value function based on the observation information and privileged information, and then calculates the return and advantage, and optimizes the policy network model. Specifically:
[0118] S11, use the current policy network to forward propagate the observation information to obtain action information (action mean, action standard deviation, etc.), value function estimation information and entropy information;
[0119] S12, adjust the learning rate, specify the target KL divergence and calculate the KL divergence difference between the current strategy and the old strategy, and adaptively adjust the learning rate based on the result;
[0120] S13. Calculate the Actor-Critic strategy network loss, including action strategy loss, entropy loss, and estimated value loss;
[0121] S14. Use the optimizer to optimize and update the Actor-Critic strategy network parameters.
[0122] In order to collect a large amount of training data in a simulation environment to guide the training and updating of the strategy model, the present invention requires a simulation environment that is as close as possible to the actual operation of the vehicle, as well as a reliable and stable model training framework. For this purpose, a reinforcement learning simulation training framework based on the CarSim vehicle dynamics simulation environment and the python-Pytorch library is built. The entire model training process can be roughly summarized as follows: Figure 3 As shown:
[0123] The model training framework is divided into two parts: the environment interaction module and the policy network training module.
[0124] The environment interaction module is mainly responsible for collecting the interaction data between the vehicle and the road environment in CarSim, using the output of the policy network as the control quantity to complete the interaction task of the current time step, and update its own state feedback information, thereby generating the observation information input required by the model. These observation information are used to design the reward function on the one hand, and as the input of the policy network module on the other hand.
[0125] The policy network module is mainly responsible for the generation of four-wheel motion control parameters. An Actor-critic policy network model is built inside the policy network module to generate actions based on the policy gradient and produce the control amount for the next time step.
[0126] The reinforcement learning training environment used by the policy network model is built based on CarSim software and Pytorch library. CarSim software is a vehicle dynamics simulation software commonly used in the industry, which is used to simulate and analyze the behavior of cars under different driving conditions. It has the advantages of high-fidelity physical models, accurate and detailed vehicle dynamics models, and rich driving scene resources and road environment resources. In addition, CarSim also provides a variety of API interfaces, which are helpful for combining with the model training environment. The policy model needs to collect massive data in the simulation environment for iterative updates. The closer the simulation environment is to reality, the more it helps to reduce the difficulty of sim2real when deploying the policy. PyTorch is a popular open source machine learning library that is widely used in deep learning projects. The present invention combines the CarSim environment with Pytorch, which can give full play to their respective advantages and complete the model training task.
[0127] First, configure the PyTorch library and its related dependencies in the Python environment, install the CarSim software and the corresponding SDK, and configure the relevant pythonAPI files so that the CarSim simulation can be controlled by Python scripts.
[0128] Then, according to the training task, create or import the vehicle model, environment model and operating parameters in CarSim, and set the relevant physical parameters and initial conditions.
[0129] Then write a Python script and use CarSim's python API to establish a connection with CarSim to achieve data interaction, including obtaining vehicle status information from CarSim and sending control instructions to CarSim.
[0130] Then, a policy model training environment is built based on Pytorch in a Python script. At each time step, the agent infers the control amount through the policy model based on the current vehicle state and environmental observation information. Then, these control instructions are sent to CarSim to update the vehicle state, and the policy model is updated with reference to the interaction data of the current time step, completing a policy iteration. This process is repeated continuously.
[0131] During the training process, the vehicle operation generates a real-time demonstration in CarSim to evaluate the control performance of the intelligent agent under the current strategy, and adjust the model parameters in time as needed to improve the learning efficiency and performance of the policy model.
[0132] The training process of the Actor-Critic strategy network is as follows Figure 4 As shown. Includes:
[0133] Initialization: Initialize training data such as learning iterations numbers, time period step information, and episode length.
[0134] Observation: The state observation information generated by the interaction between the subject and the environment in the previous iteration is passed to the Actor network, and the observation information and privileged information are passed to the Critic network.
[0135] Interaction: Generate actions based on the policy gradient, and then pass them into the entity-environment interaction module to generate new observation information, reward information, and termination flags, etc., and continuously update the status and accumulated rewards within a time segment episode.
[0136] Update: Calculate the updated reward and advantage based on the input of the Critic network. Calculate the policy network loss and optimize the updated policy network weights by maximizing the expected return, maximizing the policy entropy, and minimizing the value function loss.
[0137] Model generation: Save the stage model and the final model based on the training performance, number of iterations, and reward convergence.
[0138] As the core part of the reinforcement learning Actor-Critic policy network model training framework, the design and tuning of the reward function directly affects the performance of the Actor-Critic policy model.
[0139] Designing a reinforcement learning reward function for a vehicle chassis control system requires full consideration of the actual mission objectives, system characteristics, and multi-objective optimization requirements. Chassis control systems usually involve multiple subsystems working together, so the reward function needs to comprehensively consider performance indicators such as vehicle stability, comfort, responsiveness, and energy efficiency.
[0140] Some general reward functions for chassis control are listed below. The reward function can achieve various chassis control functions such as vehicle anti-lock, anti-skid, oversteer adjustment, and reduction of body bumps by comprehensively coordinating and guiding the control amount of each chassis actuator. These functions can be extracted from the reward function separately or used in combination.
[0141] Slippage bonus:
[0142]
[0143] where μ i is the tire-ground adhesion coefficient, The reward is used to adjust the desired change in the wheel end brake cylinder pressure to prevent the wheel from locking and losing control due to excessive wheel slip. Under stable control, the wheel slip rate λ is usually considered. d =0.5, the driver can effectively control the direction and the wheels will not get out of control.
[0144] For a braking request such as ABS or AEB, the expected slip ratio of the effective braking of the wheel is set to λ d =0.2, the slip rate bonus becomes:
[0145]
[0146] When the slip ratio is greater than 0.5, the pressure is reduced and the pressure is increased. When the slip ratio is too small, the pressure is increased and the pressure is reduced. In addition, the wheel end cylinder pressure is rewarded when the slip ratio reaches the expected slip ratio. Especially for the ground or tire with a small adhesion coefficient, the intensity of the anti-lock brake reward should be increased.
[0147] Turn to Tracking Rewards:
[0148] r steering =-exp(-v·|ε steer |)
[0149] where ε steer =σ steer -σ dsteer is the steering wheel error. The exponential function part is to encourage the strategy to reduce the direction angle error and stably track the steering curve. In order to avoid being too sensitive at high speeds, the vehicle speed v is added to the formula.
[0150] Steering-Braking Stability Bonus:
[0151]
[0152] The redundant adjustment reward of the braking system is to adjust the wheel end brake cylinder pressure to reduce the excessive yaw rate when the vehicle turns to avoid the vehicle losing control. The positive and negative relationship between the left and right and yaw rates of the vehicle can be adjusted as defined.
[0153] Steering-Suspension Stability Bonus:
[0154]
[0155] It is a redundancy adjustment reward for the suspension system, which provides lateral support for the vehicle when cornering at high speed, reduces the deviation of the vehicle's center of mass, and corrects excessive yaw rate. When the vehicle is cornering, the outer wheel suspension stiffness is increased to provide lateral support to avoid tail swinging, and the weight of this reward increases with the increase of vehicle speed.
[0156] Vertical Stability Rewards:
[0157]
[0158] Where i is the wheel code and c is a positive constant. For tires with large vertical forces, large spring damping forces are encouraged, large vertical velocities are penalized to ensure wheel end stability, and excessive vertical acceleration of the vehicle is penalized to ensure vehicle stability.
[0159] Horizontal Macro Rewards:
[0160]
[0161] This reward is a macro-control quantity with a small reward weight. It is intended to reduce the vehicle's lateral acceleration, the sideslip angle of the center of mass, and reduce the vehicle instability caused by drastic changes in the body posture.
[0162] Braking Rewards:
[0163] r brake =-|a x |-||v wheel ||-var(v wheel )
[0164] This reward is a functional reward. It is controlled when the vehicle braking function is activated, punishes the vehicle's longitudinal acceleration, regulates the wheel speed to decrease, and at the same time constrains the vehicle vibration caused by the change in wheel speed to improve comfort.
[0165] Action Smoothness Bonus:
[0166] r action =-||Y output -Y last_output ||
[0167] The network model may output discontinuous changes, causing the hardware actuator to be unable to meet the action output, thus constraining the action output amplitude.
[0168] After the Actor-Critic strategy network model is trained, the chassis integrated control system application verification is carried out. First, the control system application verification is carried out using the simulation environment.
[0169] In order to ensure that the trained chassis control strategy can be safely and reliably deployed on actual vehicles, two verification modules are built for the strategy deployment verification work, namely: the PC-CAN loop model verification module for the feasibility of the pure software-based chassis control strategy, and the bench-semi-physical chassis strategy feasibility verification module.
[0170] First, the PC-CAN loop model verification work is introduced in detail:
[0171] like Figure 5As shown in the figure, CAN (Controller Area Network) communication is an efficient and reliable serial communication protocol, which is widely used in the fields of automobiles, industrial automation, etc. for data exchange between controllers. It uses differential signal transmission, has the characteristics of multiple master stations, strong real-time performance and strong anti-interference ability, and ensures the real-time transmission of key data through the message priority mechanism. In chassis control development, the verification of the CAN control loop is an extremely important task before the control algorithm is deployed on the actual vehicle chassis. CAN loop verification can ensure the correctness and stability of communication, avoid system malfunctions due to data loss or errors, and verify whether the controller's processing of different message priorities meets the design requirements, ensuring the real-time performance of key data.
[0172] The main process is as follows: Step 1, deploy the trained chassis control strategy model on a control board with certain computing and processing capabilities and CAN communication function, and configure the CAN communication port of the control board; Step 2, configure the CarSim working environment on the PC according to the needs of functional verification, and set the input and output of the vehicle running in the simulation environment; Step 3, configure the CAN communication port on the PC, integrate the CarSim input and output signals and CAN encoding and decoding signals in the code, encode the vehicle information generated by CarSim, and decode the command information sent by the control board; Step 4, prepare a USB-to-CAN adapter and a CAN-to-USB adapter, or a CAN analyzer, connect the control board and the PC, perform communication tests, adapt the inference frequency of the strategy model in the control board and the communication frequency of the CAN loop, and ensure that the control frequency is consistent with the simulation training; Step 5, verify the control effect of the PC-CAN loop.
[0173] The following is an introduction to the bench-semi-physical loop verification module for the control system:
[0174] The bench-semi-physical chassis strategy feasibility verification module is mainly used to deploy the trained strategy network model on the chassis control verification machine. At the same time, combined with the actual dynamic response of each chassis actuator, the simulated vehicle operation is run in the simulation environment to simulate the actual control effect of each chassis actuator end under the CAN communication framework. The main process of this module is as follows Figure 6 As shown:
[0175] The main process is as follows: Step 1, deploy the trained chassis control strategy model on a control board with certain computing and processing capabilities and CAN communication function, and configure the CAN communication port of the control board; Step 2, configure the CarSim working environment on the PC according to the needs of functional verification, and set the input and output of the vehicle running in the simulation environment; Step 3, configure the industrial computer, integrate the vehicle signal output by CarSim and the feedback signal of each actuator on the test bench, generate the state observation output required for chassis control, and send the observation signal to the CAN bus; Step 4, the control board receives the input signal from the industrial computer from the CAN bus, sends it to the strategy model for reasoning, converts the reasoning result into a control quantity, and encodes it and sends it to the CAN bus; Step 5, after the industrial computer receives the CAN control signal from the control board, it solves the control instructions of each actuator on the test bench, and updates the current vehicle operation status in CarSim according to the execution of the instructions on the test bench actuator side, completing a control loop cycle; Step 6, complete the strategy control verification under the test bench-semi-physical loop working condition.
[0176] Different from PC-CAN loop verification, bench-semi-physical loop verification involves real mechanical and electronic hardware (such as ECU, sensors, actuators), which can directly reflect the dynamic response characteristics of the chassis control system, such as suspension vibration, brake lag time, etc. Bench-semi-physical loop control verification is closer to actual applications, can fully expose problems in software and hardware integration, and ensure the reliability and stability of the system under real working conditions.
[0177] The present invention also provides a data-driven intelligent chassis integrated control system, such as Figure 1 and Figure 2 As shown, the system includes:
[0178] Data acquisition module, collecting vehicle-environment interaction observation information;
[0179] The Actor-Critic strategy network module deploys the trained and verified Actor-Critic strategy network model, collects vehicle-environment interaction observation information as input, and outputs the desired control quantity signals in the XYZ directions of the actuators of each chassis subsystem;
[0180] The signal output module controls the desired state of the actuators of each chassis subsystem according to the control signal.
[0181] The present invention trains the Actor-Critic strategy network model in the simulation space, and then can deploy the control model to the vehicle chassis entity, thereby completing the intelligent control of the chassis in the XYZ three-way fusion.
[0182] The chassis integrated control system of the present invention has the following design and technical aspects: (1) The difficulty of developing control strategies is reduced: there is no need to solve the complex dynamic model of the vehicle chassis and the coupling relationship of the XYZ three-way control. It is only necessary to clarify the input and output under the fusion control framework, integrate the sensor feedback signals, estimation signals and control signals of the vehicle, driver and chassis to perform three-way fusion control of the chassis; (2) The difficulty of debugging control strategies is reduced, and the optimization and iteration process is simple and effective: only the simulator development is required to train the strategy network, and when the strategy is deployed to the vehicle chassis entity, the strategy training is continuously optimized and iterated to achieve the actual control effect; (3) Only one set of trained control strategies is required to adapt to the changing functional requirements and road conditions of the vehicle chassis: the system mainly builds an asymmetric Actor-critic strategy network and designs a set of reward functions that can cover most functional scenarios of the chassis; (4) The algorithm reasoning complexity is extremely low, and the reasoning frequency can be as high as 100HZ: in the model reasoning stage, the structure of the pre-trained model Actor network is simple, so the algorithm complexity is low, the reasoning frequency is fast, and the entity can be deployed and the control quantity can be solved in real time.
[0183] In terms of control effect, (1) XYZ three-way fusion control can improve the vehicle's maneuverability and stability: through the mutual coordination between the chassis actuators, the hardware upper limit of the single actuator function control can be broken, and the vehicle's passability under extreme working conditions, especially the cornering ability and tracking ability, can be improved; (2) Give full play to the integration advantages of the control system: the centrally integrated chassis control system can simplify the system structure, reduce cable connections, improve system stability and reliability, and the introduction of the control system makes the chassis control more professional and efficient, and can better adapt to different driving scenarios and road conditions. (3) Improve the intelligence of the whole vehicle and promote the technical integration of the whole vehicle: the chassis control system provides a control interface with the intelligent driving domain and the cockpit domain, which can more conveniently realize the combination of upper-level commands and underlying control commands, which is conducive to the improvement of the whole vehicle's autonomous driving technology.
[0184] The present invention also provides an electronic device, Figure 7 A schematic diagram of the structure of an electronic device provided by an embodiment of the present invention, such as Figure 7 As shown, the electronic device may include: a processor, a communications interface, a memory, and a communications bus, wherein the processor, the communications interface, and the memory communicate with each other via the communications bus. The processor may call the logic instructions in the memory, for example, to execute the following method:
[0185] The vehicle-environment interaction observation information is collected as input, and the trained and verified Actor-Critic strategy network is used to output the desired control quantity signals in the XYZ directions of the actuators of each chassis subsystem to control the desired states of the actuators of each chassis subsystem.
[0186] In addition, the logic instructions in the above-mentioned memory can be implemented in the form of software functional units and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on such an understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art or the part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.
[0187] An embodiment of the present invention further provides a non-transitory computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the method provided in each of the above embodiments is implemented, for example, including:
[0188] The vehicle-environment interaction observation information is collected as input, and the trained and verified Actor-Critic strategy network is used to output the desired control quantity signals in the XYZ directions of the actuators of each chassis subsystem to control the desired states of the actuators of each chassis subsystem.
[0189] The system embodiment described above is merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art may understand and implement it without creative work.
[0190] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0191] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A data-driven intelligent chassis integrated control method, characterized in that: include: The vehicle-environment interaction observation information is collected as input, and the trained and verified Actor-Critic strategy network is used to output the desired control quantity signals in the XYZ directions of the actuators of each chassis subsystem to control the desired states of the actuators of each chassis subsystem.
2. The control method according to claim 1, characterized in that: The collected vehicle-environment interaction observation information is used as input, and the vector is represented as: Among them, σ steer Input for driver's steering wheel angle; are the longitudinal, lateral and vertical accelerations of the vehicle, The yaw, roll and pitch angular velocity signals of the vehicle; is the vehicle attitude angular acceleration signal; It is the vehicle wheel speed signal; It is the estimated signal of the cylinder pressure feedback of the vehicle wheel end; It is the vertical speed signal of the suspension wheel end; Estimation signal for wheel end vertical load; Among them, the interactive observation information increases or decreases according to different tasks and application scenarios.
3. The control method according to claim 1, characterized in that: The chassis subsystems include: steering system, braking system and suspension system. The desired control quantity signals of the actuators of each subsystem in the XYZ directions are expressed as vectors: Y output =[Γ d ,Δp dwheel ,ΔF S ,ΔF D ] Among them, Γ d represents the desired steering system compensation, Δp dwheel Indicates the expected change in wheel end brake system cylinder pressure, ΔF S Denotes the desired change in suspension system spring force, ΔF D Indicates the desired change in the spring damping force of the suspension system; Among them, the desired control quantity signals in the XYZ directions of each subsystem actuator are increased or decreased according to different tasks and application scenarios.
4. The control method according to claim 1, characterized in that: The Actor-Critic strategy network includes an Actor network and a Critic network, wherein: the Actor network is responsible for generating action strategies, the Critic network uses the observation information and privileged information obtained through interaction to obtain rewards, generates action value information valuation based on the rewards and calculates the cost to guide the Actor network to perform strategy update training.
5. The control method according to claim 4, characterized in that: The privileged information is directly obtained from the simulation environment, including: the road adhesion coefficient felt by the tire Vehicle speed signal v, center of mass sideslip angle signal β and vehicle desired steering angle σ dsteer .
6. The control method according to claim 4, characterized in that: The rewards include: Slippage rate bonus, the formula is: Among them, μ i is the tire-ground adhesion coefficient, is the wheel slip rate; Turning to tracking rewards, the formula is expressed as: r steering =-exp(-v·|ε steer |) Among them, ε steer =σ steer -σ dsteer is the steering wheel steering error; Steering-brake stability reward, the formula is expressed as: in, and is the excessive yaw rate when the vehicle turns, Δp i is the cylinder pressure of the wheel end brake system; Steering-suspension stability bonus, the formula is expressed as: in, is too large yaw rate; Vertical stability reward, the formula is expressed as: Among them, i represents the wheel code, and c represents a positive constant; Indicates a tire with large vertical force, a y z i Excessive vertical acceleration; Horizontal macro rewards, the formula is expressed as: Among them, |a y |lateral acceleration, |β| sideslip angle of center of mass; Braking reward, the formula is expressed as: r brake =-|a x |-||v wheel ||-var(v wheel ) Among them, |a x | is the longitudinal acceleration; Action smoothness reward, the formula is expressed as: r action =-||Y output -Y last_output || Among them, ||Y output -Y last_output || represents the output amplitude of the constraint action.
7. The control method according to claim 6, characterized in that: When the vehicle uses ABS or AEB braking, the slip rate bonus is expressed as follows: Among them, λ d =0.2 is the expected slip rate of effective wheel braking. When the reward slip rate is greater than 0.5, the pressure is reduced. When the reward slip rate is less than 0.05, the pressure is increased. i | is the wheel end cylinder pressure.
8. The control method according to claim 4, characterized in that: The Actor-Critic strategy network training process includes: S11. Use the current policy network to forward propagate the observation information to obtain action information, value function estimation information and entropy information; S12, adjust the learning rate, specify the target KL divergence and calculate the KL divergence difference between the current strategy and the old strategy, and adaptively adjust the learning rate based on the result; S13. Calculate the Actor-Critic strategy network loss, including action strategy loss, entropy loss, and estimated value loss; S14. Use the optimizer to optimize and update the Actor-Critic strategy network parameters.
9. A control system based on the control method according to any one of claims 1 to 8, characterized in that: include: Data acquisition module, collecting vehicle-environment interaction observation information; The Actor-Critic strategy network module deploys the trained and verified Actor-Critic strategy network model, collects vehicle-environment interaction observation information as input, and outputs the desired control quantity signals in the XYZ directions of the actuators of each chassis subsystem; The signal output module controls the desired state of the actuators of each chassis subsystem according to the control signal.
10. The control system according to claim 9, characterized in that: The control system is verified by PC-CAN loop model and bench-semi-physical loop model, where: The PC-CAN ring model verification process includes: S21. Deploy the trained Actor-Critic strategy network model on the control board and configure the CAN communication port of the control board. S22, configuring the CarSim working environment on the PC to obtain the input and output of the vehicle running in the simulation environment; S23, connect the control board and the PC using a USB-to-CAN adapter and a CAN-to-USB adapter; S24, based on the input and output of the vehicle running in the simulation environment, the Actor-Critic strategy model deployed by the control board is used to verify the effect of the PC-CAN loop model; The bench-half-physical ring model verification process includes: S31. Deploy the trained Actor-Critic strategy network model on the control board and configure the CAN communication port of the control board. S32, configuring the CarSim working environment on the PC to obtain the input and output of the vehicle running in the simulation environment; S33, configure the industrial computer to integrate the vehicle signal output by CarSim and the feedback signals of each actuator on the test bench to obtain observation information, and send the observation information to the CAN bus; S34, the control board receives the observation information, sends it to the Actor-Critic strategy network model for reasoning, converts the reasoning result into a control quantity signal, and sends it to the CAN bus; S35. The industrial computer receives the control quantity signal, solves the control instructions of each actuator on the test bench, and updates the operation status of the current vehicle in the PC CarSim according to the execution status of the instructions on the test bench actuator side, completing a control loop cycle and verifying the effect of the test bench-semi-physical loop model.
Citation Information
Patent Citations
Intelligent vehicle chassis and task load integrated control method and system
CN115657645A
Cited By
Automatic driving multi-dimensional coupling cooperative control method and system considering passenger comfort
CN121893938A
Automatic driving multi-dimensional coupling cooperative control method and system considering passenger comfort
CN121893938B