Control method considering safety of intelligent suspension, medium and equipment
By introducing a multi-dimensional reward mechanism and adaptive hyperparameter adjustment in the intelligent suspension system, the problem of neglecting the actuator dynamics and physical constraints is solved, and the efficient, safe and stable control effect of the intelligent suspension system is achieved.
Patent Information
- Application Number
- CN202510775059.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-06-11
AI Technical Summary
The dynamics and time-delay characteristics of the actuator are ignored in the control of existing intelligent suspension systems, and physical constraints and actual limitations are not fully considered, which makes it difficult for the control strategy to achieve the expected results in actual applications, and may even cause system instability and security threats.
By initializing the evaluation network, control policy network and experience pool, combining multi-dimensional reward mechanisms for acceleration, dynamic stroke and tire dynamic load, hard constraints and delay mechanisms are introduced, adaptive hyperparameter adjustment and fatigue constraints are optimized to enhance robustness and adaptability.
It significantly improves the reliability and performance stability of the intelligent suspension system, enhances the feasibility of safety control and practical application, shortens the learning cycle, and improves the system efficiency and iteration speed.
Smart Images

Figure CN120269980A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of automobile dynamics control, and in particular relates to a control method, medium and equipment taking the safety of intelligent suspension into consideration. Background Art
[0002] With the vigorous development of artificial intelligence technology, autonomous driving systems and related fields are facing unprecedented opportunities for technological innovation. Compared with traditional methods, artificial intelligence technology provides a new perspective and solution to solve the control problems of key components such as suspension systems in autonomous vehicles. As a core component for improving ride comfort and driving performance, the optimization of the suspension system's performance is crucial to the overall driving experience and functional safety of autonomous vehicles.
[0003] The shortcomings of the prior art are mainly reflected in the following aspects: (1) Ignoring the actuator dynamics and time-delay characteristics: In the control of intelligent suspension systems, the actuator dynamics and time-delay effects have a crucial impact on the control effect. However, many scholars often ignore these key factors in the application of deep reinforcement learning, which makes it difficult for the control strategy to achieve the expected effect in practical applications, and may even cause system instability; (2) Failure to fully consider physical constraints and practical limitations: The deployment of deep reinforcement learning in intelligent suspension control faces many practical limitations and physical constraints. If these constraints are ignored, the intelligent agent may produce behaviors that do not conform to the laws of physics, which will not only fail to effectively control the system, but may also pose a serious threat to the safety of the vehicle. Summary of the invention
[0004] In view of this, the present invention aims to provide a control method, medium and device that take the safety of intelligent suspension into consideration, so as to strictly constrain and punish a variety of unsafe situations with comprehensive rewards, so as to effectively manage safety risks; by combining the reward mechanism for different situations, not only the robustness of the intelligent suspension control strategy is enhanced, but also its reliability and adaptability in complex and changing environments are significantly improved.
[0005] To achieve the above object, the technical solution created by the present invention is implemented as follows: A control method considering the safety of an intelligent suspension, comprising: S1: Initialize the first evaluation network, the second evaluation network, the control strategy network and the experience pool; S2: The control strategy network determines the control force output by the intelligent suspension at the current moment based on the vehicle state at the current moment; applies the control force at the current moment to the vehicle to obtain the vehicle state at the next moment; the vehicle state at the current moment includes the acceleration and speed of the vehicle during operation at the current moment, the dynamic stroke of the intelligent suspension, and the speed difference between the sprung mass and the unsprung mass of the vehicle. S3: Determine the current comprehensive reward according to the control force at the current moment and the vehicle state at the current moment; save the vehicle state at the current moment, the control force at the current moment, the current comprehensive reward, and the vehicle state at the next moment as a group of samples in the experience pool; the current comprehensive reward includes the corresponding current acceleration reward, current dynamic stroke reward, and current tire dynamic load reward corresponding to the tire dynamic load of the vehicle. S4: Repeat steps S2 - S3 multiple times, and correspondingly save multiple groups of samples in the experience pool; use multiple groups of samples to train the first evaluation network, the second evaluation network, and the control strategy network to obtain the first evaluation model, the second evaluation model, and the control strategy model. S5: Run the intelligent suspension and use the first evaluation model, the second evaluation model, and the control strategy model obtained in step S3 to control the intelligent suspension.
[0006] Further, in step S3: The current acceleration reward is: ; Wherein, represents the acceleration reward at time t, represents the acceleration of the vehicle at time t, , , and represent the acceleration reward coefficients, , and represent the acceleration thresholds; The current dynamic stroke reward is: ; Wherein, represents the dynamic stroke reward at time t, represents the dynamic stroke of the intelligent suspension at time t, , represents the stroke of the vehicle at the current moment, represents the stroke of the unsprung mass of the vehicle at time t; , and represent the dynamic stroke reward coefficients, and represent the dynamic stroke thresholds; The current tire dynamic load reward is: ; Among them, represents the tire dynamic load reward at time t, represents the wheel dynamic load of the vehicle at time t, , represents the road surface height excitation information of the vehicle at time t, 、 and represent the tire dynamic load reward coefficient, 、 、 and represent the tire dynamic load threshold, represents the intermediate variable in the range calculation of the current tire dynamic load reward, which is related to time t and is: ; Among them, represents the tire stiffness of the vehicle at time t, represents the body mass of the vehicle, represents the unsprung mass, and g represents the acceleration due to gravity; The current comprehensive reward is: ; Among them, represents the current comprehensive reward, 、 and respectively represent the current acceleration reward 、the current dynamic stroke reward and the current tire dynamic load reward weights.
[0007] Furthermore, step S2 also includes: Hard constraint on the control force output by the control policy network: Determine the control force at the current moment by using the delay mechanism according to the hard-constrained control force; Detect the fatigue degree of the control force at the current moment and adjust the control force at the current moment according to the fatigue degree.
[0008] Furthermore, the hard constraint is performed by the following formula: ; Among them, represents the control force after hard constraint on the control force output by the control policy network at time t, represents the control policy network, represents the network weights of the control policy network, represents the truncation function, and respectively represent the minimum and maximum values of the control force at the current moment, represents the state of the vehicle at time t.
[0009] Furthermore, the process of determining the control force at the current moment based on the control force after the hard constraint is as follows: ; where, represents the control force output by the intelligent suspension at time t, represents the delay time, represents at the control force output by the intelligent suspension at the moment; delay time is determined by the following formula: ; where, represents the unit quantity determined by the maximum limit control force of the intelligent suspension, represents the unit time delay.
[0010] Furthermore, in the process of detecting the fatigue degree of the control force at the current moment and adjusting the control force at the current moment according to the fatigue degree: the working intensity of the intelligent suspension is calculated by the following formula: ; where, represents the working intensity; the fatigue degree is calculated by the following formula: ; where, represents the fatigue degree, and represents the fatigue coefficient, represents the working time of the intelligent suspension; Set the fatigue degree threshold. If the fatigue degree at the current moment exceeds the fatigue degree threshold, reduce the control force at the current moment.
[0011] Furthermore, in step S4: In each time step of training, multiple groups of samples are randomly selected from the experience pool , and the corresponding target value is calculated by the following formula: ; where, represents the target value, represents the discount factor, represents the jth evaluation network, represents the network weight of the jth evaluation network, represents the state of the vehicle at time t, Denotes the control force output by the intelligent suspension at time t, Denotes the control policy network, Denotes the network weights of the control policy network; Calculate and train the loss functions of the two evaluation networks by combining with the target value: ; Among them, Denotes the loss function of the j-th evaluation network, N denotes the time step; train the two evaluation networks using the loss functions of the two evaluation networks; Train the control policy network through the following formula: .
[0012] Furthermore, during the training of the control policy network and the two evaluation networks: use the stochastic gradient descent method to train the two evaluation networks; use the stochastic gradient ascent method to train the control policy network; during the training of the control policy network and / or the two evaluation networks, every time a set number of time steps pass, evaluate the current training state and adjust the learning rate of the training according to the evaluation results; the learning rate is adjusted within the range of [0.0001, 0.01].
[0013] A readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps of the control method considering the safety of the intelligent suspension provided by the present invention.
[0014] An electronic device includes: A memory for storing a computer program; A processor for implementing the steps of the control method considering the safety of the intelligent suspension provided by the present invention when executing the computer program.
[0015] Compared with the prior art, the present invention can achieve the following beneficial effects: (1) The control method considering the safety of the intelligent suspension of the present invention enhances the safety control and practical application feasibility: by introducing an adaptive hyperparameter adjustment strategy and fatigue degree constraint, the present invention more accurately integrates actuator dynamics constraints in the deep reinforcement learning algorithm. The adaptive hyperparameter adjustment enables the model to quickly adapt to different working conditions during training and optimize the control strategy; the actuator fatigue degree constraint effectively avoids the failure of the actuator due to overuse, ensuring the safety and feasibility of the decision-making process, greatly enhancing the reliability of the control strategy in practical engineering applications, deeply verifying its control performance, and providing a solid guarantee for large-scale actual deployment; (2) The control method considering the safety of the intelligent suspension in the present invention improves the system efficiency and iteration speed: the dynamic weight allocation of multi-scenario rewards and the adaptive hyperparameter adjustment strategy work together, significantly accelerating the iteration process of the intelligent suspension in different driving scenarios. The dynamic weight allocation of multi-scenarios makes the reward function more in line with the actual needs and finds the optimal strategy faster; the adaptive hyperparameter adjustment optimizes the model training process, effectively shortening the learning cycle and improving the overall efficiency of the system, providing strong support for the rapid response and precise control of the intelligent suspension system; (3) The control method considering the safety of the intelligent suspension in the present invention improves the system reliability and performance stability: the combination of the control force constraint and the dynamic weight allocation of multi-scenario rewards fully considers the characteristics of the intelligent suspension and the vehicle driving state; the fatigue constraint and the optimized delay mechanism ensure the stable operation of the actuator, and the multi-dimensional state perception provides accurate information for the control strategy, enabling the system to maintain excellent performance and high safety in various complex road conditions and driving scenarios, significantly improving the reliability and performance stability of the system, and laying a solid foundation for the long-term stable operation of the intelligent suspension system. Brief Description of the Drawings
[0016] The drawings constituting a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings: Figure 1 It is a schematic flow chart of the control method considering the safety of the intelligent suspension according to the embodiment of the present invention; Figure 2 It is a schematic framework diagram of the control method considering the safety of the intelligent suspension according to the embodiment of the present invention; Figure 3 It is a schematic diagram of the control strategy network according to the embodiment of the present invention; Figure 4 It is a schematic diagram of the two evaluation networks according to the embodiment of the present invention; Figure 5 It is a schematic diagram of the dynamic model of the intelligent suspension system according to the embodiment of the present invention; Figure 6 It is a schematic diagram of the structure of the electronic device according to the embodiment of the present invention.
[0017] Description of the Reference Numerals: 1. Electronic device; 2. External device; 3. Processing unit; 4. Bus; 5. Network adapter; 6. Display; 7. (I / O) interface; 8. System memory; 9. Random access memory; 10. Cache memory; 11. Storage system; 12. Utility tool; 13. Program module. Detailed Embodiments
[0018] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and do not constitute a limitation to the present invention.
[0019] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.
[0020] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "center", "longitudinal", "transverse", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the accompanying drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to the present invention. In addition, the terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first", "second", etc. may explicitly or implicitly include one or more of such features. In the description of the present invention, unless otherwise specified, the meaning of "a plurality" is two or more.
[0021] In the description of the present invention, it should be noted that unless otherwise clearly defined and limited, the terms "installed", "connected", "connected" should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral connection; it may be a mechanical connection or an electrical connection; it may be directly connected, or indirectly connected through an intermediate medium, and it may be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood through specific situations.
[0022] The present invention will be described in detail below with reference to the accompanying drawings and in conjunction with the embodiments.
[0023] As Figures 1 to 2 shown, the control method considering the safety of the intelligent suspension according to the embodiment of the present invention includes: S1: Initialize the first evaluation network, the second evaluation network, the control strategy network and the experience pool.
[0024] In a certain embodiment, the control strategy network is as Figure 3As shown in the figure, the state of the vehicle is input into a fully connected layer composed of 400 neurons for processing. The processed data is then processed by the ReLU activation function and input into a fully connected layer composed of 300 neurons; the output data is processed by the ReLU activation function and then input into a fully connected layer composed of 1 neuron; finally, the output data is processed by the activation operation of the Tanh activation layer and the data scaling operation of the scaling layer to obtain the control force output by the intelligent suspension. The structures of the first evaluation network and the second evaluation network are the same. As Figure 4 shown in the figure, the state of the vehicle is input into a fully connected layer composed of 128 neurons for processing. The processed data is processed by the ReLU activation function and then input into a fully connected layer composed of 200 neurons to obtain the first feature; at the same time, the control force output by the intelligent suspension is input into a fully connected layer composed of 200 neurons to obtain the second feature. After adding the corresponding elements of the first feature and the second feature, it is processed by the ReLU activation function and then input into a fully connected layer composed of 1 neuron to obtain the evaluation result. The initialization operations include: for the network weights in the first evaluation network and the network weights in the first evaluation network , as well as the network weights of the control strategy network are randomly initialized. s represents the state of the vehicle, and a represents the control force output by the intelligent suspension.
[0025] S2: The control strategy network determines the control force of the intelligent suspension at the current moment according to the state of the vehicle at the current moment; the control force at the current moment is applied to the vehicle to obtain the state of the vehicle at the next moment. The state at the current moment includes the acceleration and speed of the vehicle running at the current moment, as well as the dynamic stroke of the intelligent suspension and the speed difference between the sprung mass and the unsprung mass of the vehicle. In the present invention, the acceleration is taken as an example of the vertical acceleration of the vehicle body.
[0026] In some embodiments, the system dynamics model of the intelligent suspension is as Figure 5 shown in the figure, Figure 5 where represents the body mass of the vehicle, represents the unsprung mass, represents the vertical displacement of the vehicle body at time t, represents the stroke of the unsprung mass of the vehicle at time t, represents the road surface height excitation information of the vehicle at time t, represents the tire stiffness of the vehicle, represents the suspension spring stiffness of the intelligent suspension, represents the damping coefficient of the intelligent suspension Represents the damping coefficient of the vehicle tire. Based on the system dynamics model of the intelligent suspension, the current state of the vehicle can be obtained , where and respectively represent the acceleration and velocity of the sprung mass in the vehicle at time t, which are important parameters for the ride comfort performance of the vehicle suspension. Represents the dynamic stroke of the intelligent suspension at time t, which is an important physical constraint in suspension control. Represents the velocity difference between the sprung mass and the unsprung mass at time t. Since the power of the intelligent suspension is P = F×v, where F represents the output force of the intelligent suspension and v is the velocity difference , the velocity difference v to some extent characterizes the physical constraints on the performance of the intelligent suspension. Among them, the velocity difference , represents the unsprung velocity of the vehicle at time t. The parameters in the current state guarantee the practical feasibility of deploying the model proposed in the present invention to the intelligent suspension from the real-world limitations. The present invention uses a state feedback law. Therefore, its performance depends on the accuracy and proper selection of the measured or estimated states. In one embodiment, in the control implementation of the intelligent suspension, the velocity is estimated by high-pass filtering and integrating the vertical acceleration , the dynamic stroke is obtained by directly measuring the displacement of the actuator of the intelligent suspension, and the velocity difference is calculated by differentiating the dynamic stroke using the hybrid smooth derivative method. In practical applications and deployments, only the values of the above variables need to be obtained.
[0027] S3: Determine the current comprehensive reward according to the current control force and the current state; save the current state, the current control force, the current comprehensive reward, and the next state as a set of samples in the experience pool.
[0028] In the application of deep reinforcement learning to the control of intelligent suspension systems, ensuring the physical safety of the operation of intelligent suspension systems is of crucial importance. Specifically, the present invention imposes physical safety constraints on three key state variables: vehicle acceleration, dynamic stroke, and tire dynamic load. The constraints on vehicle body acceleration and dynamic stroke ensure that the vehicle maintains passenger comfort and prevents safety accidents caused by excessive vibration or tilting during sudden road changes. The constraints on tire dynamic load can prevent tire damage or blowout due to overload, ensuring stability and safety during driving. In view of these considerations, the present invention refines the reward function under the guidance of physical safety constraints to minimize the impact of training termination caused by any state variable exceeding its safety limit. That is, the present invention proposes rewards for vehicle acceleration, dynamic stroke, and tire dynamic load, and combines the three rewards to achieve multi-dimensional state perception. Specifically, the current comprehensive reward includes the corresponding current acceleration reward, current dynamic stroke reward, and current tire dynamic load reward corresponding to the tire dynamic load of the vehicle.
[0029] In some embodiments, hierarchical reward constraints are imposed on acceleration. That is, according to the comfort threshold, tolerance threshold, and safety threshold, its rewards are divided into four layers. When the safety threshold is exceeded, the training episode is terminated and a high penalty is imposed to guide the agent away from behaviors that pose safety risks. Specifically, the current acceleration reward is: ; Where represents the acceleration reward at time t, , , and represent the acceleration reward coefficients, and these values are adaptively selected according to actual situations and human experience. , and represent the acceleration thresholds. It should be noted that the acceleration reward coefficient should be much smaller than the acceleration reward coefficients , and to ensure that when the acceleration exceeds the maximum acceleration threshold, there is a very small acceleration reward, that is, a large penalty is imposed. If the acceleration reward value at time t is the acceleration reward coefficient , the training is aborted. The determination process of the acceleration thresholds , and includes: Control the vehicle equipped with intelligent suspension to drive at different driving speeds, different road conditions, and different vehicle load conditions, collect acceleration data during vehicle driving, classify and count the acceleration data according to the value size, and draw a probability distribution histogram of the acceleration value. According to the distribution of the probability distribution histogram, determine the acceleration range covered by 50%, 80%, and 100% data, and obtain the corresponding acceleration thresholds respectively. , and . In one embodiment, the driving speed is set in the range of 10km / h to 120km / h, the road conditions include flat roads, rough roads, and speed bump roads, and the vehicle load conditions include no-load (the vehicle only contains the driver and necessary equipment), half-load (50% of the rated load of the vehicle) and full-load (the vehicle reaches the rated load limit). The acceleration data of the vehicle is collected during driving, and the acceleration data is classified and counted according to the numerical value, and a probability distribution histogram of the acceleration value is drawn. Determine the acceleration ranges covered by 50%, 80%, and 100% of the data, and obtain the corresponding acceleration thresholds respectively. , and The specific process is as follows: Integrate the probability distribution histogram to find the acceleration interval containing 50%, 80%, and 100% of the data, and obtain the corresponding acceleration threshold , and In addition, the acceleration bonus coefficient , , and They are -10, -20, -30 and -2000 respectively.
[0030] Taking into account vehicle design and safety margin, the safety range of the dynamic travel of the smart suspension should be set to no more than 80% of the vehicle's maximum design suspension travel to prevent damage to the suspension system under extreme conditions. The corresponding current dynamic travel reward is: ; in, represents the dynamic travel reward at time t, , and Indicates the dynamic travel reward coefficient, which is selected based on actual conditions and human experience. and Indicates the dynamic travel threshold. It should be noted that the dynamic travel reward coefficient Much smaller than the dynamic itinerary reward factor and , to ensure that when the dynamic stroke exceeds the maximum dynamic stroke threshold, there is a very small dynamic stroke reward, that is, a very large penalty is imposed. That is, if the dynamic stroke reward value at time t is the dynamic stroke reward coefficient , terminate the training. The dynamic stroke threshold and are obtained in the same way as the acceleration threshold , and . That is, control the vehicle equipped with an intelligent suspension to drive in different environments (different driving speeds, different road surface conditions, and different vehicle load conditions), collect the dynamic stroke data during the vehicle driving process, classify and statistically analyze the dynamic stroke data according to the numerical size, and draw the probability distribution histogram of the dynamic stroke value. According to the distribution of the probability distribution histogram, determine the acceleration ranges covered by 100% + 20% (that is, adding 20% of the probability distribution histogram to the entire probability distribution histogram) and 100% + 50% (that is, adding 50% of the probability distribution histogram to the entire probability distribution histogram) data, and respectively obtain the dynamic stroke thresholds and . In a certain embodiment, the driving environment of the vehicle when obtaining the dynamic stroke threshold is the same as that of the acceleration threshold, and the corresponding acceleration thresholds and are 0.12 and 0.15 respectively. In addition, set the dynamic stroke reward coefficients , and to be -10, -20 and -100 respectively.
[0031] The load of the vehicle refers to the maximum weight that the vehicle can safely support as specified by the manufacturer, including the weight of the vehicle itself, passengers and goods. When designing the load-bearing capacity of the wheels and tires, both static load and dynamic load are considered. Dynamic load refers to the additional forces applied to the tires during vehicle movement, and these forces are caused by conditions such as acceleration, emergency braking, or driving on an uneven surface. To prevent wheel slip or tire damage under these conditions, under normal driving conditions, the total load-bearing capacity of the wheels should neither be lower than 60% of the vehicle design load nor exceed 150%. The corresponding current tire dynamic load reward is: ; wherein, represents the tire dynamic load reward at time t, represents the wheel dynamic load of the vehicle at time t, , , and represent the tire dynamic load reward coefficients, and these values are adaptively selected according to the actual situation and human experience. , , and represent the tire dynamic load threshold, represents an intermediate variable in the calculation of the current tire dynamic load reward range, which is related to time t and is: ; wherein, represents the tire stiffness of the vehicle at time t, represents the body mass of the vehicle, represents the unsprung mass, and g represents the acceleration due to gravity.
[0032] It should be noted that the tire dynamic load reward coefficient should be much smaller than the dynamic stroke reward coefficient and so as to ensure that when the tire dynamic load exceeds the range of the tire dynamic load threshold, there is a very small tire dynamic load reward, that is, a large penalty is imposed. That is, if the tire dynamic load reward value at time t is the tire dynamic load reward coefficient , the training is aborted.
[0033] The tire dynamic load threshold , , and are obtained in the same way as the acceleration threshold , and , that is, controlling a vehicle equipped with an intelligent suspension to drive in different environments (different driving speeds, different road surface conditions, and different vehicle load conditions), collecting the dynamic stroke data during the vehicle driving process, classifying and statistically analyzing the dynamic stroke data according to the numerical size, and drawing a probability distribution histogram of the dynamic stroke value. According to the distribution of the probability distribution histogram, determine the acceleration ranges covered by 80%, 100% + 20%, 30% and 100% + 50% of the data, and respectively obtain the tire dynamic load thresholds , , and . In a certain embodiment, the driving environment of the vehicle is the same as the acceleration threshold when obtaining the tire dynamic load threshold, and the corresponding tire dynamic load thresholds , , and are 0.8, 1.2, 0.3 and 1.5 respectively. In addition, the tire dynamic load reward coefficients , and are -10, -20 and -2000 respectively.
[0034] Current acceleration reward , current dynamic stroke reward and current tire dynamic load reward , the current comprehensive reward is obtained as follows: ; Wherein, represents the current comprehensive reward, , and respectively represent the weights of the current acceleration reward , current dynamic stroke reward and current tire dynamic load reward . In practical applications, the weights are dynamically adjusted according to different driving scenarios and vehicle driving states. When the vehicle frequently starts and stops on urban roads, is adjusted to 0.6, is adjusted to 0.15, is adjusted to 0.25, focusing on the impact of vehicle body acceleration on ride comfort; when the vehicle is driving on off-road roads, is adjusted to 0.3, is adjusted to 0.35, is adjusted to 0.35, strengthening the constraints of tire dynamic load and suspension deflection to ensure vehicle passability and safety.
[0035] In some embodiments, for the design of the control force output by the intelligent suspension, the requirements for the safety of autonomous vehicles and the physical limitations brought by the power of the actuator need to be considered, so as to reduce related risks. The present invention realizes imposing higher penalties on dangerous states and behaviors that may endanger system safety through a hard constraint mechanism combined with a control force constraint. Specifically, step S2 further includes: S21: Perform a hard constraint on the control force output by the control strategy network. In some embodiments, the hard constraint is performed by the following formula: ; Wherein, represents the control force after performing a hard constraint on the control force output by the control strategy network at time t, represents the control strategy network, represents the network weights of the control strategy network, represents the truncation function, and respectively represent the minimum and maximum values that limit the control force at the current moment. The minimum value and the maximum value of the control force at the current moment are obtained according to the limitation of the maximum force of the actuator. In one embodiment, the minimum value and the maximum value Take values of -3 kN and 3 kN respectively.
[0036] S22: Determine the control force at the current moment using the delay mechanism based on the control force after hard constraints. Since in the control system, the time delay of the intelligent suspension is often closely related to its actual actuation ability, the present invention is used to establish a fuzzy rule-based delay mechanism to determine the actual input system action at the current moment. Specifically, in some embodiments, the process of determining the control force at the current moment using the delay mechanism based on the control force after hard constraints is as follows: ; Where, represents the control force output by the intelligent suspension at time t, represents the delay time, and its value is taken as a sampling time, represents at the control force output by the intelligent suspension at the moment; Delay time is determined by the following formula: ; Where, represents the unit quantity determined by the maximum limit control force of the intelligent suspension, represents the unit time delay. In one embodiment, the unit quantity determined by the maximum limit control force takes a value of 1 kN, and the unit time delay takes a value of 10 ms.
[0037] S23: Detect the fatigue degree of the control force at the current moment and adjust the control force at the current moment according to the fatigue degree. In some embodiments, the working intensity of the intelligent suspension can be measured by the magnitude and change frequency of the control force, that is, the working intensity of the intelligent suspension is calculated by the following formula: ; Where, represents the working intensity. The fatigue degree is calculated by the following formula: ; Where, represents the fatigue degree, represents the working time of the intelligent suspension, and represent the fatigue coefficients. The fatigue coefficients and are determined according to the characteristics of the intelligent suspension. Specifically, for an intelligent suspension made of high-strength materials and with a reasonable structural design, its fatigue coefficients and Relatively small, because such a suspension is less likely to cause fatigue under the same working conditions; conversely, for an intelligent suspension with general material properties or certain limitations in structural design, the fatigue coefficient will increase accordingly. Set a fatigue threshold. If the fatigue at the current moment exceeds the fatigue threshold, reduce the control force at the current moment to complete the reduction of the load on the intelligent suspension. The setting of this threshold is determined comprehensively based on various factors such as the design life of the intelligent suspension, safety performance requirements, and actual usage scenarios. For example, in some scenarios with high safety requirements, such as vehicles traveling at high speeds, the fatigue threshold will be set relatively low to ensure that the intelligent suspension can be adjusted in a timely manner when signs of fatigue appear; while in some scenarios with higher comfort requirements, such as smooth driving on urban roads, the fatigue threshold can be appropriately increased. Among them, the process of reducing the control force at the current moment includes: Set the lower limit of the reduction of the control force. In order to ensure that the intelligent suspension can still meet the basic driving performance and safety requirements of the vehicle while reducing the load, the lower limit of the reduction of the control force is set. The lower limit is determined based on factors such as the type of vehicle, driving conditions, and the minimum working ability of the intelligent suspension. For example, for a small car, under normal driving conditions, the lower limit of the reduction of the control force is set to 30% of its rated maximum control force to ensure that when encountering a certain bumpy road surface, the intelligent suspension can still provide sufficient support force and buffering effect to ensure the driving stability and comfort of the vehicle; for a large bus, due to its large weight, the lower limit may be set to 40% of the rated maximum control force.
[0038] The way of reduction each time: Adopt the method of reducing according to a proportion. The specific reduction proportion is determined according to the difference between the current fatigue and the fatigue threshold and the response characteristics of the intelligent suspension. Set a proportionality coefficient , and its value range is . When the amplitude of the fatigue exceeding the threshold is small, such as the difference is within 10% of the threshold, the proportionality coefficient takes a small value, such as 0.05, that is, each time the control force at the current moment is reduced by 5% of the current value; when the amplitude of the fatigue exceeding the threshold is large, such as the difference exceeds 30% of the threshold, the proportionality coefficient takes a large value, such as 0.2, that is, each time the control force at the current moment is reduced by 20% of the current value. At the same time, after reducing the control force each time, the driving state of the vehicle and the working condition of the intelligent suspension will be monitored in real time. If it is found that the driving performance of the vehicle has decreased significantly or the intelligent suspension is working abnormally, the reduction proportion will be appropriately adjusted to achieve the best load reduction effect.
[0039] Through this innovative technology of dynamically adjusting the control force according to fatigue, the load of the intelligent suspension can be effectively reduced, its service life can be extended, and at the same time, the performance and safety of the vehicle under different driving conditions can be ensured. This precise control method can better adapt to the actual working state of the intelligent suspension compared with the traditional fixed control method, improving the reliability and practicality of the intelligent suspension system.
[0040] In one embodiment, a disturbance amount that satisfies a Gaussian random distribution is added to the control force output to the intelligent suspension That is, the control force output by the intelligent suspension at the current moment is: .
[0041] S4: Repeat steps S2 - S3 multiple times, and correspondingly save multiple groups of samples in the experience pool; use the multiple groups of samples to train the first evaluation network, the second evaluation network, and the control strategy network to obtain the first evaluation model, the second evaluation model, and the control strategy model.
[0042] In some embodiments, in each training time step, multiple groups of samples are randomly drawn from the experience pool , and the corresponding target value is calculated by the following formula: ; Where, represents the target value, represents the discount factor, represents the j - th evaluation network, represents the network weight of the j - th evaluation network, represents the state of the vehicle at time t, represents the control force output by the intelligent suspension at time t, represents the control strategy network, represents the network weight of the control strategy network; Combined with the target value, calculate the loss functions for training the two evaluation networks: ; Where, represents the loss function of the j - th evaluation network, N represents the time step; use the loss functions of the two evaluation networks to train the two evaluation networks; Train the control strategy network by the following formula: .
[0043] In one embodiment, the number of training episodes M is set to 2500, the time step N is set to 1000, and in each training time step, 256 groups of samples are randomly drawn from the experience pool , the discount factor takes a value of 0.99. For the discount factor , dynamically adjusted according to the change of the state of the intelligent suspension. When the vehicle is driving on a road with relatively stable road conditions, the discount factor is appropriately increased to make the intelligent suspension pay more attention to long-term rewards; when the road conditions are complex and changeable, the discount factor is appropriately decreased to make the intelligent suspension pay more attention to short-term immediate rewards.
[0044] In some embodiments, during the training of the control policy network and the two evaluation networks: the two evaluation networks are trained using the stochastic gradient descent method, which enables the evaluation networks to continuously adjust their own parameters to more accurately quantify and evaluate the advantages and disadvantages of the control policy. The control policy network is trained using the stochastic gradient ascent method to optimize the performance of the vehicle in terms of comfort, safety, etc. under different driving conditions. Specifically, by calculating the gradient of the objective function (such as maximizing the expected reward, etc.) of the control policy network on the current mini-batch of data with respect to the network parameters, and updating the parameters along the direction of the gradient ascent, the control policy network can continuously improve the control policy it generates to obtain better performance. During the training of the control policy network and / or the two evaluation networks, in order to ensure the stability and efficiency of the training, a mechanism for dynamically adjusting the learning rate is introduced, that is, every time a set number of time steps pass, the current training state is evaluated: if the convergence speed of the network is slow and the decrease in the loss value is small, the learning rate of the training is increased; if the network shows unstable fluctuations, the learning rate is decreased; the learning rate is adjusted within the range of [0.0001, 0.01]. The set number of time steps can be adjusted according to the actual training situation.
[0045] In one embodiment, the current training state is evaluated every 100 time steps. The evaluation indicators mainly include the convergence speed of the network and the change of the loss value. The convergence speed is measured by observing the change amplitude of the network parameters in multiple consecutive time steps. If the parameter change amplitude is small and tends to be stable, the network convergence speed is considered to be fast; conversely, if the parameter change is still large and unstable, the convergence speed is considered to be slow. The decrease amplitude of the loss value is determined by comparing the loss function values of adjacent time steps. If the loss value decreases significantly, it means that the network is learning effectively; if the loss value decreases slightly, it indicates that the network learning effect is not good. When the evaluation finds that the network converges slowly and the loss value decreases slightly, it means that the current learning rate may be too small, resulting in slow update of network parameters and failure to quickly adapt to changes in training data. At this time, increase the learning rate of training (for example, multiply the learning rate by 1.2 times, but ensure that the learning rate is in the range of [0.0001, 0.01]) to speed up the update speed of network parameters and promote faster convergence of the network. When the network is unstable, it is manifested as a large oscillation in the loss function value during training, no longer showing a steady downward trend, or the gradient changes become extremely drastic and irregular. This unstable fluctuation may be due to the learning rate being too large, resulting in too aggressive network parameter updates and the inability to accurately find the optimal solution. In this case, reducing the learning rate (for example, multiplying the learning rate by 0.8, which also needs to be within the range of [0.0001, 0.01]) makes the network parameter updates more stable, thereby stabilizing the network training process and avoiding falling into a local optimal solution or training divergence.
[0046] This innovative technology of dynamically adjusting the learning rate according to the network training status can effectively solve the problems of slow convergence and unstable training caused by fixed learning rate in traditional training methods, improve the training efficiency and performance of the intelligent suspension control network, and enable the intelligent suspension system to learn the optimal control strategy faster and more accurately, thereby better ensuring the comfort and safety of the vehicle in practical applications.
[0047] The learning rate is adjusted in the range of [0.0001, 0.01]. This range setting has been verified by a large number of experiments and practical applications. It can ensure the stability of network training while providing sufficient flexibility to adapt to different training needs.
[0048] In one embodiment, the two evaluation networks are trained using the stochastic gradient descent method, and the network weights of the two evaluation networks are adjusted using a soft update method during the training process. and To update, that is: ; in, and respectively represent the network weights of the two evaluation networks after update, represents the soft update frequency, with a value of 0.001; Use the stochastic gradient ascent method to train the control policy network. During the training process, the network weights of the control policy network are also updated using the soft update method, that is: that is: ; wherein, represents the network weight of the control policy network after update. Here, the soft update frequency still takes the value of 0.001. Set the evaluation time step to 50, that is, after every 50 time steps, evaluate the current training state.
[0049] S5: Run the intelligent suspension and use the first evaluation model, the second evaluation model, and the control policy model obtained in step S3 to control the intelligent suspension.
[0050] Based on the control method considering the safety of the intelligent suspension provided by the present invention, a control system considering the safety of the intelligent suspension is also provided. The system includes an intelligent agent, a vehicle or a test bench equipped with an intelligent suspension and related sensors, a state observer and state estimator, and an endogenous reward function. Among them, the intelligent agent controls the intelligent suspension, and the control method considering the safety of the intelligent suspension provided by the present invention is integrated in the intelligent agent. The related sensors include an acceleration sensor, a displacement sensor, and an inertial measurement unit (IMU). The operation process of the intelligent suspension control system is as follows: The intelligent agent obtains data such as vehicle body acceleration, vehicle body speed, suspension dynamic deflection, and the derivative of suspension dynamic deflection from the vehicle system through sensors, and combines these data into the current state information of the vehicle through a state observer and a state estimator. Then, the intelligent agent decides how much control force the intelligent suspension should output in the current state based on this state information. Under the action of the control force, the vehicle state changes, and the intelligent agent generates a reward value for evaluating its control action according to the new system state and the endogenous reward function. The intelligent agent performs self-iteration and control policy optimization with reference to the reward value according to the control method considering the safety of the intelligent suspension provided by the present invention.
[0051] Figure 6 It is a schematic structural diagram of an electronic device 1 provided in an embodiment of the present invention. Figure 6 It shows a block diagram of an exemplary electronic device 1 suitable for implementing the embodiments of the present invention. Figure 6 The displayed electronic device 1 is only an example and should not bring any limitations to the functions and usage ranges of the embodiments of the present invention.
[0052] Such as Figure 6As shown, the electronic device 1 is presented in the form of a general-purpose computing device. The electronic device 1 is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are only examples and are not intended to limit the implementation of the invention described and / or claimed herein.
[0053] The components of the electronic device 1 may include, but are not limited to: one or more processors or processing units 3, a system memory 8, and a bus 4 that couples different system components (including the system memory 8 and the processing unit 3).
[0054] The bus 4 represents one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any of a variety of bus structures. By way of example, these architectures include, but are not limited to, Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MAC) bus, Enhanced ISA bus, Video Electronics Standards Association (VESA) local bus, and Peripheral Component Interconnect (PCI) bus.
[0055] The electronic device 1 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by the electronic device 1, including volatile and nonvolatile media, removable and non-removable media.
[0056] The system memory 8 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 9 and / or cache memory 10. The electronic device 1 may further include other removable / non-removable, volatile / nonvolatile computer system storage media. By way of example only, a storage system 11 can be used for reading and writing on non-removable, nonvolatile magnetic media ( Figure 6 not shown, typically called a "hard disk drive"). Although Figure 6 not shown in the figure, a disk drive for reading and writing on a removable nonvolatile disk (e.g., a "floppy disk") and an optical disk drive for reading and writing on a removable nonvolatile optical disk (e.g., a CD-ROM, a DVD-ROM, or other optical media) can be provided. In these cases, each drive can be connected to the bus 4 through one or more data media interfaces. The system memory 8 may include at least one program product having a set (e.g., at least one) of program modules that are configured to perform the functions of the embodiments of the present invention.
[0057] A program / utilities 12 having a set (at least one) of program modules 13 can be stored, for example, in the system memory 8. Such program modules 13 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment. The program modules 13 generally execute the functions and / or methods in the embodiments described in the present invention.
[0058] The electronic device 1 can also communicate with one or more external devices 2 (such as a keyboard, a pointing device, a display 6, etc.), can also communicate with one or more devices that enable a user to interact with the electronic device 1, and / or can communicate with any device that enables the electronic device 1 to communicate with one or more other computing devices (such as a network card, a modem, etc.). Such communication can be carried out through the input / output (I / O) interface 7. Moreover, the electronic device 1 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through the network adapter 5. As Figure 6 shown, the network adapter 5 communicates with other modules of the electronic device 1 through the bus 4. It should be understood that, although not shown in the figure, other hardware and / or software modules can be used in combination with the electronic device 1, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0059] The processing unit 3 executes various functional applications and data processing by running the programs stored in the system memory 8, such as implementing the control method considering intelligent suspension safety provided in the embodiments of the present invention.
[0060] The embodiments of the present invention also provide a non-transitory computer-readable storage medium storing computer instructions, on which a computer program is stored. When the program is executed by a processor, it is the control method considering intelligent suspension safety provided in all the embodiments of the present application.
[0061] The computer storage medium of the embodiments of the present invention may adopt any combination of one or more computer-readable media. The computer-readable media may be computer-readable signal media or computer-readable storage media. More specific examples (a non-exhaustive list) of the computer-readable storage media include: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, apparatus, or device.
[0062] The computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries the computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium may also be any computer-readable medium other than the computer-readable storage medium, and the computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device.
[0063] The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to wireless, wire, optical cable, RF, etc., or any suitable combination of the above. The computer program code for performing the operations of the present invention can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages - such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).
[0064] The embodiments of the present invention further provide a computer program product, including a computer program, and the computer program realizes the control method considering intelligent suspension safety according to the above when being executed by a processor.
[0065] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in the disclosure of the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution disclosed in the present invention can be achieved, and no limitation is imposed herein.
[0066] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub - combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A control method considering the safety of intelligent suspension, characterized in that, Including: S1: Initialize the first evaluation network, the second evaluation network, the control strategy network, and the experience pool; S2: The control strategy network determines the current control force output by the intelligent suspension according to the current state of the vehicle; Apply the current control force to the vehicle to obtain the next state of the vehicle; The current state includes the acceleration and speed of the vehicle running at the current moment, the dynamic stroke of the intelligent suspension, and the speed difference between the unsprung mass and the sprung mass of the vehicle; S3: Determine the current comprehensive reward according to the current control force and the current state; Save the current state, the current control force, the current comprehensive reward, and the next state as a set of samples in the experience pool; The current comprehensive reward includes the corresponding current acceleration reward, the current dynamic stroke reward, and the current tire dynamic load reward corresponding to the tire dynamic load of the vehicle; S4: Repeat steps S2 to S3 multiple times, and save multiple sets of samples in the experience pool correspondingly; Use multiple sets of samples to train the first evaluation network, the second evaluation network, and the control strategy network to obtain the first evaluation model, the second evaluation model, and the control strategy model; S5: Run the intelligent suspension, and use the first evaluation model, the second evaluation model, and the control strategy model obtained in step S3 to control the intelligent suspension.
2. The control method considering the safety of the intelligent suspension according to claim 1, wherein, In step S3: The current acceleration reward is: ; Among them, represents the acceleration reward at time t, represents the acceleration of the vehicle at time t, , , and represent the acceleration reward coefficients, , and represent the acceleration thresholds; The current dynamic stroke reward is: ; Among them, represents the dynamic stroke reward at time t, represents the dynamic stroke of the intelligent suspension at time t, , represents the stroke of the vehicle at the current moment, represents the stroke of the unsprung mass of the vehicle at time t; , and represent the dynamic stroke reward coefficient, and represent the dynamic stroke threshold; The current tire dynamic load reward is: ; Among them, represents the tire dynamic load reward at time t, represents the wheel dynamic load of the vehicle at time t, , represents the road surface height excitation information of the vehicle at time t, , and represent the tire dynamic load reward coefficient, , , and represent the tire dynamic load threshold, represents the intermediate variable calculated in the range of the current tire dynamic load reward, which is related to time t and is: ; Among them, represents the tire stiffness of the vehicle at time t, represents the body mass of the vehicle, represents the unsprung mass, and g represents the acceleration due to gravity; The current comprehensive reward is: ; Among them, represents the current comprehensive reward, , and respectively represent the weights of the current acceleration reward , the current dynamic stroke reward and the current tire dynamic load reward .
3. The control method considering the safety of the intelligent suspension according to claim 2, characterized in that, Step S2 further includes: Perform hard constraints on the control force output by the control strategy network: Determine the current control force according to the control force after the hard constraint by using a delay mechanism; Detect the fatigue degree of the current control force, and adjust the current control force according to the fatigue degree.
4. The control method considering the safety of the intelligent suspension according to claim 3, characterized in that, Perform the hard constraint through the following formula: ; Among them, represents the control force after hard constraint on the control force output by the control policy network at time t, represents the control policy network, represents the network weights of the control policy network, represents the truncation function, and respectively represent the minimum and maximum values that limit the control force at the current moment, represents the state of the vehicle at time t.
5. The control method considering the safety of the intelligent suspension according to claim 4, characterized in that The process of determining the current control force according to the control force after the hard constraint is as follows: ; Among them, represents the control force output by the intelligent suspension at time t, represents the delay time, represents at the control force output by the intelligent suspension at the moment; Delay time Determined by the following formula: ; Among them, represents the unit quantity determined by the maximum limit control force of the intelligent suspension, represents the unit time delay.
6. The control method considering the safety of the intelligent suspension according to claim 5, wherein During the process of detecting the fatigue degree of the current control force and adjusting the current control force according to the fatigue degree: Calculate the working intensity of the intelligent suspension through the following formula: ; Among them, represents the working intensity; Calculate the fatigue degree through the following formula: ; Among them, represents the fatigue degree, and represents the fatigue coefficient, represents the working time of the intelligent suspension; Set a fatigue degree threshold. If the fatigue degree at the current moment exceeds the fatigue degree threshold, reduce the current control force.
7. The control method considering the safety of the intelligent suspension according to claim 2, characterized in that, In step S4: In each time step of training, multiple groups of samples are randomly drawn from the experience pool , and the corresponding target value is calculated by the following formula: ; Among them, represents the target value, represents the discount factor, represents the j-th evaluation network, represents the network weight of the j-th evaluation network, represents the state of the vehicle at time t, represents the control force output by the intelligent suspension at time t, represents the control strategy network, represents the network weight of the control strategy network; Combine the target value to calculate the loss functions for training the two evaluation networks; ; Among them, represents the loss function of the j-th evaluation network, and N represents the time step; the two evaluation networks are trained using the loss functions of the two evaluation networks. Train the control strategy network through the following formula: 。 8. The control method for considering the safety of an intelligent suspension according to claim 7, characterized in that, During the training process of the control strategy network and the two evaluation networks: Use the stochastic gradient descent method to train the two evaluation networks; Use the stochastic gradient ascent method to train the control strategy network; During the training process of the control strategy network and / or the two evaluation networks, evaluate the current training state every set number of time steps, and adjust the learning rate of the training according to the evaluation result; The learning rate is adjusted within the range of [0.0001, 0.01].
9. A readable storage medium, characterized in that, A computer program is stored on the readable storage medium, and when the computer program is executed by a processor, the steps of the control method for considering the safety of the intelligent suspension as described in any one of claims 1 to 8 are implemented.
10. An electronic device, characterized in that, It includes: a memory for storing a computer program; a processor for implementing the steps of the control method for considering the safety of the intelligent suspension as described in any one of claims 1 to 8 when executing the computer program.
Citation Information
Patent Citations
Semi-active suspension multi-target control method and device based on vehicle-road cooperation
CN118906726A
Vehicle following behavior decision-making method based on improved DDPG algorithm
CN118928397A
Suspension system and vehicle
CN119636321A
Energy feedback type active suspension control method based on multi-agent reinforcement learning control
CN119821065A
Self-driving vehicle steering and suspension cooperative control method based on reinforcement learning
CN119975527A