Control methods, media, and equipment considering intelligent suspension safety

Through adaptive hyperparameter adjustment and fatigue constraint optimization of intelligent suspension control strategy, the problem of neglecting actuator dynamics and physical constraints is solved, and the safe and reliable control of the intelligent suspension system in complex environments is achieved.

CN120269980BActive Publication Date: 2025-08-19JILIN UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510775059.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-11
Publication Date
2025-08-19
Estimated Expiration
2045-06-11

AI Technical Summary

Technical Problem

The actuator dynamics and time-delay characteristics are ignored in the control of existing intelligent suspension systems, and physical constraints and actual limitations are not fully considered, resulting in unstable control strategies and safety risks.

Method used

By introducing adaptive hyperparameter adjustment strategies and fatigue constraints, combining multi-scene reward and delay mechanisms, the control policy network is optimized to ensure the safety and reliability of the intelligent suspension system in complex environments.

Benefits of technology

It enhances the robustness and adaptability of intelligent suspension control, improves the reliability and performance stability of the system, ensures the safety and feasibility of the actuator, and adapts to the iteration process of different driving scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120269980B_ABST
    Figure CN120269980B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of vehicle dynamics control technology, and more particularly to a control method, medium, and device that considers the safety of an intelligent suspension. The method includes determining, by a control strategy network, a current control force output by the intelligent suspension based on the current state of the vehicle; applying the current control force to the vehicle to obtain the vehicle's next state; determining a comprehensive reward based on the current control force and the current state, taking into account different situations; training a first evaluation network, a second evaluation network, and a control strategy network using multiple sets of current states, current control forces, current comprehensive rewards, and next state to obtain a first evaluation model, a second evaluation model, and a control strategy model; and operating the intelligent suspension and controlling it using the first evaluation model, the second evaluation model, and the control strategy model. The present invention uses comprehensive rewards to strictly constrain and punish various unsafe situations, thereby effectively managing safety risks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of automobile dynamics control, and in particular relates to a control method, medium and device taking into account the safety of intelligent suspension. Background Art

[0002] With the rapid development of artificial intelligence (AI), autonomous driving systems and related fields are experiencing unprecedented opportunities for technological innovation. Compared to traditional approaches, AI offers a fresh perspective and solution for addressing the control issues of key components in autonomous vehicles, such as the suspension system. As a core component for enhancing ride comfort and driving performance, optimizing the performance of the suspension system is crucial to the overall driving experience and functional safety of autonomous vehicles.

[0003] The shortcomings of the existing technology are mainly reflected in the following aspects:

[0004] (1) Ignoring actuator dynamics and time-delay characteristics: In the control of intelligent suspension systems, the dynamic characteristics and time-delay effects of actuators have a crucial impact on the control effect. However, many scholars often ignore these key factors in the application of deep reinforcement learning, resulting in the control strategy being difficult to achieve the expected effect in practical applications, and may even cause system instability;

[0005] (2) Insufficient consideration of physical constraints and practical limitations: The deployment of deep reinforcement learning in intelligent suspension control faces many practical limitations and physical constraints. If these constraints are ignored, the intelligent agent may behave inconsistently with the laws of physics, which will not only fail to effectively control the system but also pose a serious threat to vehicle safety. Summary of the Invention

[0006] In view of this, the present invention aims to provide a control method, medium and equipment that take the safety of intelligent suspension into consideration, and use comprehensive rewards to strictly constrain and punish various unsafe situations, so as to effectively manage safety risks; by combining the reward mechanism for different situations, not only the robustness of the intelligent suspension control strategy is enhanced, but also its reliability and adaptability in complex and changing environments are significantly improved.

[0007] To achieve the above object, the technical solution created by the present invention is implemented as follows:

[0008] A control method considering the safety of an intelligent suspension includes:

[0009] S1: Initialize the first evaluation network, the second evaluation network, the control strategy network and the experience pool;

[0010] S2: The control strategy network determines the current control force output by the intelligent suspension based on the current state of the vehicle. The current control force is applied to the vehicle to determine the vehicle's next state. The current state includes the vehicle's acceleration and speed at the current moment, the dynamic travel of the intelligent suspension, and the velocity difference between the vehicle's sprung mass and unsprung mass.

[0011] S3: Determine the current comprehensive reward based on the current control force and current state; save the current state, current control force, current comprehensive reward, and next state as a set of samples in the experience pool; the current comprehensive reward includes the corresponding current acceleration reward, current dynamic range reward, and the current tire dynamic load reward corresponding to the vehicle's tire dynamic load;

[0012] S4: Repeat steps S2 to S3 multiple times, and save multiple groups of samples in the experience pool; use the multiple groups of samples to train the first evaluation network, the second evaluation network, and the control strategy network to obtain the first evaluation model, the second evaluation model, and the control strategy model;

[0013] S5: Run the intelligent suspension and use the first evaluation model, the second evaluation model, and the control strategy model obtained in step S3 to control the intelligent suspension.

[0014] Furthermore, in step S3:

[0015] The current acceleration rewards are:

[0016] ;

[0017] in, represents the acceleration reward at time t, represents the acceleration of the vehicle at time t, 、 、 and represents the acceleration reward coefficient, 、 and Indicates the acceleration threshold;

[0018] The current dynamic itinerary rewards are:

[0019] ;

[0020] in, represents the dynamic travel reward at time t, represents the dynamic travel of the smart suspension at time t, , Indicates the current travel distance of the vehicle. represents the travel of the unsprung mass of the vehicle at time t; 、 and Indicates the dynamic travel reward coefficient, and Indicates the dynamic travel threshold;

[0021] The current tire dynamic load bonus is:

[0022] ;

[0023] in, represents the tire dynamic load reward at time t, represents the dynamic wheel load of the vehicle at time t, , represents the road height incentive information of the vehicle at time t, 、 and represents the tire dynamic load bonus coefficient, 、 、 and Indicates the tire dynamic load threshold, The intermediate variable representing the range calculation in the current tire dynamic load reward, related to time t, is:

[0024] ;

[0025] in, represents the tire stiffness of the vehicle at time t, Indicates the vehicle's body mass, represents the unsprung mass, g represents the acceleration due to gravity;

[0026] The current comprehensive rewards are:

[0027] ;

[0028] in, Indicates the current comprehensive reward, 、 and Represents the current acceleration reward , Current dynamic itinerary rewards and the current tire dynamic load bonus The weight of .

[0029] Furthermore, step S2 further includes:

[0030] Hard constraints are imposed on the control force output by the control strategy network:

[0031] According to the control force after hard constraint, the current control force is determined by using the delay mechanism;

[0032] Detect the fatigue level of the control force at the current moment and adjust the control force at the current moment according to the fatigue level.

[0033] Furthermore, hard constraints are performed using the following formula:

[0034] ;

[0035] in, It represents the control force after hard constraint on the control force output by the control strategy network at time t, represents the control strategy network, represents the network weight of the control strategy network, represents the interception function, and Respectively represent the minimum and maximum values of the control force at the current moment, Represents the state of the vehicle at time t.

[0036] Furthermore, based on the hard-constrained control force, the process of determining the current control force is as follows:

[0037] ;

[0038] in, represents the control force output by the intelligent suspension at time t, Indicates the delay time, Indicates The control force output by the intelligent suspension at all times;

[0039] Delay time Determined by the following formula:

[0040] ;

[0041] in, It represents the unit quantity determined by the maximum limit control force of the intelligent suspension. Represents unit time lag.

[0042] Furthermore, in the process of detecting the fatigue level of the control force at the current moment and adjusting the control force at the current moment according to the fatigue level:

[0043] The working strength of the smart suspension is calculated using the following formula:

[0044] ;

[0045] in, Indicates work intensity;

[0046] Fatigue is calculated using the following formula:

[0047] ;

[0048] in, Indicates fatigue. and represents the fatigue coefficient, Indicates the working time of the smart suspension;

[0049] Set a fatigue threshold. If the current fatigue exceeds the fatigue threshold, reduce the current control force.

[0050] Furthermore, in step S4:

[0051] In each training time step, multiple groups of samples are randomly drawn from the experience pool , calculate the corresponding target value by the following formula:

[0052] ;

[0053] in, represents the target value, represents the discount factor, represents the j-th evaluation network, represents the network weight of the j-th evaluation network, represents the state of the vehicle at time t, represents the control force output by the intelligent suspension at time t, represents the control strategy network, represents the network weight of the control strategy network;

[0054] Combine the target values to calculate the loss function for training the two evaluation networks:

[0055] ;

[0056] in, Denotes the loss function of the j-th evaluation network, N denotes the time step; the two evaluation networks are trained using their loss functions;

[0057] The control strategy network is trained as follows:

[0058] .

[0059] Furthermore, in the process of training the control policy network and the two evaluation networks: the two evaluation networks are trained using the stochastic gradient descent method; the control policy network is trained using the stochastic gradient ascent method; in the process of training the control policy network and / or the two evaluation networks, the current training state is evaluated after each set number of time steps, and the learning rate of the training is adjusted according to the evaluation results; the learning rate is adjusted within the range of [0.0001, 0.01].

[0060] A readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the control method considering the safety of an intelligent suspension provided by the present invention.

[0061] An electronic device, comprising:

[0062] memory for storing computer programs;

[0063] The processor is configured to implement the steps of the control method considering the safety of the intelligent suspension provided by the present invention when executing a computer program.

[0064] Compared with the prior art, the present invention can achieve the following beneficial effects:

[0065] (1) The control method for intelligent suspension safety created by the present invention enhances safety control and practical application feasibility: by introducing adaptive hyperparameter adjustment strategy and fatigue constraint, the present invention more accurately integrates actuator dynamics constraints in the deep reinforcement learning algorithm. Adaptive hyperparameter adjustment enables the model to quickly adapt to different working conditions during training and optimize the control strategy; actuator fatigue constraint effectively avoids actuator failure due to overuse, ensures the safety and feasibility of the decision-making process, greatly enhances the reliability of the control strategy in actual engineering applications, deeply verifies its control performance, and provides solid guarantees for large-scale actual deployment;

[0066] (2) The control method for intelligent suspension safety created by the present invention improves system efficiency and iteration speed: the dynamic weight allocation of multi-scenario rewards and the adaptive hyperparameter adjustment strategy work together to significantly accelerate the iteration process of intelligent suspension in different driving scenarios. The multi-scenario dynamic weight allocation makes the reward function more in line with actual needs and finds the optimal strategy more quickly; the adaptive hyperparameter adjustment optimizes the model training process, effectively shortens the learning cycle, improves the overall efficiency of the system, and provides strong support for the rapid response and precise control of the intelligent suspension system;

[0067] (3) The control method considering the safety of the intelligent suspension created by the present invention improves the reliability and performance stability of the system: the combination of control force constraints and dynamic weight distribution of multi-scenario rewards fully considers the characteristics of the intelligent suspension and the driving state of the vehicle; fatigue constraints and optimized delay mechanisms ensure the stable operation of the actuator, and multi-dimensional state perception provides accurate information for the control strategy, so that the system can maintain excellent performance and high safety in various complex road conditions and driving scenarios, significantly improving the reliability and performance stability of the system, and laying a solid foundation for the long-term stable operation of the intelligent suspension system. BRIEF DESCRIPTION OF THE DRAWINGS

[0068] The accompanying drawings, which constitute part of the present invention, are intended to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are intended to explain the present invention and do not constitute an undue limitation of the present invention. In the accompanying drawings:

[0069] Figure 1 A flow chart of a control method considering the safety of an intelligent suspension according to an embodiment of the present invention;

[0070] Figure 2 A schematic diagram of a framework of a control method considering the safety of an intelligent suspension according to an embodiment of the present invention;

[0071] Figure 3 A schematic diagram of a control strategy network according to an embodiment of the present invention;

[0072] Figure 4 Schematic diagram of two evaluation networks according to an embodiment of the present invention;

[0073] Figure 5 A schematic diagram of a dynamic model of an intelligent suspension system according to an embodiment of the present invention;

[0074] Figure 6 A schematic structural diagram of an electronic device according to an embodiment of the present invention.

[0075] Description of reference numerals:

[0076] 1. Electronic device; 2. External device; 3. Processing unit; 4. Bus; 5. Network adapter; 6. Display; 7. (I / O) interface; 8. System memory; 9. Random access memory; 10. Cache memory; 11. Storage system; 12. Utility; 13. Program module. DETAILED DESCRIPTION

[0077] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and do not constitute a limitation of the present invention.

[0078] It should be noted that, in the absence of conflict, the embodiments of the present invention and the features in the embodiments can be combined with each other.

[0079] In the description of the present invention, it should be understood that the terms "center", "longitudinal", "lateral", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside" and the like indicate orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present invention. In addition, the terms "first", "second" and the like are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, features defined as "first", "second" and the like may explicitly or implicitly include one or more of the features. In the description of the present invention, unless otherwise specified, "multiple" means two or more.

[0080] In the description of the present invention, it should be noted that, unless otherwise expressly specified or limited, the terms "installed," "connected," and "connected" should be understood in a broad sense. For example, they can refer to fixed connections, detachable connections, or integral connections; they can refer to mechanical connections or electrical connections; they can refer to direct connections or indirect connections through an intermediate medium; and they can refer to internal connections between two components. Those skilled in the art can understand the specific meanings of the above terms in the present invention based on specific circumstances.

[0081] The present invention will be described in detail below with reference to the accompanying drawings and in combination with embodiments.

[0082] like Figures 1 to 2 As shown, the control method considering the safety of the intelligent suspension according to the embodiment of the present invention includes:

[0083] S1: Initialize the first evaluation network, the second evaluation network, the control strategy network and the experience pool.

[0084] In one embodiment, the control strategy network is as follows: Figure 3 As shown in the figure, the vehicle state is input into a fully connected layer consisting of 400 neurons for processing. The processed data is processed by the ReLU activation function and then input into a fully connected layer consisting of 300 neurons. The output data is processed by the ReLU activation function and then input into a fully connected layer consisting of 1 neuron. Finally, the output data is activated by the Tanh activation layer and scaled by the scaling layer to obtain the control force output by the intelligent suspension. The structures of the first evaluation network and the second evaluation network are the same, as shown in the figure. Figure 4As shown in the figure, the vehicle state is input into a fully connected layer composed of 128 neurons for processing. The processed data is processed by the ReLU activation function and then input into a fully connected layer composed of 200 neurons to obtain the first feature. At the same time, the control force output by the intelligent suspension is input into a fully connected layer composed of 200 neurons to obtain the second feature. After the corresponding elements of the first and second features are added, they are processed by the ReLU activation function and then input into a fully connected layer composed of 1 neuron to obtain the evaluation result. The initialization operation includes: The network weights in , First Rating Network The network weights in , and control strategy network The network weight Perform random initialization, s represents the state of the vehicle, and a represents the control force output by the intelligent suspension.

[0085] S2: The control strategy network determines the current control force output by the intelligent suspension based on the vehicle's current state. This control force is then applied to the vehicle to determine the vehicle's next state. The current state includes the vehicle's acceleration and velocity at the current moment, the dynamic travel of the intelligent suspension, and the velocity difference between the vehicle's sprung and unsprung masses. The acceleration used in this disclosure is based on the vertical acceleration of the vehicle body.

[0086] In some embodiments, the system dynamics model of the intelligent suspension is as follows: Figure 5 As shown, Figure 5 middle, Indicates the vehicle's body mass, represents the unsprung mass, represents the vertical displacement of the vehicle body at time t, represents the travel of the vehicle's unsprung mass at time t, represents the road height incentive information of the vehicle at time t, Indicates the vehicle's tire stiffness, represents the suspension spring stiffness of the smart suspension, Indicates the damping coefficient of the smart suspension Represents the damping coefficient of the vehicle tire. Based on the system dynamics model of the intelligent suspension, the current state of the vehicle can be obtained. ,in, and They represent the acceleration and velocity of the sprung mass of the vehicle at time t, and are important parameters for the smoothness performance of the vehicle suspension. represents the dynamic travel of the intelligent suspension at time t, which is an important physical constraint in suspension control. It represents the velocity difference between the sprung mass and the unsprung mass at time t, because the power of the smart suspension is P=F×v, F represents the output force of the smart suspension, and v is the velocity difference , the speed difference v characterizes the physical constraints on the performance of the smart suspension to a certain extent. , Indicates the unsprung speed of the vehicle at time t. Current state The parameters in the above formula guarantee the feasibility of the proposed model to be deployed in intelligent suspensions under real-world constraints. The present invention uses a state feedback law. Therefore, its performance depends on the accuracy and appropriate selection of the measured or estimated state. In one embodiment, in the control implementation of the intelligent suspension, the vertical acceleration is directly affected by high-pass filtering and the Integrate and estimate the velocity , dynamic itinerary The speed difference is obtained by directly measuring the actuator displacement of the smart suspension. The dynamic stroke is calculated by using the mixed smooth derivative method In actual application and deployment, it is only necessary to obtain the values of the above variables.

[0087] S3: Determine the current comprehensive reward based on the current control power and the current state; save the current state, current control power, current comprehensive reward and next state as a set of samples in the experience pool.

[0088] In the application of deep reinforcement learning to the control of intelligent suspension systems, it is crucial to ensure the physical safety of the operation of intelligent suspension systems. Specifically, the present invention imposes physical safety constraints on three key state variables: vehicle acceleration, dynamic travel, and tire dynamic load. Constraints on vehicle body acceleration and dynamic travel ensure that the vehicle maintains passenger comfort and prevents safety accidents caused by excessive vibration or tilt when the road changes suddenly; constraints on tire dynamic loads can prevent tire damage or blowouts due to overload, ensuring stability and safety during driving. In response to these considerations, the present invention refines the reward function through the guidance of physical safety constraints to minimize the impact of training termination caused by any state variable exceeding its safety limit. That is, the present invention proposes rewards for acceleration, dynamic travel, and tire dynamic load of vehicles, and combines the three rewards to achieve multi-dimensional state perception. Specifically, the current comprehensive reward includes the corresponding current acceleration reward, current dynamic travel reward, and current tire dynamic load reward corresponding to the vehicle's tire dynamic load.

[0089] In some embodiments, acceleration is subject to hierarchical reward constraints, that is, its reward is divided into four levels based on the comfort threshold, tolerance threshold, and safety threshold. When the safety threshold is exceeded, the training episode is terminated and a high penalty is imposed to guide the agent away from behaviors that pose safety risks. Specifically, the current acceleration reward is:

[0090] ;

[0091] in, represents the acceleration reward at time t, 、 、 and Indicates the acceleration reward coefficient, which is selected based on the actual situation and human experience. 、 and Indicates the acceleration threshold. It should be noted that the acceleration reward coefficient Much smaller than the acceleration bonus coefficient 、 and , in order to ensure that when the acceleration exceeds the maximum acceleration threshold, there is a small acceleration reward, that is, a large penalty is imposed. If the acceleration reward at time t is the acceleration reward coefficient When the acceleration threshold is reached, the training is terminated. 、 and The determination process includes:

[0092] Control a vehicle equipped with an intelligent suspension to drive at different speeds, different road conditions, and different vehicle load conditions. Collect acceleration data during vehicle driving, classify and count the acceleration data according to the value, and draw a probability distribution histogram of the acceleration value. Based on the distribution of the probability distribution histogram, determine the acceleration ranges with 50%, 80%, and 100% data coverage, and obtain the corresponding acceleration thresholds. 、 and In one embodiment, the driving speed is set to a range of 10km / h to 120km / h, the road conditions include flat roads, rough roads, and speed bump roads, and the vehicle load conditions include no-load (the vehicle only contains the driver and necessary equipment), half-load (50% of the vehicle's rated load), and full-load (the vehicle reaches the rated load limit). Acceleration data is collected during vehicle driving, and the acceleration data is classified and counted according to the numerical value, and a probability distribution histogram of the acceleration value is drawn. Determine the acceleration ranges covered by 50%, 80%, and 100% of the data, and obtain the corresponding acceleration thresholds respectively. 、 and The specific process is as follows: perform integral calculation on the probability distribution histogram, find the acceleration interval containing 50%, 80%, and 100% of the data, and obtain the corresponding acceleration threshold 、 and In addition, the acceleration bonus coefficient 、 、 and They are -10, -20, -30 and -2000 respectively.

[0093] Taking into account vehicle design and safety margins, the safe range of the smart suspension's dynamic travel should be set to no more than 80% of the vehicle's maximum design suspension travel to prevent damage to the suspension system under extreme conditions. The corresponding current dynamic travel bonus is:

[0094] ;

[0095] in, represents the dynamic travel reward at time t, 、 and Indicates the dynamic travel reward coefficient, which is selected based on the actual situation and human experience. and Indicates the dynamic travel threshold. It should be noted that the dynamic travel reward coefficient Much smaller than the dynamic travel reward coefficient and , in order to ensure that when the dynamic travel exceeds the maximum dynamic travel threshold, there is a small dynamic travel reward, that is, a large penalty is imposed, that is, if the dynamic travel reward at time t is the dynamic travel reward coefficient When the training is terminated, the dynamic stroke threshold and Acquisition method and acceleration threshold 、 and The acquisition method is consistent, that is, controlling the vehicle equipped with the intelligent suspension to drive in different environments (different driving speeds, different road conditions, and different vehicle load conditions), collecting dynamic travel data during vehicle driving, classifying and counting the dynamic travel data according to the value size, and drawing a probability distribution histogram of the dynamic travel value. According to the distribution of the probability distribution histogram, the acceleration range covered by the data of 100%+20% (that is, adding a probability distribution histogram of 20% to the total probability distribution histogram) and 100%+50% (that is, adding a probability distribution histogram of 50% to the total probability distribution histogram) is determined, and the corresponding dynamic travel threshold values are obtained respectively. and In one embodiment, when obtaining the dynamic travel threshold, the driving environment of the vehicle is consistent with the acceleration threshold, and the corresponding acceleration threshold is obtained. and 0.12 and 0.15 respectively. In addition, set the dynamic travel reward coefficient 、 and They are -10, -20 and -100 respectively.

[0096] The load of a vehicle is the maximum weight that the vehicle can safely support, as specified by the manufacturer, including the weight of the vehicle itself, passengers and cargo. When designing the load-bearing capacity of wheels and tires, both static and dynamic loads are taken into account. Dynamic loads are the additional forces exerted on the tires when the vehicle is in motion, which are caused by conditions such as acceleration, emergency braking or driving on uneven surfaces. To prevent wheel slippage or tire damage under these conditions, the total load-bearing capacity of the wheels should neither be less than 60% nor exceed 150% of the vehicle's design load under normal driving conditions. The corresponding current tire dynamic load bonuses are:

[0097] ;

[0098] in, represents the tire dynamic load reward at time t, represents the dynamic wheel load of the vehicle at time t, , 、 and Indicates the tire dynamic load bonus coefficient, which is selected based on actual conditions and human experience. 、 、 and Indicates the tire dynamic load threshold, The intermediate variable representing the range calculation in the current tire dynamic load reward, related to time t, is:

[0099] ;

[0100] in, represents the tire stiffness of the vehicle at time t, Indicates the vehicle's body mass, represents the unsprung mass and g represents the acceleration due to gravity.

[0101] It should be noted that the tire dynamic load bonus factor Much smaller than the dynamic travel reward coefficient and , in order to ensure that when the tire dynamic load exceeds the tire dynamic load threshold, there is a small tire dynamic load reward, that is, a large penalty is imposed, that is, if the tire dynamic load reward at time t is the tire dynamic load reward coefficient When the training is stopped.

[0102] Tire dynamic load threshold 、 、 and Acquisition method and acceleration threshold 、 and The acquisition method is consistent with that of the previous one, that is, controlling a vehicle equipped with an intelligent suspension to drive in different environments (different driving speeds, different road conditions, and different vehicle load conditions), collecting dynamic travel data during vehicle driving, classifying and counting the dynamic travel data according to the value size, and drawing a probability distribution histogram of the dynamic travel value. According to the distribution of the probability distribution histogram, the acceleration ranges covered by the data at 80%, 100%+20%, 30%, and 100%+50% are determined, and the corresponding tire dynamic load thresholds are obtained respectively. 、 、 and In one embodiment, when obtaining the tire dynamic load threshold, the vehicle's driving environment is consistent with the acceleration threshold, and the corresponding tire dynamic load threshold is obtained. 、 、 and They are 0.8, 1.2, 0.3 and 1.5 respectively. In addition, set the tire dynamic load bonus coefficient 、 and They are -10, -20 and -2000 respectively.

[0103] Based on current acceleration reward , Current dynamic itinerary rewards and the current tire dynamic load bonus , the current comprehensive reward is:

[0104] ;

[0105] in, Indicates the current comprehensive reward, 、 and Represents the current acceleration reward , Current dynamic itinerary rewards and the current tire dynamic load bonus In actual application, the weight is adjusted dynamically according to different driving scenarios and vehicle driving status. When the vehicle starts and stops frequently on urban roads, Adjusted to 0.6, Adjust to 0.15, Adjusted to 0.25, focusing on the impact of vehicle acceleration on ride comfort; when the vehicle is driving on an off-road road, Adjust to 0.3, Adjusted to 0.35, Adjusted to 0.35, it strengthens the tire dynamic load and suspension deflection constraints to ensure vehicle passability and safety.

[0106] In some embodiments, the design of the control force output by the intelligent suspension needs to consider the safety requirements of autonomous vehicles and the physical limitations of actuator power to reduce related risks. The present invention combines a hard constraint mechanism with control force constraints to achieve higher penalties for dangerous states and behaviors that may endanger system safety. Specifically, step S2 also includes:

[0107] S21: Hard constraint is applied to the control force output by the control strategy network. In some embodiments, the hard constraint is applied by the following formula:

[0108] ;

[0109] in, It represents the control force after hard constraint on the control force output by the control strategy network at time t, represents the control strategy network, represents the network weight of the control strategy network, represents the interception function, and Respectively represent the minimum and maximum values of the control force at the current moment. The minimum value of the control force at the current moment and maximum value The value of is obtained according to the maximum force limit of the actuator. In one embodiment, the minimum and maximum value The values are -3kN and 3kN respectively.

[0110] S22: Based on the hard-constrained control force, a delay mechanism is used to determine the current control force. Since the time lag of an intelligent suspension in a control system is often closely related to its actual actuation capability, the present invention establishes a fuzzy rule-based delay mechanism to determine the actual input system action at the current moment. Specifically, in some embodiments, the process of determining the current control force based on the hard-constrained control force using the delay mechanism is as follows:

[0111] ;

[0112] in, represents the control force output by the intelligent suspension at time t, Indicates the delay time, its value is used as a sampling time, Indicates The control force output by the intelligent suspension at all times;

[0113] Delay time Determined by the following formula:

[0114] ;

[0115] in, It represents the unit quantity determined by the maximum limit control force of the intelligent suspension. Indicates unit time delay, a unit quantity determined by the maximum limit control force in a certain embodiment The value is 1kN, unit time lag The value is 10ms.

[0116] S23: Detecting the fatigue of the control force at the current moment and adjusting the control force at the current moment according to the fatigue. In some embodiments, the working intensity of the smart suspension can be measured by the magnitude of the control force and the frequency of change, that is, the working intensity of the smart suspension can be calculated by the following formula:

[0117] ;

[0118] in, Indicates work intensity. Fatigue is calculated using the following formula:

[0119] ;

[0120] in, Indicates fatigue. Indicates the working time of the smart suspension. and Indicates fatigue coefficient, fatigue coefficient and Determined according to the characteristics of the intelligent suspension. Specifically, for an intelligent suspension that uses high-strength materials and has a reasonable structural design, its fatigue coefficient and It is relatively small because this type of suspension is less likely to fatigue under the same working conditions; on the contrary, for smart suspensions with general material properties or certain limitations in structural design, the fatigue coefficient will increase accordingly. Set a fatigue threshold. If the fatigue at the current moment exceeds the fatigue threshold, reduce the control force at the current moment to complete the reduction of the load on the smart suspension. The setting of this threshold is based on a comprehensive determination of multiple factors such as the design life of the smart suspension, safety performance requirements, and actual usage scenarios. For example, in some scenarios with higher safety requirements, such as high-speed vehicles, the fatigue threshold will be set relatively low to ensure that the smart suspension can be adjusted in time when signs of fatigue appear; and in some scenarios with higher comfort requirements, such as smooth driving on urban roads, the fatigue threshold can be appropriately increased. Among them, the process of reducing the control force at the current moment includes:

[0121] Set a lower limit for reducing control force. To ensure that the intelligent suspension can still meet the basic driving performance and safety requirements of the vehicle while reducing the load, a lower limit for reducing control force is set. The lower limit is determined based on factors such as the type of vehicle, driving conditions, and the minimum working capacity of the intelligent suspension. For example, for small cars, under normal driving conditions, the lower limit of control force reduction is set to 30% of its rated maximum control force to ensure that when encountering certain bumpy roads, the intelligent suspension can still provide sufficient support and cushioning effects to ensure the vehicle's driving stability and comfort; for large buses, due to their larger weight, the lower limit may be set to 40% of the rated maximum control force.

[0122] The method of each reduction is to reduce the fatigue level in proportion. The specific reduction ratio is determined by the difference between the current fatigue level and the fatigue threshold and the response characteristics of the intelligent suspension. Set a proportional coefficient. , its value range is When the fatigue exceeds the threshold by a small amount, such as the difference is within 10% of the threshold, the proportional coefficient The value is small, for example, 0.05, which means that the control force at the current moment is reduced by 5% of the current value each time; when the fatigue exceeds the threshold by a large margin, such as the difference exceeds 30% of the threshold, the proportional coefficient A larger value, such as 0.2, means that the control force is reduced by 20% of the current value each time. At the same time, after each control force reduction, the vehicle's driving status and the working condition of the smart suspension are monitored in real time. If the vehicle's driving performance deteriorates significantly or the smart suspension is malfunctioning, the reduction ratio will be adjusted appropriately to achieve the best load reduction effect.

[0123] This innovative technology, which dynamically adjusts control force based on fatigue, effectively reduces the load on the smart suspension, extending its service life while ensuring vehicle performance and safety under various driving conditions. Compared to traditional fixed control methods, this precise control approach better adapts to the actual operating conditions of the smart suspension, improving the reliability and practicality of the smart suspension system.

[0124] In one embodiment, the control force output by the intelligent suspension is added with a Gaussian random distribution. The disturbance amount, that is, the current control force output by the intelligent suspension is:

[0125] .

[0126] S4: Repeat steps S2 to S3 multiple times, and save multiple groups of samples in the experience pool accordingly; use the multiple groups of samples to train the first evaluation network, the second evaluation network, and the control strategy network to obtain the first evaluation model, the second evaluation model, and the control strategy model.

[0127] In some embodiments, at each training time step, multiple groups of samples are randomly drawn from the experience pool. , calculate the corresponding target value by the following formula:

[0128] ;

[0129] in, represents the target value, represents the discount factor, represents the j-th evaluation network, represents the network weight of the j-th evaluation network, represents the state of the vehicle at time t, represents the control force output by the intelligent suspension at time t, represents the control strategy network, represents the network weight of the control strategy network;

[0130] Combine the target values to calculate the loss function for training the two evaluation networks:

[0131] ;

[0132] in, Denotes the loss function of the j-th evaluation network, N denotes the time step; the two evaluation networks are trained using their loss functions;

[0133] The control strategy network is trained as follows:

[0134] .

[0135] In one embodiment, the number of training episodes M is set to 2500, the time step N is set to 1000, and in each training time step, 256 groups of samples are randomly selected from the experience pool. , discount factor The value is 0.99. , dynamically adjusted based on changes in the smart suspension's state. When the vehicle is on a relatively stable road, the discount factor is appropriately increased, allowing the smart suspension to prioritize long-term rewards. When the road conditions are complex and changeable, the discount factor is appropriately reduced, allowing the smart suspension to prioritize short-term, immediate rewards.

[0136] In some embodiments, during the training of the control policy network and the two evaluation networks, the two evaluation networks are trained using stochastic gradient descent, enabling the evaluation networks to continuously adjust their parameters to more accurately quantify the performance of the control policy. The control policy network is trained using stochastic gradient ascent to optimize various aspects of vehicle performance, such as comfort and safety, under different driving conditions. Specifically, the gradient of the control policy network's objective function (e.g., maximizing expected reward) on the current mini-batch of data with respect to network parameters is calculated, and the parameters are updated along the gradient ascent. This allows the control policy network to continuously improve its generated control policy to achieve better performance. During the training of the control policy network and / or the two evaluation networks, a dynamic learning rate adjustment mechanism is introduced to ensure training stability and efficiency. Specifically, the current training status is evaluated after a set number of time steps. If the network converges slowly and the loss decreases only slightly, the learning rate is increased; if the network exhibits unstable fluctuations, the learning rate is decreased. The learning rate is adjusted within the range [0.0001, 0.01]. The set number of time steps can be adjusted based on the actual training situation.

[0137] In one embodiment, the current training status is evaluated every 100 time steps. Evaluation metrics primarily include the network's convergence speed and loss change. Convergence speed is measured by observing the magnitude of network parameter changes over multiple consecutive time steps. If the parameter changes are small and stable, the network is considered to be converging quickly; conversely, if the parameter changes remain large and unstable, the convergence is considered slow. The magnitude of the loss decrease is determined by comparing the loss function values of adjacent time steps. A significant decrease in the loss indicates effective network learning; a small decrease in the loss indicates poor learning. If the evaluation indicates slow network convergence and a small decrease in the loss, the current learning rate may be too low, resulting in slow network parameter updates and an inability to quickly adapt to changes in the training data. In this case, increasing the learning rate (for example, multiplying the learning rate by 1.2, but ensuring that the learning rate is within the range [0.0001, 0.01]) can accelerate network parameter updates and promote faster network convergence. When the network experiences unstable fluctuations, this is manifested by large fluctuations in the loss function during training, instead of a steady downward trend, or by unusually drastic and erratic gradient changes. This unstable fluctuation may be caused by an excessively high learning rate, which results in overly aggressive network parameter updates and an inability to accurately find the optimal solution. In this case, reducing the learning rate (for example, multiplying it by 0.8, still within the range [0.0001, 0.01]) can smoother network parameter updates and stabilize the network training process, avoiding local optima or training divergence.

[0138] This innovative technology, which dynamically adjusts the learning rate based on the network training status, can effectively address the problems of slow convergence and unstable training caused by a fixed learning rate in traditional training methods. It improves the training efficiency and performance of the intelligent suspension control network, enabling the intelligent suspension system to learn the optimal control strategy more quickly and accurately, thereby better ensuring vehicle comfort and safety in practical applications.

[0139] The learning rate is adjusted in the range of [0.0001, 0.01]. This range has been verified through a large number of experiments and practical applications. It can ensure the stability of network training while providing sufficient flexibility to adapt to different training needs.

[0140] In one embodiment, the two evaluation networks are trained using the stochastic gradient descent method, and the network weights of the two evaluation networks are adjusted using the soft update method during the training process. and To update, that is:

[0141] ;

[0142] in, and Represent the updated network weights of the two evaluation networks, Indicates the soft update frequency, the value is 0.001;

[0143] The control strategy network is trained using the stochastic gradient ascent method, and the network weights of the control strategy network are also adjusted using the soft update method during the training process. To update, that is:

[0144] ;

[0145] in, Indicates the network weight after the control strategy network is updated. The soft update frequency here is The value is still 0.001. Set the evaluation time step to 50, that is, the current training state is evaluated every 50 time steps.

[0146] S5: Run the intelligent suspension and use the first evaluation model, the second evaluation model, and the control strategy model obtained in step S3 to control the intelligent suspension.

[0147] Based on the control method for intelligent suspension safety, the present invention also provides a control system for intelligent suspension safety. The system includes an intelligent agent, a car or test bench equipped with an intelligent suspension and related sensors, a state observer and state estimator, and an intrinsic reward function. The intelligent agent controls the intelligent suspension, and the intelligent agent integrates the control method for intelligent suspension safety provided by the present invention. The related sensors include an acceleration sensor and a displacement sensor inertial measurement unit (IMU). The operation process of the intelligent suspension control system is as follows:

[0148] The intelligent agent obtains data such as vehicle body acceleration, vehicle body speed, suspension dynamic deflection and derivative of suspension dynamic deflection from the entire vehicle system through sensors, and combines this data into the current state information of the vehicle through a state observer and a state estimator. The intelligent agent then decides how much control force the intelligent suspension should output in the current state based on this state information. Under the action of the control force, the state of the entire vehicle changes. The intelligent agent generates a reward value to evaluate its control action based on the new system state and the intrinsic reward function. The intelligent agent performs self-iteration and control strategy optimization with reference to the reward value in accordance with the control method considering the safety of the intelligent suspension provided by the present invention.

[0149] Figure 6 Schematic diagram of the structure of an electronic device 1 provided in an embodiment of the present invention. Figure 6 A block diagram of an exemplary electronic device 1 suitable for implementing embodiments of the present invention is shown. Figure 6The electronic device 1 shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present invention.

[0150] like Figure 6 As shown, electronic device 1 is represented in the form of a general-purpose computing device. Electronic device 1 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. Electronic device may also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0151] The components of the electronic device 1 may include, but are not limited to: one or more processors or processing units 3, a system memory 8, and a bus 4 connecting different system components (including the system memory 8 and the processing unit 3).

[0152] Bus 4 represents one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processor, or a local bus using any of a variety of bus architectures. Examples of these architectures include, but are not limited to, the Industry Standard Architecture (ISA) bus, the Micro Channel Architecture (MAC) bus, the Enhanced ISA bus, the Video Electronics Standards Association (VESA) local bus, and the Peripheral Component Interconnect (PCI) bus.

[0153] The electronic device 1 typically includes a variety of computer system readable media, which can be any available media that can be accessed by the electronic device 1, including volatile and non-volatile media, removable and non-removable media.

[0154] The system memory 8 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 9 and / or cache memory 10. The electronic device 1 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, the storage system 11 may be used to read and write non-removable, non-volatile magnetic media ( Figure 6 Not shown, often called a "hard drive"). Although Figure 6Not shown, a magnetic disk drive for reading and writing to a removable non-volatile magnetic disk (e.g., a "floppy disk"), and an optical disk drive for reading and writing to a removable non-volatile optical disk (e.g., a CD-ROM, DVD-ROM, or other optical media) may be provided. In these cases, each drive may be connected to bus 4 via one or more data medium interfaces. System memory 8 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of various embodiments of the present invention.

[0155] A program / utility 12 having a set (at least one) of program modules 13 may be stored, for example, in system memory 8. Such program modules 13 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data, each of which, or some combination thereof, may include an implementation of a network environment. Program modules 13 generally implement the functions and / or methods of the embodiments described herein.

[0156] The electronic device 1 may also communicate with one or more external devices 2 (e.g., a keyboard, a pointing device, a display 6, etc.), one or more devices that enable a user to interact with the electronic device 1, and / or any device that enables the electronic device 1 to communicate with one or more other computing devices (e.g., a network card, a modem, etc.). Such communication may be performed via an input / output (I / O) interface 7. Furthermore, the electronic device 1 may also communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) via a network adapter 5. Figure 6 As shown, the network adapter 5 communicates with other modules of the electronic device 1 via the bus 4. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device 1, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0157] The processing unit 3 executes various functional applications and data processing by running programs stored in the system memory 8 , such as implementing the control method considering the safety of the intelligent suspension provided in the embodiment of the present invention.

[0158] An embodiment of the present invention further provides a non-transitory computer-readable storage medium storing computer instructions, on which a computer program is stored. When the program is executed by a processor, the control method considering the safety of the intelligent suspension provided in all the inventive embodiments of the present application is implemented.

[0159] The computer storage medium of the embodiment of the present invention can adopt any combination of one or more computer-readable media. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. More specific examples (non-exhaustive list) of computer-readable storage media include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, apparatus or device.

[0160] A computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0161] The program code that comprises on the computer-readable medium can be transmitted with any appropriate medium, includes but not limited to wireless, electric wire, optical cable, RF etc., or above-mentioned any suitable combination.Can write the computer program code that is used to carry out the operation of the present invention with one or more programming languages or its combination, described programming language comprises object-oriented programming language such as Java, Smalltalk, C++, also comprises conventional procedural programming language--such as " C " language or similar programming language.Program code can be carried out on user's computer completely, partly on user's computer, carry out as an independent software package, partly on user's computer partly on remote computer, or carry out completely on remote computer or server.In the situation that relates to remote computer, remote computer can comprise local area network (LAN) or wide area network (WAN) to be connected to user's computer by the network of any kind, perhaps, can be connected to external computer (for example, utilize Internet service provider to come to connect by Internet).

[0162] An embodiment of the present invention further provides a computer program product, including a computer program, which, when executed by a processor, implements the above-mentioned control method considering the safety of the intelligent suspension.

[0163] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in the present disclosure can be achieved. This is not limited herein.

[0164] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.

Claims

1. A control method considering the safety of intelligent suspension, characterized in that: include: S1: Initialize the first evaluation network, the second evaluation network, the control strategy network and the experience pool; S2: The control strategy network determines a current control force output by the intelligent suspension according to the current state of the vehicle; applies the current control force to the vehicle to obtain a next state of the vehicle; The current state includes the acceleration and speed of the vehicle at the current moment, the dynamic travel of the smart suspension, and the speed difference between the sprung mass and the unsprung mass of the vehicle; S3: Determine the current comprehensive reward based on the current control power and the current state; The current state, the current control power, the current comprehensive reward, and the next state are stored as a set of samples in the experience pool; The current comprehensive reward includes the corresponding current acceleration reward, the current dynamic range reward, and the current tire dynamic load reward corresponding to the tire dynamic load of the vehicle; The current acceleration reward is a piecewise function, and the acceleration reward coefficient in the last stage is much smaller than the acceleration reward coefficients in all previous stages; The current process for determining acceleration thresholds at different stages of the acceleration reward includes: controlling a vehicle equipped with intelligent suspension at different speeds, road conditions, and vehicle loads, collecting acceleration data during driving, classifying and statistically analyzing the acceleration data by value, and plotting a probability distribution histogram of the acceleration values. Based on the distribution of the probability distribution histogram, the acceleration ranges with 50%, 80%, and 100% data coverage are determined, corresponding to the acceleration thresholds. The current dynamic travel reward is a piecewise function, and the dynamic travel reward coefficient in the final stage is significantly smaller than that in all previous stages. The process for determining the dynamic travel thresholds for different stages of the current dynamic travel reward includes: controlling a vehicle equipped with intelligent suspension to drive in different environments, collecting dynamic travel data during driving, classifying and statistically analyzing the dynamic travel data by numerical value, and plotting a probability distribution histogram of the dynamic travel values. Based on the distribution of the probability distribution histogram, the acceleration ranges covered by 100% + 20% and 100% + 50% of the data are determined, corresponding to the dynamic travel thresholds. The current tire dynamic load reward is a piecewise function, and the tire dynamic load reward coefficient in the final stage is significantly smaller than the tire dynamic load reward coefficients in all previous stages. The process for determining the tire dynamic load thresholds for different stages of the current tire dynamic load reward includes: controlling a vehicle equipped with an intelligent suspension to drive in different environments, collecting dynamic travel data during driving, classifying and statistically analyzing the dynamic travel data by numerical value, and plotting a probability distribution histogram of the dynamic travel values. Based on the distribution of the probability distribution histogram, the acceleration ranges covered by the data at 80%, 100% + 20%, 30%, and 100% + 50% are determined, corresponding to the tire dynamic load thresholds. S4: Repeat steps S2 to S3 multiple times to store multiple sets of samples in the experience pool; use the multiple sets of samples to train the first evaluation network, the second evaluation network, and the control strategy network to obtain a first evaluation model, a second evaluation model, and a control strategy model; S5: running the intelligent suspension, and controlling the intelligent suspension using the first evaluation model, the second evaluation model, and the control strategy model obtained in step S3.

2. The control method considering the safety of intelligent suspension according to claim 1, characterized in that: In step S3: The current acceleration reward is: ; in, represents the acceleration reward at time t, represents the acceleration of the vehicle at time t, 、 、 and represents the acceleration reward coefficient, 、 and Indicates the acceleration threshold; The current dynamic itinerary rewards are: ; in, represents the dynamic travel reward at time t, represents the dynamic travel of the smart suspension at time t, , represents the current travel distance of the vehicle, represents the travel of the unsprung mass of the vehicle at time t; 、 and Indicates the dynamic travel reward coefficient, and Indicates the dynamic travel threshold; The current tire dynamic load reward is: ; in, represents the tire dynamic load reward at time t, represents the dynamic wheel load of the vehicle at time t, , represents the road height incentive information of the vehicle at time t, 、 and represents the tire dynamic load bonus coefficient, 、 、 and Indicates the tire dynamic load threshold, The intermediate variable representing the range calculation of the current tire dynamic load reward, related to time t, is: ; in, represents the tire stiffness of the vehicle at time t, represents the body mass of the vehicle, represents the unsprung mass, g represents the acceleration due to gravity; The current comprehensive rewards are: ; in, represents the current comprehensive reward, 、 and Represents the current acceleration reward , Current dynamic itinerary rewards and the current tire dynamic load bonus The weight of .

3. The control method considering the safety of intelligent suspension according to claim 2, characterized in that: Step S2 further includes: A hard constraint is imposed on the control force output by the control strategy network: Determining the current control force using a delay mechanism according to the hard-constrained control force; The fatigue degree of the control force at the current moment is detected, and the control force at the current moment is adjusted according to the fatigue degree.

4. The control method considering the safety of the intelligent suspension according to claim 3, characterized in that: The hard constraint is performed by the following formula: ; in, represents the control force after hard constraints are applied to the control force output by the control strategy network at time t, represents the control strategy network, represents the network weight of the control strategy network, represents the interception function, and Respectively represent the minimum and maximum values of the control force at the current moment, Represents the state of the vehicle at time t.

5. The control method considering the safety of the intelligent suspension according to claim 4, characterized in that: According to the control force after the hard constraint, the process of determining the control force at the current moment is as follows: ; in, represents the control force output by the intelligent suspension at time t, Indicates the delay time, Indicates The control force output by the intelligent suspension at the moment; Delay time Determined by the following formula: ; in, represents the unit quantity determined by the maximum limit control force of the intelligent suspension, Represents unit time lag.

6. The control method considering the safety of the intelligent suspension according to claim 5, characterized in that: In the process of detecting the fatigue degree of the control force at the current moment and adjusting the control force at the current moment according to the fatigue degree: The working strength of the intelligent suspension is calculated by the following formula: ; in, Indicates the intensity of the work; The fatigue level is calculated by the following formula: ; in, represents the fatigue level, and represents the fatigue coefficient, Indicates the working time of the intelligent suspension; A fatigue threshold is set, and if the fatigue level at a current moment exceeds the fatigue threshold, the control force at the current moment is reduced.

7. The control method considering the safety of intelligent suspension according to claim 2, characterized in that: In step S4: In each training time step, multiple groups of samples are randomly drawn from the experience pool , calculate the corresponding target value by the following formula: ; in, represents the target value, represents the discount factor, represents the j-th evaluation network, represents the network weight of the j-th evaluation network, represents the state of the vehicle at time t, represents the control force output by the intelligent suspension at time t, represents the control strategy network, represents the network weight of the control strategy network; The loss function for training the two evaluation networks is calculated based on the target values: ; in, represents the loss function of the j-th evaluation network, and N represents the time step; the two evaluation networks are trained using their loss functions; The control strategy network is trained by the following formula: 。 8. The control method considering the safety of the intelligent suspension according to claim 7, characterized in that: During the training of the control strategy network and the two evaluation networks: the two evaluation networks are trained using the stochastic gradient descent method; the control strategy network is trained using the stochastic gradient ascent method; during the training of the control strategy network and / or the two evaluation networks, the current training state is evaluated after each set number of time steps, and the training learning rate is adjusted according to the evaluation results; the learning rate is adjusted within the range of [0.0001, 0.01].

9. A readable storage medium, characterized in that: The readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the control method considering the safety of the intelligent suspension according to any one of claims 1 to 8.

10. An electronic device, characterized in that: include: memory for storing computer programs; A processor is configured to implement the steps of the control method considering the safety of the intelligent suspension as claimed in any one of claims 1 to 8 when executing the computer program.