Intelligent suspension control method considering suspension system time lag, medium and equipment
By introducing meta-learning modules and fuzzy delay control rules into the intelligent suspension system, the time delay problem in deep reinforcement learning is solved, the universality and robustness of suspension control are improved, and the driving performance and ride comfort of the vehicle under complex road conditions are ensured.
Patent Information
- Application Number
- CN202510775062.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-06-11
Smart Images

Figure CN120269982A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of artificial intelligence, and particularly relates to an intelligent suspension control method, medium and device considering the time delay of the suspension system. Background Technique
[0002] With the continuous progress and innovation of microprocessor technology, sensor technology and actuator technology, intelligent suspension systems have become a research hotspot in the field of vehicle engineering. An intelligent suspension system can, according to driving operations and real-time road conditions, accurately control the actuator and adaptively adjust the output force, thereby significantly improving the ride comfort and driving performance of the vehicle, meeting the high standards of modern vehicles for comfort and safety. However, since an intelligent suspension system itself is a closed-loop system including sensors, controllers and actuators, there is inevitably an inherent time delay in the signal measurement or actuation process. In most cases, this time delay may be negligible and have little impact on the system performance. But in certain specific situations, when the time delay is comparable to or even longer than the control period, its impact cannot be ignored and must be carefully considered. The existence of time delay often weakens the control effect and may even cause instability in the control system, posing a potential threat to the driving safety of the vehicle.
[0003] Although the application of deep reinforcement learning in the field of suspension control has shown initial results, the research on the time delay problem is still insufficient. It should be noted that in sequential decision-making processes such as deep reinforcement learning, the negative impact of the time delay problem on the actual control effect is particularly significant. Therefore, exploring how to effectively address the time delay problem within the framework of deep reinforcement learning remains a new and urgent issue to be solved.
[0004] Current research mainly focuses on the feasibility analysis of deep reinforcement learning technology in suspension control, and is still in a relatively superficial exploration stage. It does not deeply consider the impact of the time delay of the intelligent suspension system on the control effect, nor does it propose an effective solution to the time delay problem. Summary of the Invention
[0005] In view of this, the present invention aims to provide an intelligent suspension control method, medium and device considering the time delay of the suspension system. The present invention incorporates a delay link, aiming to guide the agent to explore and obtain a more robust control strategy in a time-delay environment, thereby effectively suppressing the negative impact of uncertain delays on the performance of the intelligent suspension system and providing strong support for the efficient and stable control of the intelligent suspension system.
[0006] To achieve the above object, the technical solution of the present invention is realized as follows: An intelligent suspension control method considering the time delay of the suspension system, including: S1: Use meta - learning to pre - train the first evaluation network, the second evaluation network, and the control strategy network, respectively obtaining the first evaluation model, the second evaluation model, and the control strategy model, and initialize the experience pool; S2: The control strategy model obtained in step S1 determines the control force at the current moment output by the intelligent suspension according to the vehicle state at the current moment; Apply the control force at the current moment to the vehicle to obtain the vehicle state at the next moment; S3: Determine the current reward according to the control force at the current moment and the state at the current moment; Save the state at the current moment, the control force at the current moment, the current reward, and the state at the next moment as a set of samples in the experience pool; S4: Repeat steps S2 - S3 multiple times, and correspondingly save multiple sets of samples in the experience pool; Use multiple sets of samples to train the first evaluation model, the second evaluation model, and the control strategy model obtained in step S1 again to obtain the final first evaluation model, the final second evaluation model, and the final control strategy model; S5: Run the intelligent suspension, and use the final first evaluation model, the final second evaluation model, and the final control strategy model obtained in step S4 to control the intelligent suspension.
[0007] Further, step S1 includes: Initialize the first evaluation network, the second evaluation network, the control strategy network, and the experience pool; Construct a meta - training set, and the initialized first evaluation network, second evaluation network, and control strategy network perform meta - learning according to the meta - training set to respectively obtain the first evaluation model, the second evaluation model, and the control strategy model.
[0008] Further, in step S2, the process of applying the control force at the current moment to the vehicle includes: Perform hard constraints on the control force output by the control strategy model to obtain the control force at the current moment; Set the low - speed range, medium - speed range, and high - speed range of the vehicle; Set the low - acceleration range, medium - acceleration range, and high - acceleration range of the vehicle; Set the low - control - force change range, medium - control - force change range, and high - control - force change range corresponding to the control force output by the control strategy model; Measure the current speed and current acceleration of the vehicle, and calculate the control - force change rate between the control force at the current moment and the control force at the previous moment; Compare the current speed with the low - speed range, medium - speed range, and high - speed range, compare the current acceleration with the low - acceleration range, medium - acceleration range, and high - acceleration range, and compare the control - force change rate with the low - control - force change range, medium - control - force change range, and high - control - force change range. Determine the time delay when the control strategy model outputs the control force at the current moment according to the three comparison results; The control strategy model controls the intelligent suspension to apply the control force at the current moment to the vehicle according to the time delay.
[0009] Further, perform hard constraints on the control force output by the control strategy model through the following formula: ; Among them, represents the control strategy model, represents the network weight of the control strategy model, represents the truncation function, and respectively represent the minimum and maximum values that limit the control force at the current moment, represents the control force at the current moment.
[0010] Furthermore, in step S3: the state at the current moment Among them, and respectively represent the acceleration and speed of the vehicle at time t, represents the dynamic stroke of the intelligent suspension at time t, represents the vertical body displacement of the vehicle at time t, represents the displacement of the unsprung mass of the vehicle at time t; represents the speed difference between the sprung mass and the unsprung mass of the vehicle at time t, , represents the unsprung speed of the vehicle at time t; The current reward is ; Among them, represents the reward at time t, , , and represent the reward coefficients, represents the dynamic wheel load of the vehicle at time t, , represents the road surface displacement of the vehicle at time t.
[0011] Furthermore, the dynamic stroke satisfies: ; Among them, represents the suspension limit dynamic deflection of the intelligent suspension.
[0012] Furthermore, satisfies: ; Among them, represents the tire stiffness of the vehicle; represents the static tire load of the vehicle, which is: ; Among them, represents the acceleration due to gravity, represents the body mass of the vehicle, represents the unsprung mass.
[0013] Further, in step S4: In each training time step, multiple groups of samples are randomly drawn from the experience pool , and the corresponding target value is calculated by the following formula: ; where, represents the target value, represents the discount factor, represents the j-th evaluation network, represents the network weights of the j-th evaluation network; The loss functions of the two evaluation networks are calculated and trained in combination with the target value: ; where, represents the loss function of the j-th evaluation network, and N represents the time step; The two evaluation networks are trained using the loss functions of the two evaluation networks; When the training processes of the two evaluation networks meet the trigger conditions of the preset training control strategy model, the control strategy model is trained by the following formula: .
[0014] A readable storage medium has a computer program stored thereon, and when the computer program is executed by a processor, it implements the steps of the intelligent suspension control method considering the time delay of the suspension system provided by the present invention.
[0015] An electronic device includes: a memory for storing a computer program; a processor for implementing the steps of the intelligent suspension control method considering the time delay of the suspension system provided by the present invention when executing the computer program.
[0016] Compared with the prior art, the present invention can achieve the following beneficial effects: (1) In the intelligent suspension control method considering the time delay of the suspension system of the present invention, a meta-learning module for complex road conditions is introduced, and the evaluation network and the policy network are pre-trained in the initial stage of training. By constructing a meta-training set covering a variety of typical complex road condition scenarios, the network is made to learn the common features and control strategies under different complex road conditions using the meta-learning algorithm, greatly enhancing the universality of the control strategy. During actual driving, in the face of various complex road conditions (such as potholes, bumps, muddy sections, etc.), the control strategy of the present invention can quickly adapt, effectively improving the driving performance and ride comfort of the vehicle under complex road conditions, and further broadening the application scope of the present invention under different road conditions; (2) In the intelligent suspension control method considering the time delay of the suspension system according to the present invention, the time delay characteristics of the suspension system, including various situations such as deterministic delay, semi-regular delay, and uncertain delay, are innovatively considered in the control strategy design. By constructing multiple groups of simulation experiments to test the control performance of the algorithm under different time delay conditions, various time delay characteristics of the control system and the uncertainties under their influence are fully considered, making the simulation results closer to the actual working conditions. At the same time, the uniquely designed action delay mechanism, especially the adoption of fuzzy delay control rules, can dynamically adjust the action delay time according to the vehicle driving state, significantly improving the robustness and reliability of the control strategy. This advantage enables the present invention to maintain more stable and effective suspension control performance when dealing with complex and changeable driving environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The drawings forming a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments and descriptions thereof of the present invention are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings: Figure 1 is a schematic flow chart of the intelligent suspension control method considering the time delay of the suspension system according to the embodiment of the present invention; Figure 2 is a schematic framework diagram of the intelligent suspension control method considering the time delay of the suspension system according to the embodiment of the present invention; Figure 3 is a schematic diagram of the control strategy network according to the embodiment of the present invention; Figure 4 is a schematic diagram of the two evaluation networks according to the embodiment of the present invention; Figure 5 is a schematic diagram of the dynamic model of the intelligent suspension system according to the embodiment of the present invention; Figure 6 is a schematic diagram of the structure of the electronic device according to the embodiment of the present invention.
[0018] DESCRIPTION OF THE REFERENCE NUMERALS: 1, electronic device; 2, external device; 3, processing unit; 4, bus; 5, network adapter; 6, display; 7, (I / O) interface; 8, system memory; 9, random access memory; 10, cache memory; 11, storage system; 12, utility tool; 13, program module. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0019] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and do not constitute a limitation to the present invention.
[0020] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments may be combined with each other.
[0021] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "center", "longitudinal", "transverse", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus cannot be understood as a limitation to the present invention. In addition, the terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first", "second", etc. may explicitly or implicitly include one or more of such features. In the description of the present invention, unless otherwise specified, the meaning of "a plurality" is two or more.
[0022] In the description of the present invention, it should be noted that unless otherwise clearly specified and defined, the terms "installed", "connected", "connected" should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral connection; it may be a mechanical connection or an electrical connection; it may be directly connected, or indirectly connected through an intermediate medium, and it may be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood through specific situations.
[0023] The present invention will be described in detail below with reference to the drawings and in combination with embodiments.
[0024] As Figures 1 to 4 shown, the intelligent suspension control method considering the time delay of the suspension system described in the embodiment of the present invention includes: S1: Use meta - learning to pre - train the first evaluation network, the second evaluation network, and the control strategy network, respectively obtain the first evaluation model, the second evaluation model, and the control strategy model, and initialize the experience pool.
[0025] Meta - learning, as a method of learning how to learn, can enable the model to quickly adapt to new tasks and new environments. Under complex road conditions, different road surface conditions (such as potholes, bumps, muddy sections, etc.) have huge differences in the control requirements for the suspension system. In order to improve the universality of the control strategy, the present invention introduces a meta - learning module to pre - train the evaluation network and the strategy network in the initial stage of training.
[0026] In some embodiments, step S1 includes: S11: Initialize the first evaluation network, the second evaluation network, the control strategy network, and the experience pool.
[0027] In a certain embodiment, the control strategy network is as Figure 3 shown. The state of the vehicle is input into a fully connected layer composed of 128 neurons for processing. The processed data is processed by the ReLU activation function and then input into a fully connected layer composed of 256 neurons; the output data is processed by the ReLU activation function and then input into a fully connected layer composed of 1 neuron; finally, the output data is activated by the Tanh activation layer and scaled by the data scaling layer to obtain the control force output by the intelligent suspension. The structures of the first evaluation network and the second evaluation network are as Figure 4 shown. The state of the vehicle is input into a fully connected layer composed of 128 neurons for processing. The processed data is processed by the ReLU activation function and then input into a fully connected layer composed of 256 neurons to obtain the first feature; at the same time, the output of the control strategy network is processed by a fully connected layer composed of 256 neurons to obtain the second feature. After adding the corresponding elements of the first feature and the second feature and processing by the ReLU activation function, it is input into a fully connected layer composed of 1 neuron, and then the evaluation result is output. The initialization operations of the three networks include: the network weights in the first evaluation network , the network weights in the first evaluation network , and the network weights of the control strategy network are randomly initialized. s represents the state of the vehicle, and a represents the control force output by the intelligent suspension.
[0028] S12: Construct a meta-training set. The initialized first evaluation network, second evaluation network, and control strategy network perform meta-learning according to the meta-training set to obtain the corresponding first evaluation model, second evaluation model, and control strategy model.
[0029] In a certain embodiment, a large amount of vehicle driving data under different complex road conditions is collected, including but not limited to vehicle body acceleration, vehicle body speed, suspension dynamic deflection, road condition information, etc., to construct a meta-training set. The meta-training set covers a variety of typical complex road condition scenarios to fully simulate various situations that may be encountered in actual driving; the first evaluation network, the second evaluation network, and the control strategy network are pre-trained using a gradient-based meta-learning algorithm (MAML: Model-Agnostic Meta-Learning). The goal of pre-training is to find a set of initial network parameters so that the network can quickly adjust the parameters with a small number of samples when facing new complex road conditions and achieve a better control effect.
[0030] During the pre-training process, set the meta-learning rate to 0.001 and the meta-training steps to 500 steps. In each meta-training iteration, randomly select multiple subtasks with different road conditions from the meta-training set. Each subtask contains a certain number of training samples. For each subtask, use the evaluation network and the policy network to train on this subtask, calculate the loss, and update the network parameters. Through multiple iterations, enable the network to gradually learn the common features and control strategies under different complex road conditions and complete the pre-training process.
[0031] S2: The control policy model obtained in step S1 determines the control force at the current moment output by the intelligent suspension according to the current moment state of the vehicle; apply the control force at the current moment to the vehicle to obtain the next moment state of the vehicle.
[0032] In some embodiments, the process of applying the control force at the current moment to the vehicle in step S2 includes: Perform a hard constraint on the control force output by the control policy model to obtain the control force at the current moment; Set the low-speed interval, medium-speed interval, and high-speed interval of the vehicle; Set the low acceleration interval, medium acceleration interval, and high acceleration interval of the vehicle; Set the low control force change interval, medium control force change interval, and high control force change interval corresponding to the control force output by the control policy model; Measure the current speed and current acceleration of the vehicle, and calculate the control force change rate between the control force at the current moment and the control force at the previous moment; compare the current speed with the low-speed interval, medium-speed interval, and high-speed interval, compare the current acceleration with the low acceleration interval, medium acceleration interval, and high acceleration interval, and compare the control force change rate with the low control force change interval, medium control force change interval, and high control force change interval. Determine the time delay when the control policy model outputs the control force at the current moment according to the three comparison results; The control policy model controls the intelligent suspension to apply the control force at the current moment to the vehicle according to the time delay.
[0033] In a certain embodiment, the hard constraint on the control force output by the control policy model is performed by the following formula: ; Where represents the truncation function, and respectively represent the minimum value and the maximum value that limit the control force at the current moment, represents the control force at the current moment. Among them, the minimum value of the control force at the current moment and the maximum value The value is obtained according to the limit of the maximum force that the intelligent suspension can output. In a certain embodiment, the minimum value is set , the maximum value .
[0034] In a certain embodiment, in the process of determining the time delay when the control strategy model outputs the control force at the current moment: Set the low-speed range of the vehicle as [0, 45 km / h], the medium-speed range (45 km / h, 90 km / h], and the high-speed range as (90 km / h, ∞); The low acceleration of the vehicle can reflect the bumpiness of the road surface. Therefore, set three acceleration ranges of the vehicle to reflect different bumpiness degrees of the road surface, that is, set the low acceleration range of the vehicle as [0, 0.5 m / s 2 , corresponding to a slight bumpiness degree; the medium acceleration range is (0.5 m / s 2 , 1.5 m / s 2 , corresponding to a medium bumpiness degree; the high acceleration range is (1.5 m / s 2 , ∞), corresponding to a severe bumpiness degree; The control force change rate can reflect the change degree of the output control force and can be obtained as follows: ; Wherein, represents the control force change rate at time t, represents the time interval. Based on the above formula, it can be set The low control force change range is [0, 10 kN / s], corresponding to a slow change of the control force; the medium control force change range is (10 kN / s, 30 kN / s], corresponding to a medium change of the control force; and the high control force change range is (30 kN / s, ∞), corresponding to a severe change of the control force.
[0035] After completing the above range settings for speed, acceleration, and control force change rate, measure the current speed and current acceleration of the vehicle, and calculate the control force change rate between the control force at the current moment and the control force at the previous moment; compare the current speed with the low-speed range, medium-speed range, and high-speed range, compare the current acceleration with the low acceleration range, medium acceleration range, and high acceleration range, and compare the control force change rate with the low control force change range, medium control force change range, and high control force change range. Determine the time delay when the control strategy model outputs the control force at the current moment according to the three comparison results. The specific process is as follows: If the current vehicle speed is low, the road surface bumpiness is slightly bumpy, and the control signal change rate is slowly changing, the delay time is a short delay, that is, a delay of 0.01 s. The reason is that in the case of low speed, slightly bumpy road surface and slowly changing control signal, the system response is relatively easy, and a short delay time can meet the control requirements while ensuring the smoothness of the vehicle; If the current vehicle speed is low, the road surface bumpiness is slightly bumpy, and the control signal change rate is moderately changing, the delay time is a short delay, that is, a delay of 0.015 s. Although the change of the control signal has accelerated at this time, the low speed and slightly bumpy road conditions still allow the system to make an effective response within a short delay, and the delay time is appropriately increased to cope with the change of the control signal; If the current vehicle speed is low, the road surface bumpiness is slightly bumpy, and the control signal change rate is rapidly changing, the delay time is a medium delay, that is, a delay of 0.02 s. Since the control signal changes rapidly, it takes a certain amount of time for the system to adapt. At the same time, the low speed and slightly bumpy road conditions will not cause serious consequences due to a slightly longer delay, so it is set as a medium delay; If the current vehicle speed is low, the road surface bumpiness is moderately bumpy, and the control signal change rate is slowly changing, the delay time is a medium delay, that is, a delay of 0.02 s. Moderate bumpiness requires the system to adjust more promptly, but the system has more time to process when driving at low speed, so a medium delay is adopted to balance the control effect and system stability; If the current vehicle speed is low, the road surface bumpiness is moderately bumpy, and the control signal change rate is moderately changing, the delay time is a medium delay, that is, a delay of 0.025 s. In this case, both moderate bumpiness and a moderately changing control signal rate require the system to have a certain response time, and the delay time is appropriately increased to ensure the accuracy of control; If the current vehicle speed is low, the road surface bumpiness is moderately bumpy, and the control signal change rate is rapidly changing, the delay time is a long delay, that is, a delay of 0.03 s. The rapid change of the control signal and moderate bumpiness require a higher response from the system. A long delay helps the system better integrate information and make more appropriate control decisions; If the current vehicle speed is low, the road surface bumpiness is severely bumpy, and the control signal change rate is slowly changing, the delay time is a medium delay, that is, a delay of 0.025 s. Severe bumpiness requires the system to respond quickly, but the control signal changes slowly, so a medium delay is adopted to balance system stability and the suppression of bumpiness; If the current vehicle speed is low, the road surface bumpiness is severely bumpy, and the control signal change rate is moderately changing, the delay time is a long delay, that is, a delay of 0.035 s. Severe bumpiness and a moderately changing control signal rate require the system to have enough time to adjust, and a long delay can better coordinate the influence of both; If the current vehicle speed is low, the road surface bumpiness is severely bumpy, and the control signal change rate is changing violently, the delay time is a long delay, that is, a delay of 0.04 s. In this case, the system faces greater challenges, and the long delay helps the system to respond more accurately to severe bumpiness and violent control signal changes; If the current vehicle speed is medium, the road surface bumpiness is slightly bumpy, and the control signal change rate is changing slowly, the delay time is a short delay, that is, a delay of 0.012 s. Driving at a medium speed and slightly bumpy road surface endow the system with a certain response ability, and the control signal changes slowly, so a short delay slightly longer than that in the case of low speed, slightly bumpy, and slow change is adopted; If the current vehicle speed is medium, the road surface bumpiness is slightly bumpy, and the control signal change rate is changing moderately, the delay time is a medium delay, that is, a delay of 0.02 s. Under the road conditions of medium speed and slightly bumpy road surface, a medium control signal change rate requires the system to have a certain reaction time, and the medium delay can balance the control accuracy and response speed; If the current vehicle speed is medium, the road surface bumpiness is slightly bumpy, and the control signal change rate is changing violently, the delay time is a medium delay, that is, a delay of 0.025 s. Although the road surface is slightly bumpy, the violent control signal change requires the system to have enough time to process, and the medium delay can ensure the stable response of the system; If the current vehicle speed is medium, the road surface bumpiness is moderately bumpy, and the control signal change rate is changing slowly, the delay time is a medium delay, that is, a delay of 0.022 s. Moderate bumpiness and medium-speed driving require the system to have a certain response speed, and the control signal changes slowly. The medium delay can meet the control requirements and ensure the vehicle's smoothness; If the current vehicle speed is medium, the road surface bumpiness is moderately bumpy, and the control signal change rate is changing moderately, the delay time is a medium delay, that is, a delay of 0.028 s. In this case, both moderate bumpiness and moderate control signal change rate require the system to respond in a timely manner. Appropriately increasing the delay time can improve the control effect; If the current vehicle speed is medium, the road surface bumpiness is moderately bumpy, and the control signal change rate is changing violently, the delay time is a long delay, that is, a delay of 0.035 s. The violent control signal change and moderate bumpiness require high requirements for the system, and the long delay helps the system to make more appropriate control decisions; If the current vehicle speed is medium, the road surface bumpiness is severely bumpy, and the control signal change rate is changing slowly, the delay time is a long delay, that is, a delay of 0.03 s. Severe bumpiness requires the system to respond quickly. Even if the control signal changes slowly, the long delay also helps the system to better cope with the bumpiness; If the current vehicle speed is medium, the road surface bumpiness is severely bumpy, and the control signal change rate is changing moderately, the delay time is a long delay, that is, a delay of 0.04 s. Severe bumpiness and moderate control signal change rate require the system to have sufficient time to adjust, and the long delay can improve the system's adaptability; If the current vehicle speed is medium, the road bumps are severe, and the control signal change rate changes dramatically, the delay time is long, that is, 0.045s. In this extreme case, the long delay allows the system to fully integrate information and make more accurate control actions; If the current vehicle speed is high, the road bumps are slightly bumpy, and the control signal change rate is slow, the delay time is medium, that is, 0.02s. When driving at high speed, the system response time is relatively tight. Even with slight bumps and slow control signal changes, a certain delay is required to ensure the accuracy of control; If the current vehicle speed is high, the road bumps are slightly bumpy, and the control signal change rate is medium, then the delay time is medium delay, that is, 0.025s. Under high speed and slightly bumpy road conditions, the medium control signal change rate requires the system to have an appropriate reaction time. Medium delay can balance control and vehicle stability; If the current vehicle speed is high, the road bumps are slightly bumpy, and the control signal change rate is drastic, the delay time is long, that is, 0.03s. Drastic control signal changes require more time to process when driving at high speeds, and long delays help the system respond stably; If the current vehicle speed is high, the road bumps are moderate, and the control signal change rate is slow, the delay time is long, that is, 0.035s. High speed and moderate bumps require high system response, and the slow change of the control signal cannot be ignored. Long delay can ensure that the system responds effectively; If the current vehicle speed is high, the road bumps are moderate, and the control signal change rate is moderate, then the delay time is long, that is, 0.04s. In this case, high speed, moderate bumps, and moderate control signal change rate all require the system to have sufficient time to respond, and long delay can improve the control effect; If the current vehicle speed is high, the road bumps are moderate, and the control signal change rate is drastic, the delay time is long, that is, 0.05s. Drastic control signal changes and high-speed, moderately bumpy road conditions pose great challenges to the system, and long delays allow the system to adjust the suspension more accurately; If the current vehicle speed is high, the road bumps are severe, and the control signal change rate is slow, the delay time is long, that is, 0.04s. Severe bumps and high-speed driving require the system to respond quickly. Even if the control signal changes slowly, a long delay helps stabilize the system. If the current vehicle speed is high, the road bumps are severe, and the control signal change rate is medium, the delay time is long, that is, 0.05s. Severe bumps, high-speed driving, and medium control signal change rates require the system to be fully adjusted, and long delays can ensure the adaptability of the system; If the current vehicle speed is high, the road surface bumpiness is severe, and the control signal change rate is highly variable, the delay time is a long delay, i.e., 0.06 s. This is the most complex working condition, and the long delay allows the system to better handle the influence of various factors and ensure the driving safety and comfort of the vehicle.
[0036] The centroid method is used as the defuzzification method to convert the delay time level of the fuzzy output into an actual delay time value. Specifically, after a series of judgments based on conditions such as vehicle speed, acceleration, and control force change rate, a fuzzy output level of the delay time (such as short delay, medium delay, long delay, etc.) will be obtained. The centroid method is a commonly used defuzzification method for converting fuzzy outputs into precise numerical values. Specifically, for each delay time level, a corresponding membership function (which describes the degree to which a certain value belongs to this level) will be preset. Taking the delay time as an example, short delay, medium delay, and long delay respectively correspond to different membership function curves. When the system outputs a fuzzy delay time level, a "centroid" position will be calculated according to these membership functions. The abscissa value of this "centroid" position is the actual delay time value obtained through the centroid method. For example, assuming that the time range covered by the membership function corresponding to the short delay is from 0 to 0.02 s, the medium delay corresponds to 0.02 s to 0.04 s, and the long delay corresponds to 0.04 s to 0.06 s. When the system's fuzzy output is a result between the short delay and the medium delay, by calculating the centroid of the membership functions of these two levels, an accurate delay time value, such as 0.025 s, can be obtained. Through the above fuzzy delay control rules, the action delay time is dynamically adjusted according to the current driving state of the vehicle, enabling the system to more reasonably handle the time delay problem under different working conditions and improving the control effect.
[0037] In one embodiment, a disturbance quantity that satisfies a Gaussian random distribution is added to the control force output to the intelligent suspension That is, the control force at the current moment output by the intelligent suspension is: .
[0038] S3: Determine the current reward according to the control force at the current moment and the state at the current moment; save the state at the current moment, the control force at the current moment, the current reward, and the state at the next moment as a set of samples in the experience pool.
[0039] In some embodiments, the dynamic model of the intelligent suspension system is as Figure 5 shown, Figure 5 where represents the vertical displacement of the vehicle body at the current moment, represents the displacement below the spring at the current moment of the vehicle, represents the body mass of the vehicle, Represents the unsprung mass, represents the road surface displacement at the current moment of the vehicle, represents the tire stiffness of the vehicle, represents the suspension spring stiffness of the intelligent suspension, represents the damping coefficient of the intelligent suspension. State at the current moment , where, and respectively represent the acceleration and speed of the vehicle at time t, which are obtained by taking the first derivative and the second derivative of the vertical body displacement of the vehicle at time t respectively, represents the dynamic stroke of the intelligent suspension at time t; , represents the unsprung speed of the vehicle at time t, which is obtained by taking the first derivative of the unsprung displacement of the vehicle at time t . In one embodiment, in order to improve the generalization performance of the control strategy model, the state is normalized, that is, the state at the current moment , where , , and represent the normalization coefficients, which are adaptively selected and adjusted according to the actual situation, and can be set to 2, 0.2, 0.4 and 0.15 in one embodiment.
[0040] In some embodiments, the current reward is: ; where, represents the reward at time t, , , and represent the reward coefficients, represents the wheel dynamic load of the vehicle at time t. Specifically, . Reward coefficient , , and are adaptively selected and adjusted according to the actual situation. In one embodiment, the reward coefficients , , and are respectively 0.7, 0.1, 0.1 and 0.1. It should be noted that the reward in the training stage refers to the state information of the system after applying the control force with a delay.
[0041] Among them, the dynamic stroke and respectively satisfy: ; ; Among them, represents the suspension limit dynamic deflection of the intelligent suspension; represents the static load of the vehicle's tire, which is: ; Among them, represents the acceleration due to gravity. In a certain embodiment, the suspension limit dynamic deflection is 0.15 m.
[0042] S4: Repeat steps S2 - S3 multiple times, and correspondingly save multiple groups of samples in the experience pool; use the multiple groups of samples to retrain the first evaluation model, the second evaluation model, and the control strategy model obtained in step S1 to obtain the final first evaluation model, the final second evaluation model, and the final control strategy model.
[0043] In some embodiments, the training process in step S4 includes: In each training time step, set the number of training episodes to M, and the number of time steps for each episode to N. Randomly extract multiple groups of samples from the experience pool , and calculate the corresponding target value through the following formula: ; Among them, represents the target value, represents the discount factor, represents the j - th evaluation network, represents the network weight of the j - th evaluation network; Combine the target value to calculate the loss functions for training the two evaluation networks: ; Among them, represents the loss function of the j - th evaluation network, and N represents the time step; Use the loss functions of the two evaluation networks to train the two evaluation networks; When the training process of the two evaluation networks meets the trigger condition for presetting the training control strategy model, train the control strategy model through the following formula: .
[0044] In a certain embodiment, the number of training episodes M is set to 4000, the number of time steps N is set to 2000. In each training time step, randomly extract 256 groups of samples from the experience pool , and the discount factor takes a value of 0.99. Use the stochastic gradient descent method, and combine the loss function to train the two evaluation networks as follows: ; During the training process, the network weights of the two evaluation networks are updated using the soft update method, i.e.: And Are updated, that is: ; Among them, And Respectively represent the network weights of the two evaluation networks after update, Represents the soft update frequency, with a value of 0.001; The preset trigger condition is that the two evaluation networks are updated 2 times, that is, when both evaluation networks are updated 2 times, the control strategy network is trained using the stochastic gradient ascent method. During the training process, the network weights of the control strategy network are also updated using the soft update method, i.e.: Are updated, that is: ; Among them, Represents the network weight of the control strategy network after update. Here, the soft update frequency Still takes the value of 0.001.
[0045] S5: Run the intelligent suspension and use the final first evaluation model, the final second evaluation model, and the final control strategy model obtained in step S4 to control the intelligent suspension.
[0046] Based on the intelligent suspension control method considering the suspension system time delay provided by the present invention, an intelligent suspension control system considering the suspension system time delay is also provided. The system includes an intelligent agent, a vehicle or test bench equipped with an intelligent suspension and related sensors, a state observer and state estimator, and an endogenous reward function. Among them, the intelligent agent controls the intelligent suspension, and the intelligent suspension control method considering the suspension system time delay provided by the present invention is integrated in the intelligent agent. The related sensors include an acceleration sensor and a displacement sensor inertial measurement unit (IMU). The operation process of the intelligent suspension control system is as follows: The intelligent agent obtains data such as vehicle body acceleration, vehicle body speed, suspension dynamic deflection, and the derivative of suspension dynamic deflection from the vehicle system through sensors, and combines these data into the current state information of the vehicle through a state observer and a state estimator. Then, the intelligent agent decides how much control force the intelligent suspension should output in the current state based on this state information. Under the action of the control force, the vehicle state changes, and the intelligent agent generates a reward value for evaluating its control action according to the new system state and the endogenous reward function. The intelligent agent performs self-iteration and control strategy optimization with reference to the reward value according to the intelligent suspension control method considering the suspension system time delay provided by the present invention.
[0047] Figure 6 It is a schematic structural diagram of an electronic device 1 provided in an embodiment of the present invention. Figure 6A block diagram of an exemplary electronic device 1 suitable for implementing the embodiments of the present invention is shown. Figure 6 The electronic device 1 shown is merely an example and should not impose any limitation on the functions and usage scope of the embodiments of the present invention.
[0048] As Figure 6 shown, the electronic device 1 is presented in the form of a general-purpose computing device. The electronic device 1 is intended to represent various forms of digital computers, such as, for example, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as, for example, personal digital processors, cellular telephones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the invention described and / or claimed herein.
[0049] The components of the electronic device 1 may include, but are not limited to: one or more processors or processing units 3, a system memory 8, and a bus 4 that couples different system components (including the system memory 8 and the processing unit 3).
[0050] The bus 4 represents one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any of a variety of bus structures. By way of example, these architectures include, but are not limited to, Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MAC) bus, Enhanced ISA bus, Video Electronics Standards Association (VESA) local bus, and Peripheral Component Interconnect (PCI) bus.
[0051] The electronic device 1 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by the electronic device 1, including volatile and non-volatile media, removable and non-removable media.
[0052] The system memory 8 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 9 and / or cache memory 10. The electronic device 1 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, a storage system 11 may be used for reading and writing on a non-removable, non-volatile magnetic medium ( Figure 6 not shown, commonly referred to as a "hard disk drive"). Although Figure 6Not shown in the figure, a disk drive for reading and writing to a removable non-volatile disk (such as a "floppy disk") and an optical disk drive for reading and writing to a removable non-volatile optical disk (such as a CD-ROM, DVD-ROM, or other optical medium) can be provided. In these cases, each drive can be connected to the bus 4 through one or more data medium interfaces. The system memory 8 can include at least one program product having a set (such as at least one) of program modules configured to perform the functions of the embodiments of the present invention.
[0053] A program / utility 12 having a set (at least one) of program modules 13 can be stored, for example, in the system memory 8. Such program modules 13 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment. The program modules 13 generally perform the functions and / or methods in the embodiments described in the present invention.
[0054] The electronic device 1 can also communicate with one or more external devices 2 (such as a keyboard, a pointing device, a display 6, etc.), and can also communicate with one or more devices that enable a user to interact with the electronic device 1, and / or communicate with any device that enables the electronic device 1 to communicate with one or more other computing devices (such as a network card, a modem, etc.). Such communication can be carried out through the input / output (I / O) interface 7. And the electronic device 1 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through the network adapter 5. As Figure 6 shown, the network adapter 5 communicates with other modules of the electronic device 1 through the bus 4. It should be understood that although not shown in the figure, other hardware and / or software modules can be used in combination with the electronic device 1, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0055] The processing unit 3 executes various functional applications and data processing by running the programs stored in the system memory 8, such as implementing the intelligent suspension control method considering the suspension system time delay provided by the embodiments of the present invention.
[0056] In the embodiments of the present invention, a non-transitory computer-readable storage medium storing computer instructions is also provided, on which a computer program is stored. When the program is executed by a processor, it is the intelligent suspension control method considering the suspension system time delay provided by all the embodiments of the present application.
[0057] The computer storage medium of the embodiments of the present invention may adopt any combination of one or more computer-readable media. The computer-readable media may be computer-readable signal media or computer-readable storage media. More specific examples (a non-exhaustive list) of the computer-readable storage media include: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In this document, the computer-readable storage media may be any tangible medium that contains or stores a program, and the program may be used by or in combination with an instruction execution system, apparatus, or device.
[0058] The computer-readable signal media may include a data signal propagated in a baseband or as part of a carrier wave, in which computer-readable program code is carried. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. The computer-readable signal media may also be any computer-readable media other than the computer-readable storage media, and the computer-readable media may send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device.
[0059] The program code contained on the computer-readable media may be transmitted by any appropriate medium, including but not limited to wireless, wire, optical cable, RF, etc., or any suitable combination of the foregoing. The computer program code for performing the operations of the present invention may be written in one or more programming languages or combinations thereof, and the programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages - such as the "C" language or similar programming languages. The program code may be executed entirely on the user computer, partially on the user computer, executed as an independent software package, partially on the user computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user computer through any type of network including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).
[0060] The embodiments of the present invention further provide a computer program product, including a computer program, and the computer program, when executed by a processor, implements the intelligent suspension control method considering the suspension system time lag as described above.
[0061] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in the disclosure of the present invention can be executed in parallel, sequentially, or in different orders, as long as the desired results of the technical solution disclosed in the present invention can be achieved, and no limitations are imposed herein.
[0062] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. An intelligent suspension control method considering the time delay of the suspension system, characterized in that, Including: S1: Use meta - learning to pre - train the first evaluation network, the second evaluation network, and the control strategy network, respectively obtain the first evaluation model, the second evaluation model, and the control strategy model, and initialize the experience pool; S2: The control strategy model obtained in step S1 determines the control force at the current moment output by the intelligent suspension according to the state of the vehicle at the current moment; Apply the control force at the current moment to the vehicle to obtain the state of the vehicle at the next moment; S3: Determine the current reward according to the control force at the current moment and the state at the current moment; Save the state at the current moment, the control force at the current moment, the current reward, and the state at the next moment as a group of samples in the experience pool; S4: Repeat steps S2 - S3 multiple times, and correspondingly save multiple groups of samples in the experience pool; Use multiple groups of samples to train the first evaluation model, the second evaluation model, and the control strategy model obtained in step S1 again to obtain the final first evaluation model, the final second evaluation model, and the final control strategy model; S5: Run the intelligent suspension and use the final first evaluation model, the final second evaluation model, and the final control strategy model obtained in step S4 to control the intelligent suspension.
2. The intelligent suspension control method considering the time delay of the suspension system according to claim 1, characterized in that Step S1 includes: Initialize the first evaluation network, the second evaluation network, the control strategy network, and the experience pool; Construct a meta - training set, and the initialized first evaluation network, second evaluation network, and control strategy network perform meta - learning according to the meta - training set to respectively obtain the first evaluation model, the second evaluation model, and the control strategy model.
3. The intelligent suspension control method considering the time delay of the suspension system according to claim 1, characterized in that In step S2, the process of applying the control force at the current moment to the vehicle includes: Perform hard constraints on the control force output by the control strategy model to obtain the control force at the current moment; Set the low - speed range, medium - speed range, and high - speed range of the vehicle; Set the low - acceleration range, medium - acceleration range, and high - acceleration range of the vehicle; Set the low - control - force change range, medium - control - force change range, and high - control - force change range corresponding to the control force output by the control strategy model; Measure the current speed and current acceleration of the vehicle, and calculate the control - force change rate between the control force at the current moment and the control force at the previous moment; Compare the current speed with the low - speed range, the medium - speed range, and the high - speed range, compare the current acceleration with the low - acceleration range, the medium - acceleration range, and the high - acceleration range, and compare the control - force change rate with the low - control - force change range, the medium - control - force change range, and the high - control - force change range. Determine the time delay when the control strategy model outputs the control force at the current moment according to the three comparison results; The control strategy model controls the intelligent suspension to apply the control force at the current moment to the vehicle according to the time delay.
4. The intelligent suspension control method considering the time delay of the suspension system according to claim 3, characterized in that Perform hard constraints on the control force output by the control strategy model through the following formula: ; Among them, represents the control strategy model, represents the network weights of the control strategy model, represents the truncation function, and respectively represent the minimum and maximum values that limit the control force at the current moment, represents the control force at the current moment.
5. The intelligent suspension control method considering the time delay of the suspension system according to claim 4, characterized in that In step S3: Current moment state , where and respectively represent the acceleration and speed of the vehicle at time t, represents the dynamic stroke of the intelligent suspension at time t, represents the vertical displacement of the vehicle body at time t, represents the displacement below the spring of the vehicle at time t; represents the speed difference between the mass above the spring and the mass below the spring of the vehicle at time t, , represents the speed below the spring of the vehicle at time t; The current reward is: ; Among them, represents the reward at time t, , , and represent the reward coefficients, represents the wheel dynamic load of the vehicle at time t, , represents the road surface displacement of the vehicle at time t.
6. The intelligent suspension control method considering the time delay of the suspension system according to claim 5, characterized in that, The dynamic stroke satisfies: ; Among them, represents the suspension limit dynamic deflection of the intelligent suspension.
7. The intelligent suspension control method considering the time delay of the suspension system according to claim 5, characterized in that The said Satisfies: ; Among them, represents the tire stiffness of the vehicle; represents the static load of the tire of the vehicle, which is: ; Wherein, represents the gravitational acceleration, represents the body mass of the vehicle, represents the unsprung mass.
8. The intelligent suspension control method considering the time delay of the suspension system according to claim 5, characterized in that In step S4: In each time step of training, multiple groups of samples are randomly drawn from the experience pool , and the corresponding target values are calculated by the following formula: ; Among them, represents the target value, represents the discount factor, represents the j-th evaluation network, represents the network weight of the j-th evaluation network; Calculate the loss functions for training two evaluation networks in combination with the target values: ; Among them, represents the loss function of the j-th evaluation network, and N represents the time step; the two evaluation networks are trained using the loss functions of the two evaluation networks. When the training processes of the two evaluation networks meet the trigger conditions for presetting the training of the control policy model, train the control policy model through the following formula: 。 9. A readable storage medium, characterized in that, A computer program is stored on the readable storage medium, and when the computer program is executed by a processor, the steps of the intelligent suspension control method considering the time delay of the suspension system as described in any one of claims 1 to 8 are implemented.
10. An electronic device, characterized in that, Including: A memory for storing a computer program; A processor for implementing the steps of the intelligent suspension control method considering the time delay of the suspension system as described in any one of claims 1 to 8 when executing the computer program.
Citation Information
Patent Citations
Active suspension control method based on depth determinacy strategy gradient
CN112158045A
Structural vibration control method based on reinforcement learning, medium and equipment
CN112698572A
Self-driving vehicle steering and suspension cooperative control method based on reinforcement learning
CN119975527A
Method for optimizing PID control parameters of semi-active suspension of vehicle
WO2024125584A1