Intelligent suspension control method combined with deterministic experience tracking, medium and equipment
Through deterministic experience tracking mechanism and disturbance optimization, the problem of low data utilization in the intelligent suspension system is solved, efficient control strategy optimization and robustness improvement are achieved, and autonomous driving needs are adapted to the needs of complex environments.
Patent Information
- Application Number
- CN202510775065.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-06-11
AI Technical Summary
In existing intelligent suspension systems, the sample estimation and fitting efficiency in high-dimensional space is low, resulting in low data utilization, poor stability, and difficulty in identifying high-value data, which affects the convergence speed and stability of the control strategy.
Deterministic experience tracking mechanism is adopted to identify high-value data through auxiliary rewards, and add disturbances during the training process, combined with transfer learning optimization control strategies, improve data utilization efficiency and model robustness.
It significantly improves data utilization efficiency, accelerates the learning process, improves the optimization iteration speed of control strategies and the broad adaptability of models, and ensures the optimization effect in complex environments.
Smart Images

Figure CN120295144A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of automotive dynamics control, and particularly relates to an intelligent suspension control method, medium and device combining deterministic empirical tracking. Background Art
[0002] At present, deep reinforcement learning technology has become a research hotspot in the field of autonomous driving vehicles and has shown great potential in core aspects such as visual perception, decision-making, and control. Such technology, based on a data-driven learning paradigm, brings many advantages to the control strategy of autonomous driving vehicles. Deep reinforcement learning can generate rich data resources through continuous interaction with the environment and self-optimize control behaviors accordingly, demonstrating a strong adaptability to new environments, especially suitable for solving complex and dynamic control problems. In addition, it allows developers to guide the efficient exploration and utilization of agents (i.e., controllers) by setting clear control goals, greatly reducing the time burden of manual parameter tuning in traditional methods. In the control system of autonomous driving technology, the vertical control technology of intelligent suspension is crucial for improving ride comfort, which is one of the key factors affecting the public's acceptance of autonomous driving vehicles.
[0003] When applying deep reinforcement learning control algorithms in intelligent suspension systems, the prior art faces a core problem: sample estimation and fitting in high-dimensional spaces, namely the so-called "sample crisis". This crisis is mainly reflected in the utilization efficiency and quality control of sample data. Take the invention patent application with the Chinese patent publication number CN112078318A, publication date December 15, 2020, and patent name "An Intelligent Control Method for Automotive Active Suspension Based on Deep Reinforcement Learning Algorithm", and the invention patent application with the Chinese patent publication number CN111487863A, publication date August 4, 2020, and patent name "A Reinforcement Learning Control Method for Active Suspension Based on Deep Q Neural Network" as examples. The current technology generally adopts the method of processing the data of each independent moment and single control as a single sample. This method ignores the temporal logic and correlation between data during the control process, resulting in a large amount of valuable temporal information not being fully mined and utilized, significantly reducing the overall utilization rate of data.
[0004] In addition, the capacity limitation of a single data sample not only weakens the stability of the data but also makes it more vulnerable to external environmental interference, thereby increasing the difference between data, which poses a severe challenge to the convergence speed and stability of the control strategy. In extreme cases, the control strategy may be difficult to converge to the optimal solution due to excessive data fluctuations. More seriously, when a large amount of low-value and low-information-concentration data floods the training process, the truly high-value data that is crucial for optimizing the control strategy is often submerged and difficult to be effectively identified and utilized. This situation seriously affects the effectiveness, robustness of the intelligent suspension control method, and the optimality of the final control effect. Summary of the Invention
[0005] In view of this, the present invention aims to provide an intelligent suspension control method, medium, and device combined with deterministic experience tracking to assist the reward in completing the experience tracking memory mechanism. This mechanism is specifically designed to integrate and process the information generated in a dense reward environment, effectively identify and amplify the high-value data that makes significant contributions to optimizing the control strategy, thereby accelerating the optimization and iteration process of the value and policy networks. In addition, a perturbation amount is added to the first control force according to the network training process to further promote the convergence of the training. The method provided by the present invention effectively solves the problems of low sample utilization efficiency, poor data stability, and difficulty in mining high-value data in the prior art, and further improves the overall performance and robustness of the intelligent suspension control method.
[0006] To achieve the above object, the technical solution of the present invention is realized as follows: An intelligent suspension control method combined with deterministic experience tracking, comprising: S1: Obtain multiple groups of first samples, and use the first samples to pre-train the constructed first evaluation network, second evaluation network, and control strategy network to respectively obtain an initial first evaluation model, an initial second evaluation model, and an initial control strategy model; S2: The initial control strategy model obtained in step S1 determines the second control force at the current moment output by the intelligent suspension according to the first state of the vehicle at the current moment; apply the second control force to the vehicle to obtain the next second state of the vehicle; determine the second reward and the auxiliary reward group according to the second control force and the second state; use the second state, the second control force, the second reward, the auxiliary reward group, and the next second state as the second sample; S3: Repeat step S2 multiple times to obtain multiple groups of second samples, and use the multiple groups of second samples to train the initial first evaluation model, the initial second evaluation model, and the initial control strategy model obtained in S1 again; add a perturbation amount to the second control force during the training process and change the perturbation amount according to the training effect; after training, obtain the corresponding final first evaluation model, final second evaluation model, and final control strategy model; S4: Run the intelligent suspension and control the intelligent suspension using the final first evaluation model, the final second evaluation model, and the final control strategy model trained in step S4.
[0007] Further, in step S1, each group of first samples includes: the first state of the vehicle at a certain moment, the first control force output by the intelligent suspension at the same moment, the next first state of the vehicle after the intelligent suspension applies the first control force to the vehicle, and the first reward obtained based on the first state, the first control force, and the first state.
[0008] Further, the first state , where and respectively represent the acceleration and velocity of the sprung mass of the vehicle at time t, represents the dynamic stroke of the intelligent suspension at time t, represents the vertical body displacement of the vehicle at time t, represents the unsprung displacement of the vehicle at time t; represents the velocity difference between the sprung mass and the unsprung mass at time t, represents the unsprung velocity of the vehicle at time t; The first reward is: ; where represents the first reward at time t, , ', and represent the first reward coefficients, represents the wheel dynamic load of the vehicle at time t, , represents the road surface height excitation information at time t, represents the first control force, P represents the control trigger coefficient, represents the suspension limit dynamic deflection of the intelligent suspension.
[0009] Further, the pre-training in step S1 includes: In each time step of the pre-training, extract the first sample and calculate the corresponding first target value through the following formula: ; where represents the first target value, represents the discount factor, represents the j-th evaluation network, represents the network weight of the j-th evaluation network; represents the control policy network, represents the network weight of the control policy network, Represents the next first state at time t; Calculate and train the loss functions of two evaluation networks in combination with the first target value: ; Among them, Represents the loss function of the j-th evaluation network, Represents the pre-training time step; pre-train the two evaluation networks using the loss functions of the two evaluation networks; pre-train the control policy network through the following formula: ; Among them, Represents the first state at time t.
[0010] Furthermore, in step S2, the second state ; The second reward is: ; Among them, Represents the second reward at time t, , , and Represents the second reward coefficient, Represents the second control force at time t; The auxiliary reward group is: ; Among them, Represents the auxiliary reward group at time t, Represents the auxiliary reward, k represents the acquisition step of the auxiliary reward; the auxiliary reward is: ; Among them, , and Represents the auxiliary reward coefficient.
[0011] Furthermore, during the training process of step S3: In each training time step, multiple groups of second samples , calculate the corresponding second target value through the following formula: ; Among them, Represents the th acquisition step, Represents the second target value, Represents the deterministic experience auxiliary reward discount factor, Represents the initial j-th evaluation model, Represents the model weight of the initial j-th evaluation model; denotes the initial control policy model, denotes the model weights of the initial control policy model, denotes the next second state at time t; Calculate the loss functions for training two initial evaluation models by combining with the second target value: ; wherein, denotes the loss function of the initial j-th evaluation model, N denotes the training time step; train the two initial evaluation models using the loss functions of the two initial evaluation models; Train the initial control policy model through the following formula: ; During the training process, if the second reward grows slowly, increase the perturbation amount; if the second reward grows steadily, decrease the perturbation amount.
[0012] Furthermore, in step S1, it further includes: Increase the number of network layers in the first evaluation network and the second evaluation network, as well as the number of neurons in each layer of the network; at the same time, decrease the number of network layers in the control policy network, as well as the number of neurons in each layer of the network; pre-train the modified first evaluation network, second evaluation network and control policy network.
[0013] Furthermore, between step S3 and step S4, it further includes: Use the final control policy model trained in step S3 to control the intelligent suspension of other vehicles, and re-obtain multiple groups of second samples; repeat the training process of step S3, and use the new multiple groups of second samples to train the final first evaluation model, final second evaluation model and final control policy model again.
[0014] A readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps of the intelligent suspension control method provided by the present invention in combination with deterministic experience tracing.
[0015] An electronic device includes: A memory for storing a computer program; A processor for implementing the steps of the intelligent suspension control method provided by the present invention in combination with deterministic experience tracing when executing the computer program.
[0016] Compared with the prior art, the present invention can achieve the following beneficial effects: (1) In the intelligent suspension control method combining deterministic experience tracking of the present invention, an innovative deterministic experience tracking mechanism using auxiliary rewards is proposed. As a general off-policy method, it can accurately identify and amplify valuable data, thus effectively promoting the rapid optimization and iteration of the value and policy networks. This mechanism significantly improves the data utilization efficiency and accelerates the learning process of the algorithm. In addition, the present invention uses an exploration strategy optimization mechanism based on reinforcement learning to add perturbation to the first control force during training, further promoting the convergence of training and ensuring the generality and robustness of the model. (2) In the intelligent suspension control method combining deterministic experience tracking of the present invention, by introducing the method of transfer learning, that is, using the final control strategy model to control the intelligent suspension of other vehicles and training according to the situation of other vehicles, the obtained final control strategy model has good generalization and strong adaptability. This means that in the face of the complex and changeable control requirements of autonomous vehicles, the present invention can quickly adapt and give the optimal control strategy, laying a solid foundation for the wide application of autonomous driving technology. Description of the Drawings
[0017] The drawings constituting a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments and descriptions of the present invention are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings: Figure 1 It is a schematic flowchart of the intelligent suspension control method combining deterministic experience tracking according to an embodiment of the present invention; Figure 2 It is a schematic framework diagram of the intelligent suspension control method combining deterministic experience tracking according to an embodiment of the present invention; Figure 3 It is a schematic diagram of the control policy network according to an embodiment of the present invention; Figure 4 It is a schematic diagram of the two evaluation networks according to an embodiment of the present invention; Figure 5 It is a schematic diagram of the dynamic model of the intelligent suspension system according to an embodiment of the present invention; Figure 6 It is a schematic structural diagram of the electronic device according to an embodiment of the present invention.
[0018] Description of the Reference Numerals: 1. Electronic device; 2. External device; 3. Processing unit; 4. Bus; 5. Network adapter; 6. Display; 7. (I / O) interface; 8. System memory; 9. Random access memory; 10. Cache memory; 11. Storage system; 12. Utility tool; 13. Program module. Detailed Implementation Modes
[0019] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and do not constitute a limitation to the present invention.
[0020] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments may be combined with each other.
[0021] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "center", "longitudinal", "transverse", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the accompanying drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to the present invention. In addition, the terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first", "second", etc. may explicitly or implicitly include one or more of such features. In the description of the present invention, unless otherwise specified, the meaning of "plurality" is two or more.
[0022] In the description of the present invention, it should be noted that unless otherwise clearly defined and limited, the terms "installed", "connected", "connected" should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral connection; it may be a mechanical connection or an electrical connection; it may be directly connected, or indirectly connected through an intermediate medium, and it may be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood through specific situations.
[0023] The present invention will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.
[0024] As Figures 1 to 2 shown, the intelligent suspension control method combining deterministic experience tracking described in the embodiment of the present invention includes: S1: Obtain multiple groups of first samples, and use the first samples to pre-train the constructed first evaluation network, second evaluation network and control strategy network, and correspondingly obtain an initial first evaluation model, an initial second evaluation model and an initial control strategy model.
[0025] In one embodiment, the control strategy network is asFigure 3 As shown, the state input of the vehicle is processed by a fully connected layer composed of 128 neurons. The processed data is processed by the ReLU activation function and then input into a fully connected layer composed of 128 neurons; the output data is processed by the ReLU activation function and then input into a fully connected layer composed of 1 neuron; finally, the output data is processed by the activation operation of the Tanh activation layer and the data scaling operation of the scaling layer to obtain the control force output by the intelligent suspension. The structures of the first evaluation network and the second evaluation network are as Figure 4 shown. The state input of the vehicle is processed by a fully connected layer composed of 128 neurons; the processed data is processed by the ReLU activation function and then enters a fully connected layer composed of 200 neurons together with the control information output by the control strategy network. The output features are processed by the ReLU activation function and then input into a fully connected layer composed of 1 neuron, and then the evaluation result is output. The initialization process of the three networks includes: for the first evaluation network the network weights in, the second evaluation network the network weights , and the network weights of the control strategy network are randomly initialized. s represents the state of the vehicle, and a represents the control force output by the intelligent suspension.
[0026] In some embodiments, the dynamic model of the intelligent suspension system is as Figure 5 shown, Figure 5 wherein represents the vertical displacement of the vehicle body at time t, represents the displacement below the spring of the vehicle at time t, represents the vehicle body mass, represents the mass below the spring, represents the road surface displacement of the vehicle at time t, represents the tire stiffness of the vehicle, represents the suspension spring stiffness of the intelligent suspension, represents the damping coefficient of the intelligent suspension.
[0027] Each group of first samples includes: the first state of the vehicle at time t, the first control force output by the intelligent suspension at time t, the next first state of the vehicle after the intelligent suspension acts on the vehicle with the first control force, and the first reward obtained according to the first state , the first control force . Specifically, the first state Including the acceleration of the sprung mass of the vehicle at time t and the speed , the dynamic stroke of the intelligent suspension at time t , and the speed difference between the sprung mass and the unsprung mass of the vehicle at time t , that is . Among them, the acceleration characterizes the ride comfort of the vehicle, and the acceleration and the speed are respectively obtained by taking the second-order and first-order derivatives of the vertical displacement of the vehicle body ; the dynamic stroke characterizes the dynamic deflection of the intelligent suspension and is an important index of safety; the speed difference , represents the unsprung speed of the vehicle at time t, and the unsprung speed is obtained by taking the first derivative of the unsprung displacement .
[0028] In some embodiments, the first reward is: ; Among them, represents the wheel dynamic load of the vehicle at time t, , and the wheel dynamic load characterizes the handling stability of the autonomous driving vehicle; , ', and represent the first reward coefficients. The first reward coefficients , ', and can effectively balance the multi-objective optimization problem, and their values are adaptively adjusted according to human experience or actual situations. P represents the control trigger coefficient, represents the suspension limit dynamic deflection of the intelligent suspension. The control trigger coefficient P is a fixed value used to prevent the control strategy from being unsafe during the training process, and the value of the control trigger coefficient P is adaptively selected and adjusted according to the actual situation. The above formula can be understood as follows: if the dynamic stroke of the intelligent suspension at time t exceeds the suspension limit dynamic deflection , the control trigger coefficient P is added to the first reward , and at this time, the current round of pre-training is directly forced to stop. In a certain embodiment, the control trigger coefficient P takes a value of -500, the suspension limit dynamic deflection , and the first reward coefficients , ', and Take values of 0.7, 0.1, 0.1, and 0.1 respectively.
[0029] In one embodiment, the first control force is hard-constrained by the following formula: ; where represents the truncation function, and represent the minimum and maximum values of the first control force respectively. The minimum value and the maximum value are obtained according to the limit of the maximum force that the intelligent suspension can output. In one embodiment, the minimum value and the maximum value are set.
[0030] In some embodiments, the process of pre-training the first evaluation network, the second evaluation network, and the control strategy network includes: In each time step of pre-training, a first sample is drawn, and the corresponding first target value is calculated by the following formula: ; where represents the first target value, represents the discount factor, represents the j-th evaluation network, represents the network weights of the j-th evaluation network; The loss functions of the two evaluation networks are calculated in combination with the first target value: ; where represents the loss function of the j-th evaluation network, represents the time step of pre-training; the two evaluation networks are pre-trained using the loss functions of the two evaluation networks; the control strategy network is pre-trained by the following formula: .
[0031] In one embodiment, the discount factor takes the value of 0.99, the number of episodes of pre-training is 2000, and the time step of each episode of pre-training is set to ; the two evaluation networks are trained using the stochastic gradient descent method, and the network weights and of the two evaluation networks are updated using the soft update method during the training, that is: ; where and respectively represent the network weights of the two evaluation networks after update, represents the soft update frequency, with a value of 0.001; Use the stochastic gradient ascent method to train the control policy network. During the training process, the network weights of the control policy network are also updated using the soft update method, that is: That is: ; Among them, represents the network weight of the control policy network after update. The soft update frequency here still takes the value of 0.001.
[0032] S2: The initial control policy model obtained in step S1 determines the second control force at the current moment output by the intelligent suspension according to the second state of the vehicle at the current moment; apply the second control force to the vehicle to obtain the next second state of the vehicle; determine the second reward and the auxiliary reward group according to the second control force and the second state; use the second state, the second control force, the second reward, the auxiliary reward group, and the next second state as the second sample.
[0033] In some embodiments, the second state , and the second reward is: ; Among them, represents the second reward at time t, , , and represent the second reward coefficients. The second reward coefficients , , and can also effectively balance the multi-objective optimization problem, and their values are adaptively adjusted according to human experience or actual situations. represents the second control force at time t. The above formula can be understood as that if the dynamic stroke of the intelligent suspension at time t exceeds the suspension limit dynamic deflection , add the control trigger coefficient P to the second reward
[0034] at this time, and directly force the training of the current round to stop. The auxiliary reward group is: Among them, represents the auxiliary reward group at time t, represents the auxiliary reward, and k represents the acquisition step of the auxiliary reward; the auxiliary reward is: ; Among them, , and represent the auxiliary reward coefficients, and their values are adaptively adjusted according to human experience or actual situations. In one embodiment, the control trigger coefficient P still takes the value of -500, and the suspension limit dynamic deflection , the second reward coefficient , , and take the values of 0.7, 0.1, 0.1, and 0.1 respectively, and the auxiliary reward coefficients , and take the values of 0.7, 0.1, and 0.1 respectively. The acquisition step k is set to k = [logN], where N represents the training time step in step S3, and [·] represents the rounding function. In one embodiment, an experience cache pool for storing samples is also set up. The second state, the second control force, the second reward, and the next second state are stored as intermediate samples in the experience cache pool, and the determination and calculation of the auxiliary reward group are completed in the experience cache pool. The intermediate samples and the corresponding auxiliary reward groups are jointly formed into second samples and then stored in the experience cache pool.
[0035] S3: Repeat step S2 multiple times to obtain multiple groups of second samples, and use the multiple groups of second samples to retrain the initial first evaluation model, the initial second evaluation model, and the initial control strategy model obtained in S1; during the training process, add a perturbation amount to the second control force and change the perturbation amount according to the training effect; after training, obtain the corresponding final first evaluation model, final second evaluation model, and final control strategy model. In one embodiment, the multiple groups of second samples are stored in the experience cache pool in chronological order.
[0036] In some embodiments, the training process of step S3 includes: In each training time step, use the multiple groups of second samples to calculate the corresponding second target value through the following formula: ; Among them, represents the th acquisition step, represents the second target value, represents the deterministic experience auxiliary reward discount factor, represents the initial jth evaluation model, represents the model weight of the initial jth evaluation model; represents the initial control strategy model, represents the model weight of the initial control strategy model, Represents the next first state at time t. In one embodiment, multiple groups of first samples are randomly selected from the experience cache pool , and the corresponding first target values are calculated.
[0037] In one embodiment, the second control force is hard-constrained by the following formula: ; where and represent the minimum and maximum values that limit the second control force respectively. The minimum value and the maximum value are also obtained according to the limitation of the maximum force that the intelligent suspension can output. In one embodiment, the minimum value and the maximum value are set.
[0038] Combine the first target value to calculate the loss functions for training two initial evaluation models: ; where represents the loss function of the initial j-th evaluation model; use the loss functions of the two initial evaluation models to train the two initial evaluation models; Train the initial control policy model by the following formula: ; During the training process, if the second reward grows slowly, increase the perturbation amount to make the obtained final model have a wider adaptability; if the second reward grows steadily, decrease the perturbation amount to promote the convergence of model training. The present invention improves the adaptability of the model in different environments and promotes the training of the model by introducing an exploration strategy optimization mechanism based on reinforcement learning. During the training process, add a perturbation amount to the second control force and change the perturbation amount according to the training effect. Since the second training of the present invention adopts the training method of reinforcement learning, the acquisition of the data participating in the training is trained episode by episode, and each episode can be adjusted before starting, and modified according to the situation of the previous episode before starting the next episode.
[0039] During the process of adding a perturbation amount to the second control force and changing the perturbation amount according to the training effect in the training of the present invention, by analyzing the change situation of the second reward in detail, the variance of the Gaussian noise (with a mean of 0) used as the perturbation amount is dynamically adjusted to optimize the training effect and adaptability of the model. The specific adjustment strategy is as follows: First, define a threshold parameter used to measure the change of the second reward, and a basic Gaussian noise variance .
[0040] The judgment of the second reward growth situation includes: If in two consecutive training iterations, the second reward obtained from the latter training and the second reward obtained from the previous training The difference Satisfy , it is determined that the second reward grows slowly. This means that in the current training state, the performance improvement of the model is not obvious, and it may be trapped in a local optimum or the convergence speed is too slow. At this time, in order to enable the model to explore a wider state space, the variance of the Gaussian noise is increased to improve the adaptability of the model.
[0041] If , and in several consecutive (e.g., 5) training iterations, the second reward shows a continuous upward trend, it is determined that the second reward grows steadily. This indicates that the model can effectively learn and improve performance under the current training settings. In order to promote the model to converge to the optimal solution faster, the variance of the Gaussian noise is reduced.
[0042] The adjustment strategy for the variance of the Gaussian noise includes: When it is determined that the second reward grows slowly, the variance of the Gaussian noise is increased in a linearly increasing manner, that is Represents the new variance, Is a fixed increment value used to control the amplitude of the variance increase. Gradually increasing the variance in this way can increase the fluctuation range of the Gaussian noise, guide the model to explore more different control strategies, and thus improve the generalization ability and adaptability of the model. In a certain embodiment, is set.
[0043] When it is determined that the second reward grows steadily, the variance of the Gaussian noise is reduced in an exponentially decaying manner, that is , where Is a decay coefficient less than 1, Is an exponential factor related to the number of training iterations, that is . As the training progresses, Gradually increases, causing the variance to decrease at an increasingly faster rate, thereby reducing noise interference and prompting the model to converge to the optimal solution faster.
[0044] Through the above precise judgment of the change situation of the second reward and the dynamic adjustment strategy of the variance of the Gaussian noise as the perturbation amount, it is possible to flexibly balance the exploration ability and convergence speed of the model according to the actual situation during the model training process, so that the finally obtained model not only has broad adaptability but also can quickly and accurately converge to the optimal intelligent suspension control strategy.
[0045] In one embodiment, the deterministic experience-assisted reward discount factor takes a value of 0.9, and the discount factor also takes a value of 0.99. The number of training episodes remains 2000, and the time steps of each training episode are set , and the perturbation amount is a Gaussian random number . Among them, the process of adding the perturbation amount to the second control force and changing the perturbation amount according to the training effect takes effect in the first 1000 episodes of training. The two initial evaluation models are trained using the stochastic gradient descent method, and the model weights and of the two initial evaluation models are updated in a soft update manner during the training process, that is: ; wherein, and respectively represent the model weights of the two initial evaluation models after update, takes a value of 0.001; The initial control policy model is trained using the stochastic gradient ascent method, and the model weights of the initial control policy model are also updated in a soft update manner during the training process, that is: ; wherein, represents the network weights of the initial control policy model after update. The soft update frequency here still takes a value of 0.001.
[0046] In some embodiments, in step S3, it further includes: In step S1, it further includes: Increase the number of network layers in the first evaluation network and the second evaluation network, as well as the number of neurons in each network layer; at the same time, reduce the number of network layers in the control policy network, as well as the number of neurons in each network layer; pre-train the modified first evaluation network, second evaluation network, and control policy network. In one embodiment, the number of network layers in the two evaluation networks, as well as the number of neurons in each network layer, are doubled. At this time, the number of network layers in the control policy network, as well as the number of neurons in each network layer, need to be halved. The present invention completes the compression and pruning of the control policy network through the above operations, greatly reducing the parameters and computational amount of the network model, and facilitating deployment on resource-constrained hardware platforms.
[0047] S4: Run the intelligent suspension, and use the final first evaluation model, final second evaluation model, and final control policy model trained in step S4 to control the intelligent suspension.
[0048] In some embodiments, between step S3 and step S4, the method further includes: Using the final control policy model trained in step S3 to control the intelligent suspension of other vehicles, and obtaining multiple sets of second samples again; repeating the training process of step S3, and using the new multiple sets of second samples to train the final first evaluation model, the final second evaluation model, and the final control policy model again. By using the method of transfer learning, the trained final control policy model is used to control the intelligent suspension of other vehicles and conduct training and learning, so that the finally obtained model can quickly adapt to new scenarios such as other similar vehicle models or road condition scenarios, reducing the retraining time and data requirements.
[0049] The intelligent suspension control method combining deterministic experience tracing provided by the present invention is compared with the deep deterministic policy gradient algorithm, the twin-delayed deterministic policy gradient algorithm, and the model predictive control algorithm. The method proposed by the present invention has achieved significant improvements of 74.92%, 64.20%, and 54.64% respectively in control performance. This data fully proves the superiority and efficiency of the present invention in the field of intelligent suspension control. In addition, experiments on the intelligent suspension control method combining deterministic experience tracing provided by the present invention under different speed conditions show that on randomly selected Class A, B, and C roads, the optimization effect of the present invention on ride comfort is close to 90%. Even on Class D roads with extremely poor road conditions, the optimization amount is stable at about 85%. This result not only verifies the robustness of the present invention but also demonstrates its excellent generalization performance under complex road conditions.
[0050] Based on the intelligent suspension control method combining deterministic experience tracing provided by the present invention, an intelligent suspension control system considering the time delay of the suspension system is further provided. The system includes an intelligent agent, a vehicle or a test bench equipped with an intelligent suspension and related sensors, a state observer and state estimator, and an endogenous reward function. Among them, the intelligent agent controls the intelligent suspension, and the intelligent suspension control method combining deterministic experience tracing provided by the present invention is integrated in the intelligent agent. The related sensors include an acceleration sensor, a displacement sensor, and an inertial measurement unit (IMU). The operation process of the intelligent suspension control system is as follows: The intelligent agent obtains data such as vehicle body acceleration, vehicle body speed, suspension dynamic deflection, and the derivative of suspension dynamic deflection from the vehicle system through sensors, and combines these data into the current state information of the vehicle through a state observer and a state estimator. Then, the intelligent agent decides how much control force the intelligent suspension should output in the current state according to this state information. Under the action of the control force, the vehicle state changes. The intelligent agent generates a reward value for evaluating its control action according to the new system state and the endogenous reward function. The intelligent agent performs self-iteration and control strategy optimization with reference to the reward value according to the intelligent suspension control method combining deterministic experience tracing provided by the present invention.
[0051] Figure 6 This is a schematic structural diagram of an electronic device 1 provided in an embodiment of the present invention. Figure 6 The figure shows a block diagram of an exemplary electronic device 1 suitable for implementing the embodiments of the present invention. Figure 6 The displayed electronic device 1 is merely an example and should not impose any limitation on the functions and usage scope of the embodiments of the present invention.
[0052] As Figure 6 shown, the electronic device 1 is presented in the form of a general-purpose computing device. The electronic device 1 is intended to represent various forms of digital computers, such as, for example, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present invention described and / or claimed herein.
[0053] The components of the electronic device 1 may include, but are not limited to: one or more processors or processing units 3, a system memory 8, and a bus 4 that connects different system components (including the system memory 8 and the processing unit 3).
[0054] The bus 4 represents one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any of a variety of bus structures. By way of example, these architectures include, but are not limited to, Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MAC) bus, Enhanced ISA bus, Video Electronics Standards Association (VESA) local bus, and Peripheral Component Interconnect (PCI) bus.
[0055] The electronic device 1 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by the electronic device 1, including volatile and non-volatile media, removable and non-removable media.
[0056] The system memory 8 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 9 and / or cache memory 10. The electronic device 1 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, a storage system 11 may be used for reading and writing on non-removable, non-volatile magnetic media ( Figure 6 not shown, commonly referred to as a "hard disk drive"). Although Figure 6Not shown in the figure, a disk drive for reading and writing to a removable non-volatile disk (such as a "floppy disk") and an optical disk drive for reading and writing to a removable non-volatile optical disk (such as a CD-ROM, DVD-ROM or other optical medium) can be provided. In these cases, each drive can be connected to the bus 4 through one or more data medium interfaces. The system memory 8 can include at least one program product having a set (such as at least one) of program modules configured to perform the functions of the embodiments of the present invention.
[0057] A program / utility 12 having a set (at least one) of program modules 13 can be stored, for example, in the system memory 8. Such program modules 13 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment. The program modules 13 generally perform the functions and / or methods in the embodiments described in the present invention.
[0058] The electronic device 1 can also communicate with one or more external devices 2 (such as a keyboard, a pointing device, a display 6, etc.), and can also communicate with one or more devices that enable a user to interact with the electronic device 1, and / or communicate with any device that enables the electronic device 1 to communicate with one or more other computing devices (such as a network card, a modem, etc.). Such communication can be carried out through the input / output (I / O) interface 7. Moreover, the electronic device 1 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN) and / or a public network, such as the Internet) through the network adapter 5. As Figure 6 shown, the network adapter 5 communicates with other modules of the electronic device 1 through the bus 4. It should be understood that although not shown in the figure, other hardware and / or software modules can be used in combination with the electronic device 1, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0059] The processing unit 3 executes various functional applications and data processing by running the programs stored in the system memory 8, such as implementing the intelligent suspension control method combining deterministic experience tracking provided by the embodiments of the present invention.
[0060] In the embodiments of the present invention, a non-transitory computer-readable storage medium storing computer instructions is also provided, on which a computer program is stored. When the program is executed by a processor, the intelligent suspension control method combining deterministic experience tracking provided by all the embodiments of the present application is implemented.
[0061] The computer storage medium of the embodiments of the present invention may adopt any combination of one or more computer-readable media. The computer-readable media may be computer-readable signal media or computer-readable storage media. More specific examples (a non-exhaustive list) of the computer-readable storage media include: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In this document, the computer-readable storage medium may be any tangible medium that contains or stores a program, which can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0062] The computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. The computer-readable signal medium may also be any computer-readable medium other than the computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0063] The program code contained on the computer-readable medium may be transmitted by any appropriate medium, including but not limited to wireless, wire, optical fiber, RF, etc., or any suitable combination of the foregoing. The computer program code for performing the operations of the present invention may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages - such as the "C" language or similar programming languages. The program code may be executed entirely on the user computer, partially on the user computer, executed as an independent software package, partially on the user computer and partially on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0064] The embodiments of the present invention further provide a computer program product, including a computer program, which when executed by a processor, implements the intelligent suspension control method according to the foregoing combined with deterministic experience tracking.
[0065] It should be understood that the various forms of processes shown above can be used, with steps reordered, added or deleted. For example, the steps described in the disclosure of the present invention can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solution disclosed in the present invention can be achieved, and no limitation is made herein.
[0066] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub - combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. An intelligent suspension control method combining deterministic experience tracking, characterized in that, Including: S1: Obtain multiple groups of first samples, and use the first samples to pre-train the constructed first evaluation network, second evaluation network, and control strategy network to correspondingly obtain an initial first evaluation model, an initial second evaluation model, and an initial control strategy model; S2: The initial control strategy model obtained in step S1 determines the second control force at the current moment output by the intelligent suspension according to the first state of the vehicle at the current moment; apply the second control force to the vehicle to obtain the next second state of the vehicle; Determine the second reward and the auxiliary reward group according to the second control force and the second state; use the second state, the second control force, the second reward, the auxiliary reward group, and the next second state as the second samples; S3: Repeat step S2 multiple times to obtain multiple groups of second samples, and use the multiple groups of second samples to train the initial first evaluation model, initial second evaluation model, and initial control strategy model obtained in S1 again; during the training process, add a perturbation amount to the second control force, and change the perturbation amount according to the training effect; After training, obtain the corresponding final first evaluation model, final second evaluation model, and final control strategy model; S4: Run the intelligent suspension, and use the final first evaluation model, final second evaluation model, and final control strategy model trained in step S4 to control the intelligent suspension.
2. The intelligent suspension control method combined with deterministic experience tracking according to claim 1, wherein In step S1, each group of the first samples includes: the first state of the vehicle at a certain moment, the first control force output by the intelligent suspension at the same moment, the next first state of the vehicle after the intelligent suspension applies the first control force to the vehicle, and the first reward obtained according to the first state, the first control force, and the first state.
3. The intelligent suspension control method combined with deterministic experience tracking according to claim 2, characterized in that The first state , wherein and respectively represent the acceleration and velocity of the unsprung mass of the vehicle at time t, represents the dynamic stroke of the intelligent suspension at time t, represents the vertical displacement of the vehicle body at time t, represents the unsprung displacement of the vehicle at time t; represents the velocity difference between the unsprung mass and the sprung mass at time t, represents the unsprung velocity of the vehicle at time t; The first reward is: ; Among them, represents the first reward at time t, , ', and represent the first reward coefficient, represents the wheel dynamic load of the vehicle at time t, , represents the road surface height excitation information at time t, represents the first control force, P represents the control trigger coefficient, represents the suspension limit dynamic deflection of the intelligent suspension.
4. The intelligent suspension control method combined with deterministic experience tracking according to claim 3, characterized in that The pre-training in step S1 includes: In each time step of pre-training, extract the first sample and calculate the corresponding first target value through the following formula: ; Among them, represents the first target value, represents the discount factor, represents the j-th evaluation network, represents the network weight of the j-th evaluation network; represents the control policy network, represents the network weight of the control policy network, represents the next first state at time t; Calculate the loss functions for training the two evaluation networks in combination with the first target value: ; Among them, represents the loss function of the j-th evaluation network, represents the time step of pre-training; pre-train the two evaluation networks using the loss functions of the two evaluation networks; pre-train the control policy network through the following formula: ; Among them, represents the first state at time t.
5. The intelligent suspension control method combining deterministic experience tracking according to claim 4, characterized in that, In step S2, The second state ; The second reward is: ; Among them, represents the second reward at time t, , , and represent the second reward coefficient, represents the second control force at time t; The auxiliary reward group is: ; Among them, represents the auxiliary reward group at time t, represents the auxiliary reward, k represents the acquisition step of the auxiliary reward; the auxiliary reward is: ; Among them, , and represent the auxiliary reward coefficients.
6. The intelligent suspension control method combining deterministic experience tracking according to claim 5, characterized in that During the training process of step S3: In each time step of training, multiple groups of second samples , calculate the corresponding second target value through the following formula: ; Among them, represents the th acquisition step, represents the second target value, represents the deterministic experience-assisted reward discount factor, represents the initial jth evaluation model, represents the model weight of the initial jth evaluation model; represents the initial control policy model, represents the model weight of the initial control policy model, represents the next second state at time t; Calculate the loss functions for training the two initial evaluation models in combination with the second target value: ; wherein, represents the loss function of the initial j-th evaluation model, and N represents the training time step; the two initial evaluation models are trained using the loss functions of the two initial evaluation models; Train the initial control strategy model through the following formula: ; During the training process, if the second reward grows slowly, increase the perturbation amount; if the second reward grows steadily, decrease the perturbation amount.
7. The intelligent suspension control method combined with deterministic experience tracking according to claim 1, wherein In step S1, it also includes: Increase the number of network layers in the first evaluation network and the second evaluation network, and the number of neurons in each layer of the network; at the same time, decrease the number of network layers in the control strategy network, and the number of neurons in each layer of the network; pre-train the modified first evaluation network, second evaluation network, and control strategy network.
8. The intelligent suspension control method combined with deterministic experience tracking according to claim 1, characterized in that, Between step S3 and step S4, it also includes: Use the finally trained control strategy model in step S3 to control the intelligent suspension of other vehicles, and obtain multiple groups of second samples again; repeat the training process in step S3, and use the new multiple groups of second samples to train the finally first evaluation model, the finally second evaluation model, and the finally control strategy model again.
9. A readable storage medium, characterized in that, A computer program is stored on the readable storage medium, and when the computer program is executed by a processor, the steps of the intelligent suspension control method combining deterministic experience tracking as described in any one of claims 1 to 8 are implemented.
10. An electronic device, characterized in that, Comprising: A memory for storing a computer program; A processor for implementing the steps of the intelligent suspension control method combining deterministic experience tracking as described in any one of claims 1 to 8 when executing the computer program.
Citation Information
Patent Citations
Active suspension reinforcement learning control method based on deep Q neural network
CN111487863A
Automobile active suspension intelligent control method based on deep reinforcement learning algorithm
CN112078318A
Multi-axis collaborative robot dynamic path planning method based on deep reinforcement learning
CN118596133A
Depth deterministic strategy gradient driven multi-constraint guidance law design method
CN119644742A
System and method for nonlinear dynamic control based on soft computing with discrete constraints
CN1672103A
Cited By
Demonstration assistance-based reinforcement learning suspension control method and system and storage medium
CN120517116A