Suspension system control method for adjusting reward in combination with RND, medium and electronic equipment

By introducing an RND reward adjustment structure with physical value correlation factors into the suspension system control, combined with the TD3 model, the problem of over-exploration of the agent is solved, and a more robust and safe control strategy of the suspension system is realized.

CN120156237AActive Publication Date: 2025-06-17JILIN UNIVERSITY
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510619000.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-06-17
Estimated Expiration
2045-05-14

AI Technical Summary

Technical Problem

In suspension system control, the existing technology fails to effectively combine the random network distillation (RND) algorithm in deep reinforcement learning with physical constraints, resulting in the agents that may over-exploration and perform actions that do not meet physical or safety requirements, affect the performance of the suspension system and cause safety hazards.

Method used

By introducing an RND reward adjustment structure with physical value correlation factors, the TD3 model can adjust the reward according to the degree of compliance between the physical value correlation factors and physical constraints, so as to maintain sensitivity to physical reality while pursuing intrinsic rewards.

Benefits of technology

It realizes effective balance of exploration and physical constraints in suspension system control, improves the stability and safety of the suspension system, enhances passenger comfort, and improves the vehicle's response ability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120156237A_ABST
    Figure CN120156237A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of automobile control, in particular to a suspension system control method for adjusting rewards in combination with RND, a medium and electronic equipment, and the method comprises the steps: determining a plurality of optimization targets for controlling a suspension system; constructing a TD3 architecture with an RND reward adjustment structure, and training the TD3 architecture according to the plurality of optimization targets; in the RND reward adjustment structure, the difference between the output of a Critic network and the output of a CriticT network in the TD3 architecture is calculated, and a physical value correlation factor is generated according to a plurality of optimization targets; combining the obtained difference and physical value correlation factor with a reward corresponding to the current environment state to obtain a current reward, and training the TD3 architecture by using the current reward to obtain a TD3 model; and finally, inputting the current environment state conforming to the optimization target into the TD3 model, and predicting the output action of the suspension system. The RND reward adjusting structure provided by the invention is combined with external rewards, so that the suspension system can better adapt to different tasks and environments, and the overall performance of reinforcement learning is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of automotive control, and particularly relates to a suspension system control method, medium and electronic device that combines RND to adjust rewards. Background Art

[0002] In 2019, Random Network Distillation (RND) proposed by Schaul et al. is an exploration mechanism based on intrinsic rewards, aiming to solve the exploration problem when external rewards are sparse in reinforcement learning. The core idea of RND is to generate intrinsic rewards by using a randomly initialized fixed target network and a trainable prediction network.

[0003] In the application of deep reinforcement learning (DRL), the RND algorithm has received wide attention for its unique intrinsic reward mechanism. This mechanism is based on prediction errors and the novelty of states, aiming to guide the agent to explore more unknown states. However, the actual application of the suspension system control system needs to strictly follow a series of physical constraints, such as safety, stability and ride comfort. These constraints are crucial for ensuring the driving experience and vehicle performance.

[0004] If these physical limitations are not explicitly modeled in the suspension system control, RND may cause the agent to over-explore. Specifically, the agent may try to perform some actions that do not meet physical or safety requirements. This situation may not only affect the performance of the suspension system, but also pose serious safety hazards. For example, during the exploration process, the agent may choose overly aggressive control strategies, causing the suspension system to face severe vibrations and instability, resulting in passenger discomfort and potential vehicle damage. Summary of the Invention

[0005] In view of this, the present invention aims to provide a suspension system control method, medium and electronic device that combines RND to adjust rewards. By introducing an RND reward adjustment structure, the obtained TD3 model can adjust the rewards according to the degree of compliance between the physical value correlation factor in the RND reward adjustment structure and the physical constraints. Furthermore, while the suspension system pursues intrinsic rewards, it can maintain sensitivity to physical reality.

[0006] To achieve the above object, the technical solution of the present invention is realized as follows: A suspension system control method that combines RND to adjust rewards, comprising: S1: Determine multiple optimization goals for controlling the suspension system; S2: Construct a TD3 architecture with an RND reward adjustment structure, and train the TD3 architecture according to the multiple optimization goals in step S1; In the RND reward adjustment structure, calculate the difference between the outputs of the Critic network and the CriticT network in the TD3 architecture, and generate a physical value correlation factor according to multiple optimization objectives; combine the obtained difference and the physical value correlation factor with the reward corresponding to the current environmental state to obtain the current reward, and use the current reward to train the TD3 architecture to obtain the TD3 model; S3: Input the current environmental state that meets the optimization objectives in step S1 into the TD3 model obtained in step S2 to predict the output action of the suspension system.

[0007] Furthermore, the RND reward adjustment structure includes a maximization module, a difference output module, a physical value correlation factor module, and a reward output module; among them, The maximization module receives the evaluation value output by the Critic network in the TD3 architecture and outputs the maximum value in the evaluation value; The difference output module receives the maximum value and the minimum value in the target evaluation value output by the CriticT network in the TD3 architecture, and calculates the difference through the following formula: ; Where, represents the difference, represents the maximum value, represents the minimum value; The physical value correlation factor module combines multiple optimization objectives and the negative gradient function to generate a physical value correlation factor; The reward output module multiplies the difference and the physical value correlation factor by the reward corresponding to the current environmental state to obtain the current reward.

[0008] Furthermore, the TD3 architecture in step S2 also includes an Actor network, an ActorT network, a Critic1 network, a Critic2 network, a CriticT1 network, and a CriticT2 network; among them, The Actor network receives the current environmental state and applies the generated execution action to the environment to generate the next environmental state; the ActorT network receives the next environmental state and generates a target action; the CriticT1 network and the CriticT2 network simultaneously receive the target action and the next environmental state, and input the minimum target evaluation value in the output target evaluation values into the RND reward adjustment structure; the Critic1 network and the Critic2 network simultaneously receive the current environmental state and the next environmental state; the Critic1 network and the Critic2 network input the evaluation value corresponding to the next environmental state into the RND reward adjustment structure; the Critic1 network and the Critic2 network generate the evaluation values participating in the training according to the current environmental state.

[0009] Furthermore, the training process in step S2 includes: S21: Initialize the TD3 architecture; S22: In the current environmental state, the Actor network generates an action and controls the suspension system to execute the action, obtaining the next environmental state and the corresponding reward generated by the action in the optimization objective; the ActorT network receives the next environmental state and generates a target action; S23: The Critic1 network and the Critic2 network simultaneously receive the current environmental state and the next environmental state; the evaluation values generated by the Critic1 network and the Critic2 network based on the next environmental state are input into the RND reward adjustment structure; S24: The CriticT1 network and the CriticT2 network simultaneously receive the target action and the next environmental state, and input the minimum target evaluation value among the generated target evaluation values into the RND reward adjustment structure; S25: The RND reward adjustment structure receives the evaluation value in step S23 and the minimum target evaluation value in step S24, and generates the current reward in combination with the reward obtained according to step S22; S26: Repeat steps S22 - S25 multiple times, and integrate the obtained actions, next environmental states, current environmental states, and current rewards into experience data and store them in the experience pool; S27: Randomly sample a batch of experience data from the experience pool, and calculate the target value using the Bellman expectation formula; calculate the loss functions of the Critic1 network and the Critic2 network in combination with the target value and the target action; S28: According to the loss functions obtained in step S27, perform parameter updates on the TD3 architecture using a delayed update strategy, and repeat steps S21 - S28 with the updated TD3 architecture until the training of the TD3 architecture is completed.

[0010] Further, in step S27, the target value is obtained through the following formula: ; where, denotes the target value, denotes the discount factor, denotes the CriticT1 network, denotes the CriticT2 network, denotes the network parameters of the CriticT1 network, denotes the network parameters of the CriticT2 network, denotes the target action, denotes the next environmental state, and r denotes the current reward.

[0011] Further, in step S27, the loss functions for training the Critic1 network and the Critic2 network are as follows: ; ; Among them, represents the loss function of the Critic1 network, represents the loss function of the Critic2 network, represents the result obtained by inputting the action and the current environmental state into the Critic1 network, represents the network parameters of the Critic1 network, represents the result obtained by inputting the action and the current environmental state into the said Critic2 network, represents the network parameters of the Critic2 network, MSE represents the mean squared error loss function, represents the target value of the Critic1 network, represents the target value of the Critic2 network.

[0012] A readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps of the suspension system control method for adjusting rewards in combination with RND provided by the present invention.

[0013] An electronic device includes: A memory for storing a computer program; A processor for implementing the steps of the suspension system control method for adjusting rewards in combination with RND provided by the present invention when executing the computer program.

[0014] Compared with the prior art, the present invention can achieve the following beneficial effects: In the suspension system control method for adjusting rewards in combination with RND according to the present invention, by introducing an RND reward adjustment structure with a physical value correlation factor, the relationship between exploration and physical constraints is effectively balanced, thereby prompting the suspension control system to implement a more robust and safe control strategy. This can not only improve the stability of the suspension system under various driving conditions but also enhance the comfort of passengers. In practical applications, this method may significantly improve the response ability of the vehicle, enabling the intelligent agent to make more accurate and safe decisions in complex environments.

[0015] Integrating the physical value correlation factor into the RND algorithm not only improves the safety and stability of suspension control but also provides a new idea for the application of deep reinforcement learning in complex engineering control tasks. This combination provides a more reasonable exploration guidance for the agent, ensuring that it always follows the physical laws in a changing environment, thereby enhancing the performance of the overall system. Brief Description of the Drawings

[0016] The drawings forming a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments and descriptions thereof of the present invention are used to explain the present invention and do not constitute an improper limitation on the present invention. In the drawings: Figure 1 It is a schematic flowchart of the suspension system control method combining RND to adjust the reward according to the embodiments of the present invention; Figure 2 It is a schematic structural diagram of the TD3 architecture according to the embodiments of the present invention; Figure 3 It is a schematic structural diagram of the RND reward adjustment structure according to the embodiments of the present invention; Figure 4 It is a schematic structural diagram of the electronic device according to the embodiments of the present invention.

[0017] Description of the Reference Numerals: 1. Electronic device; 2. External device; 3. Processing unit; 4. Bus; 5. Network adapter; 6. Display; 7. (I / O) interface; 8. System memory; 9. Random access memory; 10. Cache memory; 11. Storage system; 12. Utility tool; 13. Program module. Detailed Description of the Embodiments

[0018] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and do not constitute a limitation on the present invention.

[0019] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.

[0020] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by terms such as "center", "longitudinal", "transverse", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to the present invention. In addition, terms such as "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first", "second", etc. may explicitly or implicitly include one or more of such features. In the description of the present invention, unless otherwise stated, the meaning of "a plurality" is two or more.

[0021] In the description of the present invention, it should be noted that unless otherwise clearly specified and limited, the terms "installation", "connection", and "coupling" should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral connection; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood through specific situations.

[0022] The present invention will be described in detail below with reference to the drawings and in conjunction with embodiments.

[0023] As Figures 1 to 3 shown, the suspension system control method combining RND adjustment rewards described in the embodiments of the present invention includes: S1: Determine multiple optimization goals for controlling the suspension system.

[0024] In a certain embodiment, the optimization goals for controlling the suspension system are respectively the displacement of the sprung mass, the velocity of the sprung mass, and the acceleration of the sprung mass; the displacement of the unsprung mass, the velocity of the unsprung mass, and the acceleration of the unsprung mass.

[0025] S2: Construct a TD3 architecture with an RND reward adjustment structure and train the TD3 architecture according to the multiple optimization goals in step S1.

[0026] Specifically, the TD3 architecture includes an Actor network, an ActorT network, a Critic1 network, a Critic2 network, a CriticT1 network, a CriticT2 network, and an RND reward adjustment structure. The Actor network receives the current environmental state and applies the generated execution action to the environment to generate the next environmental state. The ActorT network receives the next environmental state and generates a target action. The CriticT1 network and the CriticT2 network simultaneously receive the target action and the next environmental state, and input the minimum target evaluation value among the output target evaluation values into the RND reward adjustment structure. The Critic1 network and the Critic2 network simultaneously receive the current environmental state and the next environmental state; the Critic1 network and the Critic2 network input the evaluation value corresponding to the next environmental state into the RND reward adjustment structure. In the RND reward adjustment structure, the difference between the outputs of the Critic network and the CriticT network in the TD3 architecture is calculated, and a physical value correlation factor is generated according to multiple optimization objectives; the obtained difference, the physical value correlation factor, and the reward corresponding to the current environmental state are combined to obtain the current reward. The Critic1 network and the Critic2 network generate evaluation values for participating in training according to the current environmental state. The TD3 architecture is trained with the current reward to obtain the TD3 model.

[0027] In a specific embodiment, the Actor network, the Actor-T network, the two CriticT networks, and the two Critic networks generally adopt a structure of 3 to 4 layers, where the structures of the Actor network and the Actor-T network are the same, and the structures of the two CriticT networks and the two Critic networks correspond to each other, that is, the structure of the CriticT1 network is the same as that of the CriticT1 network, and the structure of the CriticT2 network is the same as that of the CriticT2 network. The input layer of the Actor network can use a node to receive the state, pass through 2 to 3 hidden layers (for example, the number of nodes is 300 and 200, and the activation function uses ReLU), and the output layer generates an action. Since the generated action needs to be restricted within the range of the action space (for example, [−1,1]), the tanh activation function is usually used. If the range of the action space is larger, the output value can be extended to the target range through a linear transformation. The input layers of the two Critic networks both receive the concatenated result of the current environmental state and the action, and there should be two nodes in the input layer. The hidden layer settings are similar to those of the Actor network (also using the ReLU activation function), and the output layer is a linear output for calculating the evaluation value of the generated action. The TD3 model is updated softly to balance performance and computational cost while avoiding the problems of gradient disappearance and overfitting.

[0028] In some embodiments, the RND reward adjustment structure is as follows Figure 3 shown, including a maximization module, a difference output module, a physical value correlation factor module, and a reward output module. Among them, the maximization module receives the evaluation values corresponding to the next environmental state of the Critic1 network and the Critic2 network, and outputs the maximum value among the evaluation values. The difference output module receives the maximum value, as well as the minimum value among the target evaluation values output by the CriticT1 network and the CriticT2 network, and calculates the mean square error of the minimum value and the maximum value to obtain the difference between the outputs of the Critic network and the CriticT network. The physical value correlation factor module generates a physical value correlation factor by combining multiple optimization objectives and a negative gradient function. The reward output module multiplies the difference and the physical value correlation factor by the reward corresponding to the current environmental state to obtain the current reward.

[0029] It should be noted that the maximization module is a connection structure between the TD3 and RND algorithms, which is one of the innovation points of the present invention. Its function is to receive the outputs of the Critic1 and Critic2 networks, and output the maximum value of the two to calculate the difference in the RND algorithm, and at the same time complete the connection between the TD3 algorithm model and the RND algorithm model.

[0030] It can be understood that in the difference output module, the difference is calculated by the following formula: ; where represents the difference, represents the maximum value, represents the minimum value.

[0031] In the reward output module, the current reward is obtained by the following formula: ; where represents the current reward, represents the physical value correlation factor, represents the reward corresponding to the current environmental state. The reward output module can achieve that when the suspension system is performing an action, it will not only consider the rewards of the external environment (such as reaching the goal, obtaining items, etc.), but also internally adjust the rewards according to the exploration of the unknown state, so as to balance the role of exploration and exploitation.

[0032] It should be noted that in the traditional RND algorithm, after calculating the output difference L, it can be put into training. However, in practical engineering problems, especially in the vertical problem of suspension control, we must control the Actor network not to explore some action spaces that do not conform to physical reality, so as to reduce the waste of computing resources and improve the exploration efficiency. Therefore, to achieve this goal, the physical value correlation factor module designed in the present invention changes the ordinary hyperparameters in the traditional RND algorithm into physical value correlation factors, and changes its form into a negative gradient function for state variables , that is: ; Among them, represents the optimization objective, and n represents the number of optimization objectives. It can be understood that the negative gradient function satisfies the following conditions: ; Among them, represents taking the partial derivative of each term (T1, T2,..., T in the negative gradient function n ).

[0033] Because in the vertical control of the suspension, the state variables of each component of the suspension will not have very large values (too large state variable values often mean situations such as hitting the limit block).

[0034] The physical value correlation factor is defined by the generalized function definition method. Any negative gradient function that satisfies the above conditions can be used as the expression form of the actual physical value correlation factor η. The present invention does not limit the specific form of the negative gradient function. This condition means that when the agent installed on the suspension system explores an action space that does not conform to physical reality, it will not bring great benefits. The larger the state variable, the more unrealistic it is, and the smaller the physical value correlation factor η, the smaller the current reward, that is, the exploration intensity is restricted.

[0035] In a specific embodiment, the selected negative gradient function is an exponential function. The agent used consists of a processor and a motor, and is located between the sprung mass and the unsprung mass. It can output a vertical force to the environment / system according to the current state using the Actor network in the processor.

[0036] The RND reward adjustment structure provided by the present invention provides an intrinsic reward for the agent by combining the difference between the outputs of the Critic network and the CriticT network, encouraging the exploration of unencountered states. The RND reward adjustment structure effectively alleviates the sparse reward problem, enabling the agent to explore states more proactively in complex environments, thereby improving the learning efficiency. The RND reward adjustment structure provided by the present invention, in combination with the extrinsic reward (i.e., the corresponding reward generated by the action in the optimization objective), better adapts to different tasks and environments, improving the overall performance of reinforcement learning.

[0037] It can be understood that in the TD3 architecture provided by the present invention, the Actor network receives the current environmental state , and the generated execution action acts on the environment to generate the next environmental state . The ActorT network receives the next environmental state , and generates the target action . The CriticT1 network and the CriticT2 network simultaneously receive the target action and the next environmental state , and input the minimum target evaluation value and in the output target evaluation values into the RND reward adjustment structure. represents the network parameters of the CriticT1 network, represents the network parameters of the CriticT2 network. The Critic1 network and the Critic2 network simultaneously receive the current environmental state and the next environmental state . The Critic1 network and the Critic2 network input the evaluation values and corresponding to the next environmental state into the RND reward adjustment structure. The Critic1 network and the Critic2 network generate the evaluation values and participating in the training according to the current environmental state . represents the network parameters of the Critic1 network, represents the network parameters of the Critic2 network.

[0038] In some embodiments, the training process in step S2 includes: S21: Initialize the TD3 architecture. In some embodiments, the following parameters are randomly initialized using the Xavier initialization method. The initialization includes: Initialize the environmental state; the state of the suspension is the state corresponding to the optimization objective, i.e., including the displacement, velocity, and acceleration of the sprung mass, and the displacement, velocity, and acceleration of the unsprung mass; Randomly initialize the Critic1 network and the Critic2 network. Specifically, for the network parameters of the Critic1 network and the network parameters of the Critic2 network perform random initialization; Initialize the two target Critic networks, i.e., the CriticT1 network and the CriticT2 network. Specifically, for the network parameters of the CriticT1 network and the network parameters of the CriticT2 network perform initialization. The initialized network parameters are equal to the network parameters , and the initialized network parameters are equal to the network parameters ; Randomly initialize the Actor network. Specifically, for the network parameters of the Actor network perform random initialization; Initialize the target Actor network, i.e., the ActorT network. Specifically, for the network parameters of the ActorT network perform initialization. The initialized network parameters are equal to the network parameters ; Initialize the hyperparameters for updating the TD3 architecture. The hyperparameters include the discount factor , the soft update frequency , the policy update frequency policy_delay, and the learning rates of the Critic1 network and the Critic2 network. In one embodiment, the learning rate of the Actor network is usually set to 10 -4 , and the learning rates of the two Critic networks are both 10 -3 , and the soft update frequency is set in [0.005, 0.01].

[0039] S22: In the current environmental state, the Actor network generates an action and controls the suspension system to execute the action, obtaining the next environmental state and the corresponding reward generated by the action in the optimization objective. The ActorT network receives the next environmental state and generates a target action.

[0040] In some embodiments, in the current environmental state s, the Actor network generates an action added with exploration noise , i.e.: That is: ; Among them, represents the Actor network, and the exploration noise is Gaussian noise.

[0041] Control the suspension system to execute actions to obtain the next environmental state , and obtain the action through the following formula For the corresponding rewards generated by each optimization objective , that is: ; Among them, j represents the number of optimization objectives, represents the next environmental state corresponding to the jth optimization objective, represents the reward corresponding to the jth optimization objective.

[0042] The ActorT network receives the next environmental state , and generates the target action , that is: ; Among them, represents the ActorT network, represents the noise, and Gaussian noise is adopted in a certain embodiment.

[0043] S23: The Critic1 network and the Critic2 network simultaneously receive the current environmental state and the next environmental state; the minimum evaluation value among the evaluation values generated by the Critic1 network and the Critic2 network according to the next environmental state is input into the RND reward adjustment structure.

[0044] It can be understood that the Critic1 network and the Critic2 network obtain the corresponding evaluation values and from the evaluation values generated according to the next environmental state, and the evaluation value are input into the RND reward adjustment structure.

[0045] S24: The CriticT1 network and the CriticT2 network simultaneously receive the target action and the next environmental state, and input the minimum target evaluation value among the generated target evaluation values into the RND reward adjustment structure.

[0046] It can be understood that the CriticT1 network and the CriticT2 network simultaneously receive the target action and the next environmental state , and respectively generate the target evaluation value and the target evaluation value The target evaluation value is obtained by the following formula and the target evaluation value among them, the minimum target evaluation value : ; Input the minimum target evaluation value into the RND reward adjustment structure.

[0047] S25: The RND reward adjustment structure receives the evaluation value in step S23 and the minimum target evaluation value in step S24, and combines the reward obtained according to step S22 to generate the current reward.

[0048] Combined with the above description of the RND reward adjustment structure, it can be obtained that the evaluation value and the evaluation value are input into the maximization module in the RND reward adjustment structure, and the maximum value is obtained by the following formula : .

[0049] The minimum target evaluation value is input into the difference output module in the RND reward adjustment structure. The difference output module combines the maximum value and the minimum target evaluation value to output the difference L. The reward output module combines the difference L and the physical value correlation factor generated by the physical value correlation factor module, and obtains the current reward r according to the reward corresponding to the current environmental state.

[0050] It can be understood that the current reward r is obtained by the following formula ; Among them, represents the physical value correlation factor corresponding to the jth optimization target, and n represents the total number of optimization targets.

[0051] S26: Repeat steps S22~S25 multiple times, and integrate the obtained actions, next environmental state, current environmental state, and current reward into experience data and store them in the experience pool.

[0052] It can be understood that by repeating steps S22~S25 multiple times, the obtained actions , next environmental state , current environmental state and the current reward r are integrated into experience data and stored in the experience pool.

[0053] S27: Randomly sample a batch of experience data from the experience pool, and calculate the target value using the Bellman expectation formula; calculate the loss functions of the Critic1 network and the Critic2 network by combining the target value and the target action.

[0054] It can be understood that a batch of experience data is randomly sampled from the experience pool , and based on the experience data calculate the target value using the Bellman expectation formula , that is: .

[0055] The loss functions for training the Critic1 network and the Critic2 network are: ; ; where represents the loss function of the Critic1 network, represents the loss function of the Critic2 network, MSE represents the mean squared error loss function, represents the target value of the Critic1 network , represents the target value of the Critic2 network .

[0056] S28: According to the loss functions obtained in step S27, update the parameters of the TD3 architecture using a delayed update strategy, and repeat steps S21 - S28 with the updated TD3 architecture until the training of the TD3 architecture is completed. Among them, the delayed update strategy specifically includes: During each round of training, update the parameter and the parameter through the following formula: ; ; where represents taking the gradient of the loss function with respect to in it, represents taking the gradient of the loss function with respect to in it, represents the gradient coefficient.

[0057] According to the parameter and the parameter , update the parameter and the parameter respectively: ; ; Every policy_delay rounds of training, for the updated parameters and parameters , including: Calculate the policy loss through the following formula : ; According to the policy loss , update the parameter through the following formula: ; Wherein, represents taking the gradient of in the policy loss ; Then update the parameter through the following formula: ; Update the current environmental state to the next environmental state .

[0058] S3: Input the current environmental state that meets the optimization objective of step S1 into the TD3 model obtained in step S2, and predict the output action of the suspension system.

[0059] Correspondingly, according to the embodiments of the present invention, the present invention also provides an electronic device, a readable storage medium, and a computer program product.

[0060] Figure 4 It is a schematic structural diagram of an electronic device 1 provided in the embodiments of the present invention. Figure 4 Shows a block diagram of an exemplary electronic device 1 suitable for implementing the embodiments of the present invention. Figure 4 The shown electronic device 1 is only an example and should not bring any limitation to the functions and usage scope of the embodiments of the present invention.

[0061] As Figure 4 shown, the electronic device 1 is presented in the form of a general-purpose computing device. The electronic device 1 is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are only examples and are not intended to limit the implementation of the present invention described herein and / or claimed.

[0062] The components of the electronic device 1 may include, but are not limited to: one or more processors or processing units 3, a system memory 8, and a bus 4 that couples different system components (including the system memory 8 and the processing unit 3).

[0063] The bus 4 represents one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processor, or a local bus using any of a variety of bus structures. By way of example, these architectures include, but are not limited to, Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MAC) bus, Enhanced ISA bus, Video Electronics Standards Association (VESA) local bus, and Peripheral Component Interconnect (PCI) bus.

[0064] The electronic device 1 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by the electronic device 1, including volatile and nonvolatile media, removable and non-removable media.

[0065] The system memory 8 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 9 and / or cache memory 10. The electronic device 1 may further include other removable / non-removable, volatile / nonvolatile computer system storage media. By way of example only, a storage system 11 can be used for reading and writing on non-removable, nonvolatile magnetic media ( Figure 4 not shown and typically called a "hard disk drive"). Although Figure 4 not shown in the figure, a disk drive for reading and writing on a removable nonvolatile disk (such as a "floppy disk"), and an optical disk drive for reading and writing on a removable nonvolatile optical disk (such as a CD-ROM, DVD-ROM, or other optical media) can be provided. In these cases, each drive can be coupled to the bus 4 via one or more data media interfaces. The system memory 8 may include at least one program product having a set (e.g., at least one) of program modules that are configured to perform the functions of the embodiments of the present invention.

[0066] A program / utility 12 having a set (at least one) of program modules 13 can be stored, for example, in the system memory 8. Such program modules 13 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data, and an implementation of a network environment may be included in each or some combination of these examples. The program modules 13 typically perform the functions and / or methods described in the embodiments of the present invention.

[0067] The electronic device 1 can also communicate with one or more external devices 2 (such as a keyboard, a pointing device, a display 6, etc.), and can also communicate with one or more devices that enable a user to interact with the electronic device 1, and / or communicate with any device that enables the electronic device 1 to communicate with one or more other computing devices (such as a network card, a modem, etc.). Such communication can be carried out through an input / output (I / O) interface 7. Moreover, the electronic device 1 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter 5. As Figure 4 shown, the network adapter 5 communicates with other modules of the electronic device 1 through a bus 4. It should be understood that, although not shown in the figure, other hardware and / or software modules can be used in combination with the electronic device 1, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.

[0068] The processing unit 3 executes various functional applications and data processing by running programs stored in the system memory 8, for example, implementing the suspension system control method combining RND adjustment rewards provided by the embodiments of the present invention.

[0069] In the embodiments of the present invention, a non-transitory computer-readable storage medium storing computer instructions is also provided, on which a computer program is stored. Among them, when the program is executed by a processor, the suspension system control method combining RND adjustment rewards provided by all the embodiments of the present application is implemented.

[0070] The computer storage medium of the embodiments of the present invention can adopt any combination of one or more computer-readable media. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. More specific examples (non-exhaustive list) of the computer-readable storage medium include: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, the computer-readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, apparatus, or device.

[0071] A computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device.

[0072] The program code contained on a computer-readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing. The computer program code for performing the operations of the present invention may be written in one or more programming languages or combinations thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and also including conventional procedural programming languages - such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or alternatively, may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0073] An embodiment of the present invention also provides a computer program product, including a computer program, which when executed by a processor, implements the suspension system control method for adjusting rewards in combination with RND as described above.

[0074] It should be understood that various forms of the flow shown above may be used, reordering, adding, or deleting steps. For example, the steps recited in the disclosure of the present invention may be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution disclosed in the present invention can be achieved, and no limitation is imposed herein.

[0075] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A suspension system control method combined with RND adjustment reward, characterized in that: include: S1: Determine multiple optimization objectives for controlling the suspension system; S2: constructing a TD3 architecture with an RND reward adjustment structure, and training the TD3 architecture according to the multiple optimization objectives of step S1; In the RND reward adjustment structure, the difference between the outputs of the Critic network and the CriticT network in the TD3 architecture is calculated, and a physical value association factor is generated according to multiple optimization objectives; the obtained difference is combined with the reward corresponding to the physical value association factor and the current environmental state to obtain a current reward, and the TD3 architecture is trained with the current reward to obtain a TD3 model; S3: Inputting the current environmental state that meets the optimization target of step S1 into the TD3 model obtained in step S2 to predict the output action of the suspension system.

2. The suspension system control method combined with RND adjustment reward according to claim 1, characterized in that: The RND reward adjustment structure includes a maximization module, a difference output module, a physical value correlation factor module and a reward output module; wherein, The maximization module receives the evaluation values ​​output by the Critic network in the TD3 architecture, and outputs the maximum value among the evaluation values; The difference output module receives the maximum value and the minimum value of the target evaluation value output by the CriticT network in the TD3 architecture, and calculates the difference by the following formula: ; in, Indicates the difference, represents the maximum value, represents said minimum value; The physical value correlation factor module generates the physical value correlation factor by combining multiple optimization objectives and a negative gradient function; The reward output module multiplies the difference and the physical value association factor by the reward corresponding to the current environmental state to obtain the current reward.

3. The suspension system control method combined with RND adjustment reward according to claim 2, characterized in that: The TD3 architecture in step S2 also includes an Actor network, an ActorT network, a Critic1 network, a Critic2 network, a CriticT1 network and a CriticT2 network; wherein, The Actor network receives the current environment state and applies the generated execution action to the environment to generate the next environment state; The ActorT network receives the next environment state and generates a target action; The CriticT1 network and the CriticT2 network simultaneously receive the target action and the next environment state, and input the minimum target evaluation value among the output target evaluation values ​​into the RND reward adjustment structure; The Critic1 network and the Critic2 network simultaneously receive the current environment state and the next environment state; the Critic1 network and the Critic2 network input the evaluation value corresponding to the next environment state into the RND reward adjustment structure; the Critic1 network and the Critic2 network generate evaluation values ​​participating in training according to the current environment state.

4. The suspension system control method combined with RND adjustment reward according to claim 3, characterized in that: The training process in step S2 includes: S21: Initialize the TD3 architecture; S22: Under the current environment state, the Actor network generates an action and controls the suspension system to perform the action, thereby obtaining the next environment state and the corresponding reward generated by the action at the optimization target; the ActorT network receives the next environment state and generates the target action; S23: The Critic1 network and the Critic2 network simultaneously receive the current environment state and the next environment state; the evaluation values ​​generated by the Critic1 network and the Critic2 network according to the next environment state are input into the RND reward adjustment structure; S24: The CriticT1 network and the CriticT2 network simultaneously receive the target action and the next environment state, and input the minimum target evaluation value among the generated target evaluation values ​​into the RND reward adjustment structure; S25: The RND reward adjustment structure receives the evaluation value of step S23 and the minimum target evaluation value of step S24, and generates a current reward in combination with the reward obtained according to step S22; S26: Repeat steps S22 to S25 multiple times, integrate the obtained action, the next environment state, the current environment state and the current reward into experience data and store them in the experience pool; S27: randomly sampling a batch of experience data from the experience pool, and calculating a target value using the Bellman expectation formula; and calculating a loss function for training the Critic1 network and the Critic2 network in combination with the target value and the target action; S28: According to the loss function obtained in step S27, the TD3 architecture is updated with a delayed update strategy, and steps S21 to S28 are repeated with the updated TD3 architecture until the training of the TD3 architecture is completed.

5. The suspension system control method combined with RND adjustment reward according to claim 4, characterized in that: In step S27, the target value is obtained by the following formula: ; in, represents the target value, represents the discount factor, represents the CriticT1 network, represents the CriticT2 network, represents the network parameters of the CriticT1 network, represents the network parameters of the CriticT2 network, represents the target action, represents the next environment state, and r represents the current reward.

6. The suspension system control method combined with RND adjustment reward according to claim 5, characterized in that: In step S27, the loss function for training the Critic1 network and the Critic2 network is: ; ; in, represents the loss function of the Critic1 network, represents the loss function of the Critic2 network, Indicates that the action will and the current state of the environment Input the result of the Critic1 network, represents the network parameters of the Critic1 network, Indicates that the action will and the current state of the environment Input the result of the Critic2 network, represents the network parameters of the Critic2 network, MSE represents the mean square error loss function, represents the target value of the Critic1 network, Represents the target value of the Critic2 network.

7. A readable storage medium, characterized in that: The readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the suspension system control method combined with RND adjustment reward as described in any one of claims 1 to 6 are implemented.

8. An electronic device, characterized in that: include: Memory for storing computer programs; A processor is used to implement the steps of the suspension system control method combined with RND adjustment reward as described in any one of claims 1 to 6 when executing the computer program.

Citation Information

Patent Citations

  • Active suspension reinforcement learning control method based on deep Q neural network

    CN111487863A

  • Sub-target tree mechanical arm obstacle avoidance path planning method based on maximized curiosity

    CN116890339A

  • Self-driving vehicle steering and suspension cooperative control method based on reinforcement learning

    CN119975527A

  • Control device of vehicle

    JP2016104605A