Reward-controllable suspension system control method, medium and electronic device
By introducing a reward judgment mechanism into the TD3 algorithm, dynamically adjusting the weight of the optimization target, the problem of insufficient convergence speed and stability in the multi-objective optimization task is solved, and a more efficient multi-objective optimization effect is achieved.
Patent Information
- Application Number
- CN202510619004.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-05-14
AI Technical Summary
Traditional TD3 algorithms are difficult to effectively adapt to the complex interactions between multiple targets in multi-objective optimization tasks, resulting in inconsistent convergence speed and effect, especially in complex application scenarios, overall performance declines.
A reward judgment mechanism is introduced, and the weights of each optimization goal are dynamically adjusted according to the learning process and environmental conditions, and a TD3 architecture with a reward judgment mechanism is built. Through weight integration and loss function optimization training process, flexible adaptation of multiple goals is achieved.
The convergence speed and stability of the TD3 algorithm in multi-objective optimization tasks are improved, coordination and balance between multiple targets is achieved, and the optimization effect of the agent is improved.
Smart Images

Figure CN120156238B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of vehicle control, and particularly relates to a control method, medium and electronic device for a suspension system with controllable rewards. Background Art
[0002] TD3 (Twin Delayed Deep Deterministic Policy Gradient) algorithm is a reinforcement learning algorithm proposed by Scott Fujimoto et al. in 2018, which is specifically designed for complex control problems in continuous action spaces. The TD3 algorithm is applicable to high-dimensional continuous control tasks that require precise adjustment, such as robot control and active suspension system control of autonomous vehicles.
[0003] Traditional TD3 (Twin Delayed DDPG) algorithm shows the following deficiencies in multi-objective optimization tasks:
[0004] Firstly, the design of the reward function of this algorithm lacks detailed weight control, making it difficult to effectively adapt to the complex interactions between multiple objectives. This deficiency is particularly prominent when dealing with multi-objective tasks with conflicts, making it difficult for the algorithm to maintain consistency in the convergence speed and convergence effect of different objectives. Due to this imbalance, the optimization results of the algorithm often tend to favor certain objectives, resulting in a decline in the overall learning performance and possibly neglecting the optimization of certain key objectives.
[0005] Secondly, a single reward function is adopted in the traditional TD3 framework, making it difficult to comprehensively reflect the mutual influence relationship between multiple objectives in a complex system. Although the algorithm may achieve rapid convergence on certain objectives, the learning of other objectives may lag, ultimately leading to a decline in the overall performance. This problem is particularly significant in complex application scenarios such as engineering control, especially when the optimization objectives are diverse and there are conflict or competition relationships. The convergence and stability of the traditional TD3 algorithm are difficult to meet the actual requirements. Summary of the Invention
[0006] In view of this, the present invention aims to provide a control method, medium and electronic device for a suspension system with controllable rewards, introducing a reward determination mechanism into the TD3 algorithm. The reward determination mechanism can dynamically adjust the importance of each objective according to different stages of the learning process or changes in environmental conditions. This reward determination mechanism allows subjective weights to be introduced during the convergence process of the algorithm, making the algorithm more flexible in adapting to multi-objective optimization requirements.
[0007] To achieve the above object, the technical solution of the present invention is realized as follows:
[0008] A control method for a suspension system with controllable rewards, comprising:
[0009] S1: Determine multiple optimization objectives for controlling the suspension system and the conflict relationships among the multiple optimization objectives;
[0010] S2: Construct a TD3 architecture with a reward determination mechanism and train the TD3 architecture according to the multiple optimization objectives in step S1; wherein,
[0011] The reward determination mechanism assigns weights to the reward values corresponding to each optimization objective according to the conflict relationships and integrates the multiple reward values into a total reward value according to the weights;
[0012] Combine the total reward value to obtain a target value, update the TD3 architecture according to the target value, and obtain a TD3 model with a reward determination mechanism;
[0013] S3: Input the current environmental state that meets the optimization objectives in step S1 into the TD3 model obtained in step S2 to obtain the output action of the predicted suspension system.
[0014] Further, the TD3 architecture in step S2 includes an Actor network, an ActorT network, a Critic1 network, a Critic2 network, a CriticT1 network, and a CriticT2 network; wherein, the current environmental state is input into the Actor network, the output action of the Actor network acts on the environment, and the output action is input into the ActorT network; the action output by the ActorT network and the next environmental state are jointly input into the CriticT1 network and the CriticT2 network; the outputs of the CriticT1 network and the CriticT2 network are combined with the total reward value to obtain a target value; the current environmental state is input into the Critic1 network and the Critic2 network; the Critic1 network and the Critic2 network calculate the loss with the target value in combination with the current environmental state.
[0015] Further, the training process in step S2 includes:
[0016] S21: Initialize the TD3 architecture;
[0017] S22: In the current environmental state, the Actor network generates an action and controls the execution of the action in the suspension system to obtain the next environmental state, as well as the corresponding reward generated by the action in the optimization objective; the reward determination mechanism assigns weights to the reward and integrates the reward into a total reward value according to the weights;
[0018] S23: Repeat step S22 multiple times, integrate the obtained action, next environmental state, current environmental state, and total reward value into experience data and store it in the experience pool;
[0019] S24: Randomly sample a batch of experience data from the experience pool, and based on the experience data, use the ActorT network to generate target actions; the CriticT1 network and the CriticT2 network generate target values according to the target actions;
[0020] S25: Combine the target values and target actions generated in step S24 to calculate the loss functions for training the Critic1 network and the Critic2 network;
[0021] S26: According to the loss functions obtained in step S25, use a delayed update strategy to update the parameters of the TD3 architecture, and repeat steps S21 - S26 with the updated TD3 architecture to complete the training of the TD3 architecture.
[0022] Furthermore, in step S24, the target value is obtained through the following formula:
[0023] ;
[0024] where, represents the target value, represents the discount factor, represents the CriticT1 network, represents the CriticT2 network, represents the network parameters of the CriticT1 network, represents the network parameters of the CriticT2 network, represents the target action, represents the next environmental state.
[0025] Furthermore, in step S25, the loss functions for training the Critic1 network and the Critic2 network are:
[0026] ;
[0027] ;
[0028] where, represents the loss function of the Critic1 network, represents the loss function of the Critic2 network, represents the result obtained by inputting the action and the current environmental state into the Critic1 network, represents the network parameters of the Critic1 network, represents the result obtained by inputting the action and the current environmental state into the said Critic2 network, Denote the network parameters of Critic2 network, and MSE denote the mean square error loss function.
[0029] Furthermore, the process that the reward determination mechanism in step S2 assigns weights to the reward values corresponding to each optimization objective according to the conflict relationship includes:
[0030] There is a conflict relationship between optimization objective A and optimization objective B, and optimization objective A converges preferentially;
[0031] If the reward values of optimization objective A and optimization objective B satisfy the following trigger conditions:
[0032] ;
[0033] wherein, denotes the reward corresponding to optimization objective A, denotes the reward corresponding to optimization objective B, denotes the trigger coefficient;
[0034] The variably controllable factor is obtained by the following formula:
[0035] ;
[0036] wherein, denotes the variably controllable factor, denotes the reward corresponding to optimization objective A, denotes the control factor;
[0037] At this time, the weight of the reward of optimization objective B is h, and the weights of the rewards of other optimization objectives are 1.
[0038] A readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps of the reward controllable suspension system control method provided by the present invention.
[0039] An electronic device includes:
[0040] A memory for storing a computer program;
[0041] A processor for implementing the steps of the reward controllable suspension system control method provided by the present invention when executing the computer program.
[0042] Compared with the prior art, the present invention can achieve the following beneficial effects:
[0043] (1) In the control method of the suspension system with controllable rewards of the present invention, a reward determination mechanism is designed and introduced into the TD3 algorithm. The reward determination mechanism can dynamically adjust the importance of each optimization objective according to different stages of the learning process or changes in environmental conditions. This dynamic adjustment mechanism allows the introduction of subjective weights, enabling the TD3 algorithm to more flexibly adapt to the multi-objective optimization requirements. This innovation not only provides a new idea for the application of deep reinforcement learning in multi-objective optimization but also offers an efficient and flexible solution for complex multi-objective optimization problems in the field of engineering control.
[0044] (2) In the control method of the suspension system with controllable rewards of the present invention, the reward determination mechanism has a variable controllable factor with controllability and time-varying characteristics, which can optimize the convergence behavior of training and make the obtained TD3 model perform more balancedly in multi-objective optimization tasks. Specifically, since the rewards corresponding to the states change at all times, the variable controllable factor is coordinated with the controllable control factor to adaptively adjust according to the learning progress on each objective, thereby dynamically adjusting the reward distribution strategy. This mechanism not only improves the convergence speed of the algorithm but also significantly improves the convergence stability in a multi-objective environment, enabling the intelligent agent to achieve better coordination and balance among conflicting objectives. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] The drawings constituting a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0046] Figure 1 is a schematic flowchart of the control method of the suspension system with controllable rewards according to the embodiment of the present invention;
[0047] Figure 2 is a schematic diagram of the TD3 architecture according to the embodiment of the present invention;
[0048] Figure 3 is a schematic diagram of the structure of the electronic device according to the embodiment of the present invention.
[0049] Description of the reference numerals:
[0050] 1, electronic device; 2, external device; 3, processing unit; 4, bus; 5, network adapter; 6, display; 7, (I / O) interface; 8, system memory; 9, random access memory; 10, cache memory; 11, storage system; 12, utility tool; 13, program module. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0051] In order to make the objectives, technical solutions and advantages of the present invention more clear and understandable, the following further details the present invention in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and do not constitute a limitation to the present invention.
[0052] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments may be combined with each other.
[0053] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "center", "longitudinal", "transverse", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the accompanying drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus cannot be understood as a limitation to the present invention. In addition, the terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first", "second", etc. may explicitly or implicitly include one or more of such features. In the description of the present invention, unless otherwise specified, the meaning of "a plurality" is two or more.
[0054] In the description of the present invention, it should be noted that unless otherwise clearly specified and defined, the terms "mounted", "connected", "coupled" should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral connection; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood through specific circumstances.
[0055] The present invention will be described in detail below with reference to the accompanying drawings and embodiments.
[0056] As Figures 1 to 2 shown, the control method of the suspension system with a combined reward variable controllable factor described in the embodiment of the present invention includes:
[0057] S1: Determine a plurality of optimization objectives for controlling the suspension system and the conflict relationship between the plurality of optimization objectives.
[0058] In different application scenarios, a suspension system usually involves multiple interrelated or competing optimization objectives. In one embodiment, the optimization objectives for controlling the suspension system are respectively the displacement of the sprung mass, the velocity of the sprung mass, and the acceleration of the sprung mass; the displacement of the unsprung mass, the velocity of the unsprung mass, and the acceleration of the unsprung mass. For the convenience of subsequent analysis and elaboration, in one embodiment, without loss of generality, three optimization objectives are set: Target1, Target2, and Target3. Each objective corresponds to a specific performance index and has a direct impact on the overall performance of the system.
[0059] After clarifying each optimization objective, it is necessary to further identify the possible conflict relationships according to the actual situation of the specific research project and the experience of the researchers. In one embodiment, it is assumed that there is an obvious conflict between Target1 and Target3, that is, in the process of optimizing Target1, the performance of Target3 may decline. In addition, the convergence priority of Target1 is set higher than that of Target3, and both Target1 and Target3 are independent of the optimization of Target2, that is, the optimization of Target2 will not affect the optimization of Target1 and Target3. This analysis of the conflict and dependence structure helps to reasonably design the reward weights and convergence strategies in multi-objective optimization to achieve the overall balanced optimization of the system.
[0060] The reward types for each optimization objective include positive reward type, negative reward type (or penalty type), and positive and negative reward type. Among them, the positive reward type determines which behaviors or states should receive positive rewards in the task and only receive positive rewards. The negative reward type determines which behaviors or states should receive negative rewards (or be punished) in the task and only receive negative rewards (or be punished). The positive and negative reward type determines which behaviors or states should receive positive rewards and which behaviors or states should receive negative rewards (i.e., penalties) in the task. For example, when the suspension system completes the target task, positive rewards are given, and negative rewards should be given when it fails or makes unreasonable actions. For the convenience of elaboration, in one embodiment, the rewards all adopt the positive and negative reward type.
[0061] S2: Construct a TD3 architecture with a reward determination mechanism and train the TD3 architecture according to the multiple optimization objectives in step S1; among them, the reward determination mechanism assigns weights to the reward values corresponding to each optimization objective according to the conflict relationship and integrates the multiple reward values into a total reward value; obtain the target value by combining the total reward value, and update the TD3 architecture according to the target value to obtain a TD3 model with a reward determination mechanism.
[0062] Specifically, such as Figure 2As shown, the TD3 architecture includes an Actor network, an ActorT network, a Critic1 network, a Critic2 network, a CriticT1 network, and a CriticT2 network. The current environmental state is input into the Actor network, and the output action of the Actor network acts on the environment and is also input into the ActorT network; the action output by the ActorT network and the next environmental state are jointly input into the CriticT1 network and the CriticT2 network; the outputs of the CriticT1 network and the CriticT2 network are combined with the total reward value obtained by the reward determination mechanism to obtain the target value; the current environmental state is input into the Critic1 network and the Critic2 network; the Critic1 network and the Critic2 network calculate the loss between them and the target value in combination with the current environmental state.
[0063] In a specific embodiment, the Actor network, the ActorT network, the two CriticT networks, and the two Critic networks generally adopt a structure of 3 to 4 layers. Among them, the structures of the Actor network and the ActorT network are the same, and the structures of the two CriticT networks and the two Critic networks correspond to each other, that is, the structure of the CriticT1 network is the same as that of the CriticT1 network, and the structure of the CriticT2 network is the same as that of the CriticT2 network. The input layer of the Actor network can use a node to receive the state, pass through 2 to 3 hidden layers (for example, the number of nodes is 300 and 200, and the ReLU activation function is used), and the output layer generates an action. Since the generated action needs to be restricted within the range of the action space (for example, [−1,1]), the tanh activation function is usually used. If the range of the action space is larger, the output value can be extended to the target range through a linear transformation. The input layers of the two Critic networks both receive the concatenated result of the current environmental state and the action, and there should be two nodes in the input layer. The hidden layer settings are similar to those of the Actor network (also using the ReLU activation function), and the output layer is a linear output for calculating the evaluation value of the generated action. The soft-update TD3 model is adopted to balance performance and computational cost while avoiding the problems of gradient disappearance and overfitting.
[0064] The training process for the TD3 architecture includes:
[0065] S21: Initialize the TD3 architecture.
[0066] In a certain embodiment, the Xavier initialization method is used to randomly initialize the following parameters. The initialization includes:
[0067] Initialize the environmental state s, and the state of the suspension is the state corresponding to the optimization objective, that is, including the displacement, velocity, and acceleration of the sprung mass, and the displacement, velocity, and acceleration of the unsprung mass;
[0068] Randomly initialize the Critic1 network and the Critic2 network. Specifically, randomly initialize the network parameters of the Critic1 network and the network parameters of the Critic2 network for random initialization;
[0069] Initialize the two target Critic networks, namely the CriticT1 network and the CriticT2 network. Specifically, initialize the network parameters of the CriticT1 network and the network parameters of the CriticT2 network for initialization. The initialized network parameters are equal to the network parameters , and the initialized network parameters are equal to the network parameters ;
[0070] Randomly initialize the Actor network. Specifically, randomly initialize the network parameters of the Actor network for random initialization;
[0071] Initialize the target Actor network, namely the ActorT network. Specifically, initialize the network parameters of the ActorT network for initialization. The initialized network parameters are equal to the network parameters ;
[0072] Initialize the hyperparameters for updating the TD3 architecture. The hyperparameters include the discount factor , the soft update frequency , the policy update frequency policy_delay, and the learning rates of the Critic1 network and the Critic2 network. In one embodiment, the learning rate of the Actor network is usually set to 10 -4 , and the learning rates of the two Critic networks are both 10 -3 , and the soft update frequency is set in [0.005, 0.01].
[0073] S22: In the current environmental state, the Actor network generates an action and controls the execution of the action in the suspension system to obtain the next environmental state, as well as the corresponding reward generated by the action in the optimization objective. The reward determination mechanism assigns weights to the rewards and integrates the rewards into the total reward value according to the weights.
[0074] In some embodiments, in the current environmental state s, the Actor network generates an action added with exploration noise ; that is:
[0075] ;
[0076] wherein, denotes the Actor network, and the exploration noise is Gaussian noise.
[0077] Control the execution of the action in the suspension system , obtain the next environmental state, and the corresponding reward generated by the action in the optimization objective. The reward determination mechanism assigns weights to the rewards and integrates the rewards into a total reward value according to the weights.
[0078] In a certain embodiment, control the execution of the action in the suspension system , obtain the next environmental state , and obtain the reward values of the optimization objectives Target1, Target2, and Target3 respectively through the following formula , and :
[0079] ;
[0080] where j = 1, 2, 3, denotes the next environmental state corresponding to the jth optimization objective.
[0081] In some embodiments, the process by which the reward determination mechanism obtains the total reward value includes:
[0082] There is a conflict relationship between optimization objective A and optimization objective B, and optimization objective A converges first;
[0083] If the reward values of optimization objective A and optimization objective B satisfy the following trigger conditions:
[0084] ;
[0085] wherein, denotes the reward corresponding to optimization objective A, denotes the reward corresponding to optimization objective B, denotes the trigger coefficient;
[0086] Obtain the variable controllable factor through the following formula:
[0087] ;
[0088] wherein, denotes the variable controllable factor, represents the reward corresponding to the optimization objective A represents the control factor;
[0089] At this time, the weight of the reward for the optimization objective B is h, and the weights of the rewards for other optimization objectives are 1. Among them, the trigger coefficient and the discount factor are adaptively adjusted according to the research content requirements, the actual situation, and the experience of the researchers.
[0090] In one embodiment, since there is an obvious conflict between Target1 and Target3, and the convergence priority of Target1 is higher than that of Target3, and the trigger coefficient is set to be 10, and the discount factor is -2, then in the reward determination mechanism, it can be understood that:
[0091] If the reward values of the optimization objective Target1 and the optimization objective Target3 satisfy the following trigger conditions:
[0092] ;
[0093] The variable controllable factor is obtained through the following formula:
[0094] ;
[0095] At this time, the weight of the reward for the optimization objective Target3 is h, and the weights of the rewards for the optimization objective Target1 and Target2 are both 1, and then the total reward value r is obtained as:
[0096] .
[0097] S23: Repeat step S22 multiple times, and integrate the obtained actions, the next environmental state, the current environmental state, and the total reward value into experience data and store them in the experience pool.
[0098] It can be understood that by repeating step S22 multiple times, the obtained actions , the next environmental state , the current environmental state and the total reward value r are integrated into experience data and stored in the experience pool.
[0099] S24: Randomly sample a batch of experience data from the experience pool, and use the ActorT network to generate target actions according to the experience data; the CriticT1 network and the CriticT2 network generate target values according to the target actions.
[0100] In some embodiments, a batch of empirical data is randomly sampled from the empirical pool , and based on the empirical data , the ActorT network is used to generate the target action , that is:
[0101] ;
[0102] wherein represents the ActorT network, represents noise, and Gaussian noise is adopted in a certain embodiment.
[0103] The target value is obtained through the Bellman expectation formula:
[0104] ;
[0105] wherein represents the target value, represents the discount factor, represents the CriticT1 network, represents the CriticT2 network, represents the network parameters of the CriticT1 network, represents the network parameters of the CriticT2 network.
[0106] S25: Combine the target value and the target action generated in step S24 to calculate the loss functions for training the Critic1 network and the Critic2 network.
[0107] In a certain embodiment, the loss functions for training the Critic1 network and the Critic2 network are:
[0108] ;
[0109] ;
[0110] wherein represents the loss function of the Critic1 network, represents the loss function of the Critic2 network, represents the result obtained by inputting the action and the current environmental state into the Critic1 network, represents the result obtained by inputting the action and the current environmental state into the Critic2 network, and MSE represents the mean square error loss function.
[0111] S26: Based on the loss function obtained in step S25, perform parameter update on the TD3 architecture using a delayed update strategy, and repeat steps S21 - S26 with the updated TD3 architecture to complete the training of the TD3 architecture.
[0112] In a specific embodiment, repeat steps S21 - S26 600 - 650 times with the updated TD3 architecture. During each round of training, update the parameters through the following equations and parameter :
[0113] ;
[0114] ;
[0115] where, represents taking the gradient of in the loss function ; represents taking the gradient of in the loss function ;
[0116] Update parameter and parameter respectively according to parameter and parameter :
[0117] ;
[0118] ;
[0119] Every policy_delay rounds of training, update parameter and parameter , including:
[0120] Calculate the policy loss through the following equation:
[0121] ;
[0122] Update parameter according to the policy loss through the following equation:
[0123] ;
[0124] where, represents taking the gradient of in the policy loss ;
[0125] Then update parameter through the following equation:
[0126] ;
[0127] Update the current environmental state to the next environmental state .
[0128] S3: Input the current environmental state that meets the optimization objective of step S1 into the TD3 model obtained in step S2 to obtain the output action of the predicted suspension system.
[0129] Correspondingly, according to the embodiments of the present invention, the present invention also provides an electronic device, a readable storage medium, and a computer program product.
[0130] Figure 3 It is a schematic structural diagram of an electronic device 1 provided in the embodiments of the present invention. Figure 3 It shows a block diagram of an exemplary electronic device 1 suitable for implementing the embodiments of the present invention. Figure 3 The shown electronic device 1 is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present invention.
[0131] As Figure 3 shown, the electronic device 1 is presented in the form of a general-purpose computing device. The electronic device 1 is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are only examples and are not intended to limit the implementation of the present invention described herein and / or claimed.
[0132] The components of the electronic device 1 may include but are not limited to: one or more processors or processing units 3, a system memory 8, and a bus 4 connecting different system components (including the system memory 8 and the processing unit 3).
[0133] The bus 4 represents one or more of several types of bus structures, including a memory bus or a memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any of the multiple bus structures. For example, these architectures include but are not limited to the Industry Standard Architecture (ISA) bus, the Micro Channel Architecture (MAC) bus, the Enhanced ISA bus, the Video Electronics Standards Association (VESA) local bus, and the Peripheral Component Interconnect (PCI) bus.
[0134] The electronic device 1 typically includes various computer system readable media. These media can be any available media accessible to the electronic device 1, including volatile and non-volatile media, removable and non-removable media.
[0135] The system memory 8 can include computer system readable media in the form of volatile memory, such as random access memory (RAM) 9 and / or cache memory 10. The electronic device 1 can further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, the storage system 11 can be used for reading and writing on a non-removable, non-volatile magnetic medium ( Figure 3 not shown, commonly referred to as a "hard disk drive"). Although Figure 3 not shown in the figure, a disk drive for reading and writing on a removable non-volatile disk (such as a "floppy disk") and an optical disk drive for reading and writing on a removable non-volatile optical disk (such as a CD-ROM, DVD-ROM or other optical media) can be provided. In these cases, each drive can be connected to the bus 4 through one or more data media interfaces. The system memory 8 can include at least one program product having a set (such as at least one) of program modules configured to perform the functions of the embodiments of the present invention.
[0136] A program / utility 12 having a set (at least one) of program modules 13 can be stored, for example, in the system memory 8. Such program modules 13 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment. The program modules 13 generally perform the functions and / or methods in the embodiments described in the present invention.
[0137] The electronic device 1 can also communicate with one or more external devices 2 (such as a keyboard, a pointing device, a display 6, etc.), and can also communicate with one or more devices that enable a user to interact with the electronic device 1, and / or communicate with any device that enables the electronic device 1 to communicate with one or more other computing devices (such as a network card, a modem, etc.). Such communication can be carried out through an input / output (I / O) interface 7. And, the electronic device 1 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN) and / or a public network, such as the Internet) through a network adapter 5. As Figure 3 shown, the network adapter 5 communicates with other modules of the electronic device 1 through the bus 4. It should be understood that although not shown in the figure, other hardware and / or software modules can be used in combination with the electronic device 1, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0138] The processing unit 3 executes various functional applications and data processing by running the programs stored in the system memory 8, for example, implementing the reward-controllable suspension system control method provided by the embodiments of the present invention.
[0139] Embodiments of the present invention also provide a non-transitory computer-readable storage medium storing computer instructions, on which a computer program is stored. When the program is executed by a processor, the reward-controllable suspension system control method provided by all the embodiments of the present application is implemented.
[0140] The computer storage medium of the embodiments of the present invention may be any combination of one or more computer-readable media. The computer-readable media may be a computer-readable signal medium or a computer-readable storage medium. More specific examples (non-exhaustive list) of the computer-readable storage medium include: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, the computer-readable storage medium may be any tangible medium that contains or stores a program, which can be used by or in connection with an instruction execution system, apparatus, or device.
[0141] The computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries the computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium may also be any computer-readable medium other than the computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device.
[0142] The program code contained on a computer-readable medium can be transmitted using any suitable medium, including but not limited to wireless, wire, optical fiber cable, RF, etc., or any suitable combination of the foregoing. The computer program code for performing the operations of the present invention can be written in one or more programming languages or combinations thereof, and the programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages - such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any kind of network including a local area network (LAN) or a wide area network (WAN), or, it can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).
[0143] An embodiment of the present invention also provides a computer program product, including a computer program, which when executed by a processor implements the reward-controllable suspension system control method according to the foregoing.
[0144] It should be understood that various forms of the flow shown above can be used, with steps reordered, added, or deleted. For example, the steps described in the disclosure of the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution disclosed in the present invention can be achieved, and no limitations are imposed herein.
[0145] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A control method for a suspension system with controllable rewards, characterized in that, Including: S1: Determine multiple optimization objectives for controlling the suspension system and the conflict relationships between the multiple optimization objectives; S2: Construct a TD3 architecture with a reward determination mechanism and train the TD3 architecture according to the multiple optimization objectives in step S1; wherein, The reward determination mechanism assigns weights to the reward values corresponding to each optimization objective according to the conflict relationships and integrates the multiple reward values into a total reward value according to the weights; Obtain a target value in combination with the total reward value, update the TD3 architecture according to the target value, and obtain a TD3 model with a reward determination mechanism; Among them, the process of the reward determination mechanism assigning weights to the reward values corresponding to each optimization objective according to the conflict relationships includes: There is a conflict relationship between optimization objective A and optimization objective B, and optimization objective A converges first; If the reward values of optimization objective A and optimization objective B meet the following trigger conditions: ; Among them, represents the reward corresponding to optimization objective A, represents the reward corresponding to optimization objective B, represents the trigger coefficient; Obtain a variable controllable factor through the following formula: ; Among them, represents the variable controllable factor, represents the reward corresponding to the optimization objective A, represents the control factor; At this time, the weight of the reward of optimization objective B is h, and the weights of the rewards of other optimization objectives are 1; S3: Input the current environmental state that meets the optimization objectives in step S1 into the TD3 model obtained in step S2 to obtain the output action for predicting the suspension system.
2. The control method of the suspension system with controllable rewards according to claim 1, characterized in that, The TD3 architecture in step S2 includes an Actor network, an ActorT network, a Critic1 network, a Critic2 network, a CriticT1 network, and a CriticT2 network; wherein, The current environmental state is input into the Actor network, the output action of the Actor network acts on the environment, and the output action is input into the ActorT network; The action output by the ActorT network and the next environmental state are jointly input into the CriticT1 network and the CriticT2 network; The outputs of the CriticT1 network and the CriticT2 network are combined with the total reward value to obtain the target value; The current environmental state is input into the Critic1 network and the Critic2 network, and the Critic1 network and the Critic2 network calculate the loss with the target value in combination with the current environmental state.
3. The control method of the suspension system with controllable rewards according to claim 2, wherein The training process in step S2 includes: S21: Initialize the TD3 architecture; S22: In the current environmental state, the Actor network generates an action and controls the suspension system to execute the action to obtain the next environmental state, as well as the corresponding reward generated by the action in the optimization objective; the reward determination mechanism assigns weights to the reward and integrates the rewards into a total reward value according to the weights; S23: Repeat step S22 multiple times, integrate the obtained action, the next environmental state, the current environmental state, and the total reward value into experience data and store them in the experience pool; S24: Randomly sample a batch of experience data from the experience pool, and use the ActorT network to generate a target action according to the experience data; the CriticT1 network and the CriticT2 network generate a target value according to the target action; S25: Calculate the loss functions for training the Critic1 network and the Critic2 network by combining the target value generated in step S24 and the target action; S26: According to the loss functions obtained in step S25, perform parameter update on the TD3 architecture using a delayed update strategy, and repeat steps S21 - S26 with the updated TD3 architecture to complete the training of the TD3 architecture.
4. The reward controllable suspension system control method according to claim 3, characterized in that, In step S24, the target value is obtained through the following formula: ; wherein, represents the target value, represents the discount factor, represents the CriticT1 network, represents the CriticT2 network, represents the network parameters of the CriticT1 network, represents the network parameters of the CriticT2 network, represents the target action, represents the next environmental state.
5. The control method of the suspension system with controllable rewards according to claim 4, characterized in that In step S25, the loss functions for training the Critic1 network and the Critic2 network are: ; ; Among them, represents the loss function of the Critic1 network, represents the loss function of the Critic2 network, represents the action and the current environmental state input into the Critic1 network to obtain the result, represents the network parameters of the Critic1 network, represents the action and the current environmental state input into the Critic2 network to obtain the result, represents the network parameters of the Critic2 network, and MSE represents the mean square error loss function.
6. A readable storage medium, characterized in that, The readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps of the reward - controllable suspension system control method according to any one of claims 1 to 5.
7. An electronic device, characterized in that, Including: A memory for storing a computer program; A processor for implementing the steps of the reward - controllable suspension system control method according to any one of claims 1 to 5 when executing the computer program.
Citation Information
Patent Citations
Method for controlling vertical vibration of electric vehicle driven by hub motor based on TD3 algorithm
CN118966033A