Suspension system reward empowerment method combined with entropy weight method, medium and electronic equipment

Through the suspension system reward and empowerment method combined with the entropy weight method, the reward weight is dynamically adjusted, and the problem of insufficient reward and empowerment mechanism in the multi-objective optimization in the existing technology is solved, and the reasonable balance of multi-objectives in the suspension system control is achieved.

CN120156239AActive Publication Date: 2025-06-17JILIN UNIVERSITY
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510619005.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-06-17
Estimated Expiration
2045-05-14

AI Technical Summary

Technical Problem

The existing technology lacks a simple, effective and versatile reward and empowerment mechanism in multi-objective optimization, which makes it difficult for the agent to balance performance indicators such as comfort, stability and response speed in the control of the suspension system, affecting the overall control effect.

Method used

The suspension system reward empowerment method combined with the entropy weight method is adopted, and the weight coefficient of the reward is dynamically adjusted through the entropy weight method, and the reward weight is adaptively adjusted according to actual needs, so as to promote the network to converge faster and more stably, thereby achieving a reasonable balance of each goal.

Benefits of technology

A reasonable balance of multiple targets in suspension system control is achieved, which significantly enhances the coordination of the agent's performance indicators such as comfort, stability and response speed, and improves the overall control effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120156239A_ABST
    Figure CN120156239A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of automobile control, in particular to a suspension system reward empowerment method combined with an entropy weight method, a medium and electronic equipment, and the method comprises the steps: determining a plurality of optimization targets for controlling a suspension system; a TD3 framework with an entropy weight reward adjustment structure is constructed, the TD3 framework is trained according to the multiple optimization targets, the entropy weight reward adjustment structure periodically collects generated rewards in the training process, and reward weights corresponding to all the optimization targets are calculated through an entropy weight method; and finally, inputting the current environment state conforming to the optimization target in the step S1 into the TD3 model obtained in the step S2, and predicting the output action of the suspension system. According to the method, the weight coefficient of the rewards is dynamically adjusted in combination with the entropy weight method, and the purpose of balancing the convergence degree and convergence speed difference between conflict targets is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of vehicle control, and particularly relates to a suspension system reward weight assignment method, medium and electronic device combining the entropy weight method. Background Art

[0002] At first, the TD3 (Twin Delayed Deep Deterministic Policy Gradient) algorithm was proposed by Scott Fujimoto et al. in 2018. It is an algorithm used to solve continuous action space problems in reinforcement learning and is particularly suitable for high-dimensional continuous control tasks, such as complex systems like robot control and autonomous driving. In recent years, this algorithm has been widely applied to control problems. However, the design of the reward function in the TD3 algorithm is often a key point in various control problems, especially in the multi-objective optimization problems commonly encountered in engineering. Reasonable weight assignment can enable the network to have good simultaneous convergence ability for multiple objectives.

[0003] In the prior art, deep reinforcement learning algorithms with the Actor-Critic framework face some significant defects in multi-objective optimization. Especially in the TD3 algorithm driven by temporal difference, the weight assignment of the reward function directly affects the multi-objective convergence performance of the Critic and Actor networks. However, there is generally a lack of a simple, effective and universal weight assignment mechanism in the prior art, making it difficult to reasonably adjust the weights of each objective in practical applications, resulting in the problem of performance imbalance of the agent in multi-objective optimization. For example, in multi-objective optimization tasks involving suspension control, the weight assignment methods in the prior art are difficult to effectively balance multiple key performance indicators such as comfort, stability and response speed, thus affecting the comprehensive control effect of the suspension system.

[0004] The existing optimization weight matching methods mainly include the following two categories: one is the enumeration weight assignment method, that is, regarding the weight as a hyperparameter and finding the best weight range that meets the optimization goal through a large number of experiments. However, this method is costly and consumes a significant amount of experimental time; the other is the weight assignment method based on the evaluation system, such as the analytic hierarchy process or the advantage function method, etc. These methods often cannot fully reflect the data characteristics in multi-objective optimization, have poor universality, and are prone to failure in complex environments and are difficult to capture the core characteristics of the data flow of the agent. Summary of the Invention

[0005] In view of this, the present invention aims to provide a suspension system reward weight assignment method, medium and electronic device combining the entropy weight method. By introducing the entropy weight method to dynamically adjust the weight coefficient of the reward, the reward weight is adaptively adjusted according to actual needs during the optimization process of different physical quantities, prompting the network to converge faster and more stably, thereby achieving a reasonable balance of each objective.

[0006] To achieve the above object, the technical solution of the present invention is realized as follows: A method for rewarding and weighting a suspension system combined with the entropy weight method, comprising: S1: Determine multiple optimization objectives for controlling the suspension system; S2: Construct a TD3 architecture with an entropy weight reward adjustment structure, and train the TD3 architecture according to the multiple optimization objectives in step S1; wherein, The entropy weight reward adjustment structure periodically collects the rewards generated during the training process, and calculates the reward weights corresponding to each optimization objective by using the entropy weight method; S3: Input the current environmental state that meets the optimization objectives in step S1 into the TD3 model obtained in step S2 to predict the output action of the suspension system.

[0007] Further, the entropy weight reward adjustment structure in step S2 includes a weight assignment trigger module, a reward matrix module, a calculation trigger module, and an entropy weight calculation module; wherein: The weight assignment trigger module determines the sampling period of the rewards and monitors the number of training rounds of the TD3 architecture. When the number of training rounds meets the sampling period, it triggers the reward matrix module; The reward matrix module collects the rewards in the sampling period and forms a reward matrix for each optimization objective; The calculation trigger module performs a fast Fourier transform on the reward matrix of each optimization objective to obtain the main frequency and the main frequency amplitude of the corresponding reward matrix; compares the differences between the main frequencies corresponding to each optimization objective and the differences between the main frequency amplitudes corresponding to each optimization objective, and triggers the entropy weight calculation module according to the comparison results; The entropy weight calculation module calculates the reward weights corresponding to each optimization objective by using the entropy weight method according to the reward matrix.

[0008] Further, the reward matrix module collects the rewards corresponding to each optimization objective in the sampling period to form a data stream matrix, and performs standardization processing on the data stream matrix; normalizes the standardized data stream matrix column by column to obtain a reward matrix.

[0009] Further, in the entropy weight calculation module, the information entropy corresponding to each optimization objective is calculated by the following formula: ; wherein, represents the information entropy corresponding to the jth optimization objective, represents the element in the reward matrix; According to the information entropy, the information utility value corresponding to each optimization objective is calculated by the following formula: ; wherein, Denote the information utility value corresponding to the j-th optimization objective; Calculate the reward weight corresponding to each optimization objective through the following formula: ; where Denote the reward weight corresponding to the j-th optimization objective.

[0010] Furthermore, the TD3 architecture in step S2 further includes an Actor network, an ActorT network, a Critic1 network, a Critic2 network, a CriticT1 network, and a CriticT2 network; where The Actor network receives the current environmental state and applies the generated execution action to the environment to generate the next environmental state; The ActorT network receives the next environmental state and generates a target action; The CriticT1 network and the CriticT2 network simultaneously receive the target action and the next environmental state, and the outputs of the CriticT1 network and the CriticT2 network are combined with the current reward to obtain the target value; The Critic1 network and the Critic2 network simultaneously receive and calculate the loss with respect to the target value according to the current environmental state.

[0011] Furthermore, the training process in step S2 includes: S21: Initialize the TD3 architecture; S22: In the current environmental state, the Actor network generates an action and controls the execution action in the suspension system to obtain the next environmental state, and the corresponding rewards generated by the action in each optimization objective are combined with the corresponding reward weights to calculate the current reward; S23: Repeat step S22 multiple times, and integrate the obtained action, the next environmental state, the current environmental state, and the current reward into experience data and store them in the experience pool; S24: Randomly sample a batch of experience data from the experience pool, and use the ActorT network to generate a target action according to the experience data; the CriticT1 network and the CriticT2 network use the Bellman expectation formula to generate the target value according to the target action; S25: Combine the target value and the target action generated in step S24 to calculate the loss functions of the Critic1 network and the Critic2 network for training; S26: According to the loss functions obtained in step S25, perform parameter updates on the TD3 architecture using a delayed update strategy; S27: Determine whether the current training round meets the sampling period: If it meets, adjust the reward weights corresponding to each optimization objective using entropy-weighted rewards, and execute step S28; if it does not meet, directly proceed to step S28; S28: Repeat steps S21 - S27 with the updated TD3 architecture to complete the training of the TD3 architecture.

[0012] Further, in step S24, the target value is obtained through the following formula: ; where, represents the target value, represents the discount factor, represents the CriticT1 network, represents the CriticT2 network, represents the network parameters of the CriticT1 network, represents the network parameters of the CriticT2 network, represents the target action, represents the next environmental state, and r represents the current reward.

[0013] Further, in step S25, the loss functions for training the Critic1 network and the Critic2 network are: ; ; where, represents the loss function of the Critic1 network, represents the loss function of the Critic2 network, represents the result obtained by inputting the action and the current environmental state into the Critic1 network, represents the network parameters of the Critic1 network, represents the result obtained by inputting the action and the current environmental state into the Critic2 network, represents the network parameters of the Critic2 network, and MSE represents the mean squared error loss function.

[0014] A readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps of the suspension system reward weight assignment method combined with the entropy weight method provided by the present invention.

[0015] An electronic device includes: A memory for storing a computer program; A processor, which is used to implement the steps of the suspension system reward weighting method combined with the entropy weight method provided by the present invention when executing a computer program.

[0016] Compared with the prior art, the present invention can achieve the following beneficial effects: (1) In the suspension system reward weighting method combined with the entropy weight method of the present invention, the entropy weight method can well collect the fluctuation information of each state quantity in the actual training process, and can naturally assign large weights to the targets with large fluctuations and unstable training, and assign small weights to the targets with small fluctuations and stable training, so as to achieve the purpose of balancing the convergence degree and convergence speed difference between conflicting targets.

[0017] (2) In the suspension system reward weighting method combined with the entropy weight method of the present invention, a data flow matrix is used to construct a reward matrix, providing a new perspective for data analysis in deep reinforcement learning; (3) In the suspension system reward weighting method combined with the entropy weight method of the present invention, the data flow matrix and the entropy weight method are combined to capture the fluctuation information in the characteristics of reward data, realizing low-cost and high-efficiency reward weighting, which can significantly enhance the coordination between the performance indicators such as comfort, stability and response speed of the agent controlling the suspension system, making the algorithm easier to be transplanted into the actual engineering scenario, and at the same time can greatly improve the comprehensive control effect of the suspension system. Description of the Drawings

[0018] The drawings constituting a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings: Figure 1 It is a schematic flow chart of the suspension system reward weighting method combined with the entropy weight method according to the embodiment of the present invention; Figure 2 It is a schematic diagram of the TD3 architecture according to the embodiment of the present invention; Figure 3 It is a schematic diagram of the entropy weight reward adjustment structure according to the embodiment of the present invention; Figure 4 It is a schematic diagram of the electronic device according to the embodiment of the present invention.

[0019] Description of the reference numerals: 1. Electronic device; 2. External device; 3. Processing unit; 4. Bus; 5. Network adapter; 6. Display; 7. (I / O) interface; 8. System memory; 9. Random access memory; 10. Cache memory; 11. Storage system; 12. Utility tool; 13. Program module. Detailed Embodiments

[0020] In order to make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and do not constitute a limitation to the present invention.

[0021] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments may be combined with each other.

[0022] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "center", "longitudinal", "transverse", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the accompanying drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to the present invention. In addition, the terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the technical features indicated. Thus, the features defined with "first", "second", etc. may explicitly or implicitly include one or more of such features. In the description of the present invention, unless otherwise specified, the meaning of "plurality" is two or more.

[0023] In the description of the present invention, it should be noted that unless otherwise clearly defined and limited, the terms "mounted", "connected", "coupled" should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral connection; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood through specific situations.

[0024] The present invention will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.

[0025] As Figures 1 to 2 shown, the suspension system reward weighting method combining the entropy weight method according to the embodiment of the present invention includes: S1: Determine a plurality of optimization objectives for controlling the suspension system.

[0026] In one embodiment, the optimization objectives for controlling the suspension system are respectively the displacement of the sprung mass, the velocity of the sprung mass, the acceleration of the sprung mass; the displacement of the unsprung mass, the velocity of the unsprung mass, and the acceleration of the unsprung mass.

[0027] S2: Construct a TD3 architecture with an entropy-weighted reward adjustment structure, and train the TD3 architecture according to multiple optimization objectives in step S1.

[0028] Specifically, the TD3 architecture includes an Actor network, an ActorT network, a Critic1 network, a Critic2 network, a CriticT1 network, and a CriticT2 network, as well as an entropy-weighted reward adjustment structure. The Actor network receives the current environmental state and applies the generated execution action to the environment to generate the next environmental state. The ActorT network receives the next environmental state and generates a target action. The CriticT1 network and the CriticT2 network simultaneously receive the target action and the next environmental state, and the outputs of the CriticT1 network and the CriticT2 network are combined with the current reward to obtain a target value. The Critic1 network and the Critic2 network simultaneously receive and calculate the loss with the target value according to the current environmental state. The entropy-weighted reward adjustment structure periodically collects the generated rewards during the training process and uses the entropy-weighting method to calculate the reward weights corresponding to each optimization objective, thereby changing the rewards within the next collection.

[0029] It can be understood that in the TD3 architecture provided by the present invention, the Actor network receives the current environmental state , and applies the generated execution action to the environment to generate the next environmental state . The ActorT network receives the next environmental state and generates a target action . The CriticT1 network and the CriticT2 network simultaneously receive the target action and the next environmental state , and combine the output target evaluation value and with the current reward r to obtain the target value y, represents the network parameters of the CriticT1 network, represents the network parameters of the CriticT2 network. The Critic1 network and the Critic2 network simultaneously receive the current environmental state and the next environmental state , and the Critic1 network and the Critic2 network simultaneously receive and calculate the loss with the target value y according to the current environmental state s. The entropy-weighted reward adjustment structure periodically collects the generated rewards during the training process and uses the entropy-weighting method to calculate the reward weights corresponding to each optimization objective, thereby changing the rewards within the next collection.

[0030] In a specific embodiment, the Actor network, the Actor-T network, the two CriticT networks, and the two Critic networks generally adopt a structure of 3 to 4 layers. Among them, the structures of the Actor network and the Actor-T network are the same, and the structures of the two CriticT networks and the two Critic networks correspond to each other. That is, the structure of the CriticT1 network is the same as that of the CriticT1 network, and the structure of the CriticT2 network is the same as that of the CriticT2 network. The input layer of the Actor network can use a node to receive the state. After passing through 2 to 3 hidden layers (for example, the number of nodes is 300 and 200, and the ReLU activation function is used), the output layer generates an action. Since the generated action needs to be restricted within the range of the action space (for example, [−1,1]), the tanh activation function is usually used. If the range of the action space is larger, the output value can be extended to the target range through a linear transformation. The input layers of the two Critic networks both receive the concatenated result of the current environmental state and the action. The input layer should have two nodes. The setting of the hidden layer is similar to that of the Actor network (also using the ReLU activation function). The output layer is a linear output, which is used to calculate the evaluation value of the generated action. The soft-update TD3 model is adopted to balance the performance and the computational cost, and at the same time avoid the problems of gradient disappearance and overfitting.

[0031] In some embodiments, the entropy weight reward adjustment structure includes a weight assignment trigger module, a reward matrix module, a calculation trigger module, and an entropy weight calculation module. The weight assignment trigger module determines the sampling period of the reward and monitors the number of training rounds of the TD3 architecture. When the number of training rounds meets the sampling period, it triggers the reward matrix module; the reward matrix module collects the rewards in the sampling period and forms a reward matrix for each optimization target; the calculation trigger module performs a fast Fourier transform on the reward matrix of each optimization target to obtain the main frequency and the main frequency amplitude of the corresponding reward matrix; compares the differences between the main frequencies corresponding to each optimization target and the differences between the main frequency amplitudes corresponding to each optimization target, and triggers the entropy weight calculation module according to the comparison results; the entropy weight calculation module calculates the reward weights corresponding to each optimization target using the entropy weight method based on the reward matrix. The sampling period is determined according to the actual situation. In a specific embodiment, when the total number of training rounds is 1000, the sampling period can be set to sample once every 100 steps. The reward matrix module collects the rewards corresponding to each optimization target in the sampling period to form a data stream matrix, and performs a normalization process on the data stream matrix. The normalized data stream matrix is normalized column by column to obtain the reward matrix. In a specific embodiment, the triggering conditions for triggering the entropy weight calculation module according to the differences between the main frequencies and the differences between the main frequency amplitudes are determined according to the actual situation.

[0032] It should be noted that in the category of deep reinforcement learning algorithms, not only the TD3 algorithm has an environment, but all algorithms belonging to this category have an environment. Therefore, in order to improve the generality of the method and also to propose a new perspective on data analysis in deep reinforcement learning, a data stream will be defined in the category of deep reinforcement learning algorithms. A data stream refers to a column of data constructed based on a certain logical relationship, which can be abstracted into a 1-row and n-column matrix. The elements in each row are determined by the same logical relationship, and the data structure of its elements is not limited. A data stream matrix is an m×n matrix, where each row must be a separate data stream, and the matrix has no limit on the element data structure. For example, in the multi-objective optimization problem of deep reinforcement learning, m optimization objectives can be used as row indicators, and the n values generated by each state quantity to be optimized in the process of interacting with the environment in the causal logical order are arranged as row vectors according to the row indicators to form a data stream matrix. Corresponding to the reward matrix module of the present invention, n time-series rewards of each of the m optimization objectives in the acquisition sampling period are collected to form an m×n data stream matrix. This step can be completed using popular programming software such as python and matlab.

[0033] In addition, the fluctuations in the data will cause fluctuations in the loss function during the training process, making the gradient update unstable. High-fluctuation training data may lead to large changes in the gradient during the optimization process, thereby slowing down the convergence speed of the model. In the multi-objective optimization algorithm based on TD3, the fluctuations in the rewards of each objective to be optimized are generally concerned. When the amplitude fluctuations between the rewards corresponding to each objective to be optimized are large and the fluctuation frequencies are too different, the network will have inconsistent convergence effects on different objective quantities, and even the convergence degrees and convergence speeds will differ greatly. At this time, it is said that the multi-objective convergence performance is poor. Therefore, the calculation trigger module designed in the entropy weight reward adjustment structure of the present invention performs a fast Fourier transform on the reward matrices of different optimization objectives to obtain the main frequency and main frequency amplitude of their reward data streams, and preliminarily compares the differences in the main frequencies and amplitudes of each item.

[0034] The Entropy Weight Method is an objective method for determining the weights of indicators in multi-index decision-making problems. It uses the concept of information entropy to reflect the degree of difference between indicators, thereby giving the weights of each indicator. The core idea of the Entropy Weight Method is that the more dispersed the values of an indicator are among different samples, the greater the amount of information contained in the indicator and the higher the weight; conversely, if the values of the indicator do not vary much among samples, its weight will be lower because its discrimination ability for the evaluation results is weaker. From the above concept of the Entropy Weight Method, it can be seen that the role of the Entropy Weight Method is to evaluate the distribution of a set of data. The data stream with large fluctuations has a high weight, and the stable data stream has a low weight. For a certain optimization goal, the more dispersed the distribution of its reward data stream, the lower the training efficiency, and the more stable the reward distribution, the higher the training efficiency. To correct the difference in the target convergence effect caused by this data fluctuation difference through weight setting, the Entropy Weight Method is just right. Based on this, the present invention combines the Entropy Weight Method to determine the reward weight.

[0035] In a specific embodiment, in the entropy calculation module, the information entropy corresponding to each optimization goal is calculated by the following formula: ; where represents the information entropy corresponding to the j-th optimization goal, represents the element in the reward matrix; According to the information entropy, the information utility value corresponding to each optimization goal is calculated by the following formula: ; where represents the information utility value corresponding to the j-th optimization goal; The reward weight corresponding to each optimization goal is calculated by the following formula: ; where represents the reward weight corresponding to the j-th optimization goal. At this time, the sum of the reward weights of the optimization goals is 1.

[0036] After determining the reward weights corresponding to each optimization goal, the reward in the next sampling period is determined by the following formula: ; where represents the current reward of the j-th optimization, represents the previous reward of the j-th optimization.

[0037] In some embodiments, the training process in step S2 includes: S21: Initialize the TD3 architecture. In a specific embodiment, the following parameters are randomly initialized using the Xavier initialization method. The initialization includes: Initialize the environmental state; the state of the suspension is the state corresponding to the optimization objective, that is, including the displacement, velocity, and acceleration of the sprung mass, and the displacement, velocity, and acceleration of the unsprung mass; Randomly initialize the Critic1 network and the Critic2 network. Specifically, initialize the network parameters of the Critic1 network and the network parameters of the Critic2 network for random initialization; Initialize the two target Critic networks, namely the CriticT1 network and the CriticT2 network. Specifically, initialize the network parameters of the CriticT1 network and the network parameters of the CriticT2 network for initialization. The initialized network parameters are equal to the network parameters , and the initialized network parameters are equal to the network parameters ; Randomly initialize the Actor network. Specifically, initialize the network parameters of the Actor network for random initialization; Initialize the target Actor network, that is, the ActorT network. Specifically, initialize the network parameters of the ActorT network for initialization. The initialized network parameters are equal to the network parameters ; Initialize the hyperparameters for updating the TD3 architecture. The hyperparameters include the discount factor , the soft update frequency , the policy update frequency policy_delay, and the learning rates of the Critic1 network and the Critic2 network. In one embodiment, the learning rate of the Actor network is usually set to 10 -4 , and the learning rates of the two Critic networks are both 10 -3 , and the soft update frequency is set in [0.005, 0.01].

[0038] S22: In the current environmental state, the Actor network generates an action and controls the execution of the action in the suspension system to obtain the next environmental state, and calculates the combined corresponding rewards of the action in each optimization objective with the corresponding reward weights to obtain the current reward.

[0039] In some embodiments, in the current environmental state s, the Actor network generates an action with exploration noise added , that is: ​ ; Among them, represents the Actor network, and the exploration noise is Gaussian noise.

[0040] Control the actions performed in the suspension system , obtain the next environmental state, and calculate the combined corresponding reward of the action in the optimization objective with the corresponding reward weight to obtain the current reward r, that is: .

[0041] S23: Repeat step S22 multiple times, and integrate the obtained actions, next environmental state, current environmental state, and current reward into experience data and store them in the experience pool.

[0042] It can be understood that by repeating step S22 multiple times, the obtained actions , next environmental state , current environmental state and the total reward value r are integrated into experience data and stored in the experience pool.

[0043] S24: Randomly sample a batch of experience data from the experience pool, and use the ActorT network to generate target actions according to the experience data; the CriticT1 network and the CriticT2 network use the Bellman expectation formula to generate target values according to the target actions.

[0044] In some embodiments, randomly sample a batch of experience data from the experience pool , and according to the experience data , use the ActorT network to generate target actions , that is: ; Among them, represents the ActorT network, represents the noise, and Gaussian noise is used in a certain embodiment.

[0045] Obtain the target value through the Bellman expectation formula: ; Among them, represents the target value, represents the discount factor, represents the CriticT1 network, represents the CriticT2 network.

[0046] S25: Calculate the loss functions for training the Critic1 network and the Critic2 network by combining the target value and the target action generated in step S24.

[0047] In one embodiment, the loss functions for training the Critic1 network and the Critic2 network are: ; ; where represents the loss function of the Critic1 network, represents the loss function of the Critic2 network, represents the result obtained by inputting the action and the current environmental state into the Critic1 network, represents the result obtained by inputting the action and the current environmental state into the Critic2 network, and MSE represents the mean squared error loss function.

[0048] S26: According to the loss functions obtained in step S25, update the parameters of the TD3 architecture using a delayed update strategy.

[0049] During each round of training, update the parameters and the parameter through the following formula: ; ; where represents taking the gradient of the loss function with respect to in it, represents taking the gradient of the loss function with respect to in it.

[0050] Update the parameter and the parameter respectively according to the parameters and the parameter : ; ; Every policy_delay rounds of training, update the parameters and the parameter , including: Calculate the policy loss through the following formula: ; According to the policy loss , the parameter is updated by the following formula: ; wherein, represents taking the gradient of in the policy loss ; Then, the parameter is updated by the following formula: ; The current environmental state is updated to the next environmental state .

[0051] S27: Determine whether the current training round meets the sampling period: If it meets, use the entropy weight reward adjustment structure to adjust the reward weights corresponding to each optimization target, and execute step S28; if it does not meet, directly proceed to step S28.

[0052] S28: Repeat steps S21 - S27 with the updated TD3 architecture to complete the training of the TD3 architecture.

[0053] S3: Input the current environmental state that meets the optimization target in step S1 into the TD3 model obtained in step S2 to predict the output action of the suspension system.

[0054] Correspondingly, according to the embodiments of the present invention, the present invention also provides an electronic device, a readable storage medium, and a computer program product.

[0055] Figure 4 It is a schematic structural diagram of an electronic device 1 provided in the embodiments of the present invention. Figure 4 It shows a block diagram of an exemplary electronic device 1 suitable for implementing the embodiments of the present invention. Figure 4 The shown electronic device 1 is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present invention.

[0056] As Figure 4 shown, the electronic device 1 is presented in the form of a general - purpose computing device. The electronic device 1 is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described herein and / or claimed.

[0057] The components of the electronic device 1 may include, but are not limited to: one or more processors or processing units 3, a system memory 8, and a bus 4 that couples different system components (including the system memory 8 and the processing unit 3).

[0058] The bus 4 represents one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processor, or a local bus using any of a variety of bus structures. By way of example, these architectures include, but are not limited to, Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MAC) bus, Enhanced ISA bus, Video Electronics Standards Association (VESA) local bus, and Peripheral Component Interconnect (PCI) bus.

[0059] The electronic device 1 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by the electronic device 1, including volatile and nonvolatile media, removable and non-removable media.

[0060] The system memory 8 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 9 and / or cache memory 10. The electronic device 1 may further include other removable / non-removable, volatile / nonvolatile computer system storage media. By way of example only, a storage system 11 may be used for reading and writing on non-removable, nonvolatile magnetic media ( Figure 4 not shown and typically called a "hard disk drive"). Although Figure 4 not shown in the figure, a disk drive for reading and writing on removable nonvolatile disks (such as a "floppy disk"), and an optical disk drive for reading and writing on removable nonvolatile optical disks (such as a CD-ROM, DVD-ROM, or other optical media) may be provided. In these cases, each drive may be connected to the bus 4 through one or more data media interfaces. The system memory 8 may include at least one program product having a set (e.g., at least one) of program modules that are configured to perform the functions of the embodiments of the present invention.

[0061] A program / utility 12 having a set (at least one) of program modules 13 may be stored, for example, in the system memory 8. Such program modules 13 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment. The program modules 13 typically perform the functions and / or methods described in the embodiments of the present invention.

[0062] The electronic device 1 can also communicate with one or more external devices 2 (such as a keyboard, a pointing device, a display 6, etc.), and can also communicate with one or more devices that enable a user to interact with the electronic device 1, and / or communicate with any device that enables the electronic device 1 to communicate with one or more other computing devices (such as a network card, a modem, etc.). Such communication can be carried out through an input / output (I / O) interface 7. Moreover, the electronic device 1 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter 5. As Figure 4 shown, the network adapter 5 communicates with other modules of the electronic device 1 through a bus 4. It should be understood that although not shown in the figure, other hardware and / or software modules can be used in combination with the electronic device 1, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.

[0063] The processing unit 3 executes various functional applications and data processing by running programs stored in the system memory 8, for example, implementing the suspension system reward empowerment method combining the entropy weight method provided by the embodiments of the present invention.

[0064] The embodiments of the present invention also provide a non-transitory computer-readable storage medium storing computer instructions, on which a computer program is stored, wherein when the program is executed by a processor, the suspension system reward empowerment method combining the entropy weight method provided by all the embodiments of the present application is implemented.

[0065] The computer storage medium of the embodiments of the present invention can adopt any combination of one or more computer-readable media. The computer-readable media can be a computer-readable signal medium or a computer-readable storage medium. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, the computer-readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, apparatus, or device.

[0066] A computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device.

[0067] The program code contained on a computer-readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing. The computer program code for performing the operations of the present invention may be written in one or more programming languages or combinations thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and also including conventional procedural programming languages - such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0068] An embodiment of the present invention also provides a computer program product, including a computer program, which when executed by a processor implements the suspension system reward empowerment method combined with the entropy weight method as described above.

[0069] It should be understood that various forms of the processes shown above may be used, with steps reordered, added, or deleted. For example, the steps recited in the disclosure of the present invention may be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution disclosed in the present invention can be achieved, and no limitation is imposed herein.

[0070] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions may be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A suspension system reward weighting method combined with entropy weight method, characterized in that: include: S1: Determine multiple optimization objectives for controlling the suspension system; S2: constructing a TD3 architecture with an entropy-weighted reward adjustment structure, and training the TD3 architecture according to the multiple optimization objectives of step S1; wherein, The entropy weight reward adjustment structure periodically collects rewards generated during the training process, and uses the entropy weight method to calculate the reward weights corresponding to each optimization goal; S3: Inputting the current environmental state that meets the optimization target of step S1 into the TD3 model obtained in step S2 to predict the output action of the suspension system.

2. The suspension system reward weighting method combined with the entropy weight method according to claim 1, characterized in that: The entropy weight reward adjustment structure in step S2 includes a weighting trigger module, a reward matrix module, a calculation trigger module and an entropy weight calculation module; wherein: The weighted trigger module determines a sampling period of the reward and monitors the number of training rounds of the TD3 architecture, and triggers the reward matrix module when the number of training rounds meets the sampling period; The reward matrix module collects rewards in the sampling period and forms a reward matrix for each optimization target; The calculation trigger module performs fast Fourier transform on the reward matrix of each optimization target to obtain the main frequency and main frequency amplitude of the corresponding reward matrix; compares the difference between the main frequencies corresponding to each optimization target and the difference between the main frequency amplitudes corresponding to each optimization target, and triggers the entropy weight calculation module according to the comparison result; The entropy weight calculation module calculates the reward weights corresponding to each optimization target using the entropy weight method according to the reward matrix.

3. The suspension system reward weighting method combined with the entropy weight method according to claim 2 is characterized in that: The reward matrix module collects rewards corresponding to each optimization target in the sampling period to form a data flow matrix, and the data flow matrix is ​​standardized; the standardized data flow matrix is ​​normalized by column to obtain the reward matrix.

4. The suspension system reward weighting method combined with the entropy weight method according to claim 2, characterized in that: In the entropy weight calculation module, the information entropy corresponding to each optimization target is calculated by the following formula: ; in, represents the information entropy corresponding to the jth optimization objective, represents an element in the reward matrix; According to the information entropy, the information utility value corresponding to each optimization objective is calculated by the following formula: ; in, Represents the information utility value corresponding to the jth optimization objective; The reward weight corresponding to each optimization objective is calculated by the following formula: ; in, Represents the reward weight corresponding to the j-th optimization objective.

5. The suspension system reward weighting method combined with the entropy weight method according to claim 2, characterized in that: The TD3 architecture in step S2 also includes an Actor network, an ActorT network, a Critic1 network, a Critic2 network, a CriticT1 network and a CriticT2 network; wherein, The Actor network receives the current environment state and applies the generated execution action to the environment to generate the next environment state; The ActorT network receives the next environment state and generates a target action; The CriticT1 network and the CriticT2 network simultaneously receive the target action and the next environment state, and the outputs of the CriticT1 network and the CriticT2 network are combined with the current reward to obtain a target value; The Critic1 network and the Critic2 network simultaneously receive and calculate the loss between the current environment state and the target value according to the current environment state.

6. The suspension system reward weighting method combined with the entropy weight method according to claim 5, characterized in that: The training process in step S2 includes: S21: Initialize the TD3 architecture; S22: Under the current environment state, the Actor network generates an action and controls the suspension system to execute the action, thereby obtaining the next environment state and the corresponding rewards generated by the action at each optimization target combined with the corresponding reward weight calculation to obtain the current reward; S23: Repeat step S22 multiple times, integrate the obtained action, the next environment state, the current environment state and the current reward into experience data and store them in the experience pool; S24: randomly sampling a batch of experience data from the experience pool, and generating a target action using the ActorT network based on the experience data; the CriticT1 network and the CriticT2 network generate a target value based on the target action using the Bellman expectation formula; S25: Calculate the loss function for training the Critic1 network and the Critic2 network in combination with the target value and the target action generated in step S24; S26: According to the loss function obtained in step S25, the TD3 architecture is updated with a delayed update strategy; S27: Determine whether the current number of training rounds meets the sampling period: if so, use the entropy weight reward adjustment structure to adjust the reward weights corresponding to each optimization target, and execute step S28; if not, directly execute step S28; S28: Repeat steps S21 to S27 with the updated TD3 architecture to complete the training of the TD3 architecture.

7. The suspension system reward weighting method combined with entropy weight method according to claim 5, characterized in that: In step S24, the target value is obtained by the following formula: ; in, represents the target value, represents the discount factor, represents the CriticT1 network, represents the CriticT2 network, represents the network parameters of the CriticT1 network, represents the network parameters of the CriticT2 network, represents the target action, represents the next environment state, and r represents the current reward.

8. The suspension system reward weighting method combined with entropy weight method according to claim 5, characterized in that: In step S25, the loss function for training the Critic1 network and the Critic2 network is: ; ; in, represents the loss function of the Critic1 network, represents the loss function of the Critic2 network, Indicates that the action will and the current state of the environment Input the result obtained by the Critic1 network, represents the network parameters of the Critic1 network, Indicates that the action will and the current state of the environment Input the result of the Critic2 network, represents the network parameters of the Critic2 network, and MSE represents the mean square error loss function.

9. A readable storage medium, characterized in that: The readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the suspension system reward weighting method combined with the entropy weight method as described in any one of claims 1 to 8 are implemented.

10. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the suspension system reward weighting method combined with the entropy weight method as described in any one of claims 1 to 8 when executing the computer program.

Citation Information

Patent Citations

  • Multi-objective optimization method for single suspension system of vehicle

    CN118586280A

  • Method for controlling vertical vibration of electric vehicle driven by hub motor based on TD3 algorithm

    CN118966033A

  • Energy feedback type active suspension control method based on multi-agent reinforcement learning control

    CN119821065A

  • Method for optimizing PID control parameters of semi-active suspension of vehicle

    WO2024125584A1