Reward Weight Assignment Method, Medium and Electronic Device for Suspension System Combining Entropy Weight Method
By introducing the entropy weight method to dynamically adjust the reward weight in the suspension system, the problem of performance imbalance in the suspension system in multi-objective optimization tasks is solved, and faster and more stable convergence and comprehensive control effects are achieved.
Patent Information
- Application Number
- CN202510619005.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-05-14
AI Technical Summary
In the prior art, it is difficult for the suspension system to effectively balance performance indicators such as comfort, stability and response speed in multi-objective optimization tasks, resulting in an unbalanced performance of the agent during multi-objective optimization. The existing reward function empowerment methods are expensive or have poor universality, making it difficult to capture data characteristics in complex environments.
The entropy weight method is used to dynamically adjust the reward weight. By constructing a TD3 architecture with an entropy weight reward adjustment structure, rewards are periodically collected and reward weights for each optimization target are calculated, and the data flow matrix and fast Fourier transform are combined to achieve a reasonable balance of each target.
It achieves faster and more stable convergence in the suspension system, enhances the coordination between performance indicators such as comfort, stability and response speed, and improves the comprehensive control effect of the suspension system.
Smart Images

Figure CN120156239B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of vehicle control, and particularly relates to a suspension system reward weighting method, medium and electronic device combining the entropy weight method. Background Art
[0002] At first, the TD3 (Twin Delayed Deep Deterministic Policy Gradient) algorithm was proposed by Scott Fujimoto et al. in 2018. It is an algorithm used to solve continuous action space problems in reinforcement learning and is particularly suitable for high-dimensional continuous control tasks, such as complex systems like robot control and autonomous driving. In recent years, this algorithm has been widely applied to control problems. However, the design of the reward function in the TD3 algorithm is often a key point in various control problems, especially in multi-objective optimization problems commonly encountered in engineering. Reasonable weighting can enable the network to have good simultaneous convergence ability for multiple objectives.
[0003] In the prior art, deep reinforcement learning algorithms with the Actor-Critic framework face some significant defects in multi-objective optimization. Especially in the TD3 algorithm driven by temporal difference, the weight assignment of the reward function directly affects the multi-objective convergence performance of the Critic and Actor networks. However, there is generally a lack of a simple, effective and universal weighting mechanism in the prior art, making it difficult to reasonably adjust the weights of each objective in practical applications, resulting in the problem of performance imbalance of the agent in multi-objective optimization. For example, in multi-objective optimization tasks involving suspension control, the existing weighting methods are difficult to effectively balance multiple key performance indicators such as comfort, stability and response speed, thus affecting the comprehensive control effect of the suspension system.
[0004] The existing optimization weight matching methods mainly include the following two categories: one is the enumeration weighting method, that is, regarding the weight as a hyperparameter and finding the best weight range that meets the optimization goal through a large number of experiments. However, this method is costly and consumes a significant amount of experimental time; the other is the weighting method based on the evaluation system, such as the analytic hierarchy process or the advantage function method, etc. These methods often cannot fully reflect the data characteristics in multi-objective optimization, have poor universality, and are prone to failure in complex environments and are difficult to capture the core characteristics of the agent's data flow. Summary of the Invention
[0005] In view of this, the present invention aims to provide a suspension system reward weighting method, medium and electronic device combining the entropy weight method. By introducing the entropy weight method to dynamically adjust the weight coefficient of the reward, the reward weight is adaptively adjusted according to actual needs during the optimization process of different physical quantities, prompting the network to converge faster and more stably, so as to achieve a reasonable balance of each objective.
[0006] To achieve the above object, the technical solution of the present invention is realized as follows:
[0007] A method for rewarding and weighting a suspension system combined with the entropy weight method, comprising:
[0008] S1: Determine multiple optimization objectives for controlling the suspension system;
[0009] S2: Construct a TD3 architecture with an entropy weight reward adjustment structure, and train the TD3 architecture according to the multiple optimization objectives in step S1; wherein,
[0010] The entropy weight reward adjustment structure periodically collects the generated rewards during the training process, and uses the entropy weight method to calculate the reward weights corresponding to each optimization objective;
[0011] S3: Input the current environmental state that meets the optimization objectives in step S1 into the TD3 model obtained in step S2 to predict the output action of the suspension system.
[0012] Further, the entropy weight reward adjustment structure in step S2 includes a weight assignment trigger module, a reward matrix module, a calculation trigger module, and an entropy weight calculation module; wherein:
[0013] The weight assignment trigger module determines the sampling period of the rewards and monitors the number of training rounds of the TD3 architecture. When the number of training rounds meets the sampling period, it triggers the reward matrix module;
[0014] The reward matrix module collects the rewards in the sampling period and forms a reward matrix for each optimization objective;
[0015] The calculation trigger module performs a fast Fourier transform on the reward matrix of each optimization objective to obtain the main frequency and the main frequency amplitude of the corresponding reward matrix; compares the differences between the main frequencies corresponding to each optimization objective and the differences between the main frequency amplitudes corresponding to each optimization objective, and triggers the entropy weight calculation module according to the comparison results;
[0016] The entropy weight calculation module calculates the reward weights corresponding to each optimization objective according to the reward matrix by using the entropy weight method.
[0017] Further, the reward matrix module collects the rewards corresponding to each optimization objective in the sampling period to form a data stream matrix, and performs standardization processing on the data stream matrix; normalizes the standardized data stream matrix column by column to obtain the reward matrix.
[0018] Further, in the entropy weight calculation module, the information entropy corresponding to each optimization objective is calculated by the following formula:
[0019] ;
[0020] Wherein, denotes the information entropy corresponding to the j-th optimization objective, denotes an element in the reward matrix;
[0021] According to the information entropy, the information utility value corresponding to each optimization objective is calculated by the following formula:
[0022] ;
[0023] where denotes the information utility value corresponding to the j-th optimization objective;
[0024] The reward weight corresponding to each optimization objective is calculated by the following formula:
[0025] ;
[0026] where denotes the reward weight corresponding to the j-th optimization objective.
[0027] Furthermore, the TD3 architecture in step S2 further includes an Actor network, an ActorT network, a Critic1 network, a Critic2 network, a CriticT1 network, and a CriticT2 network; where
[0028] The Actor network receives the current environmental state and applies the generated execution action to the environment to generate the next environmental state;
[0029] The ActorT network receives the next environmental state and generates a target action;
[0030] The CriticT1 network and the CriticT2 network simultaneously receive the target action and the next environmental state, and the outputs of the CriticT1 network and the CriticT2 network are combined with the current reward to obtain the target value;
[0031] The Critic1 network and the Critic2 network simultaneously receive and calculate the loss with respect to the target value according to the current environmental state.
[0032] Furthermore, the training process in step S2 includes:
[0033] S21: Initialize the TD3 architecture;
[0034] S22: In the current environmental state, the Actor network generates an action and controls the execution action in the suspension system to obtain the next environmental state, and the current reward is calculated by combining the corresponding rewards generated by the action in each optimization objective with the corresponding reward weights;
[0035] S23: Repeat step S22 multiple times, integrate the obtained actions, next environmental states, current environmental states, and current rewards into experience data, and store them in the experience pool;
[0036] S24: Randomly sample a batch of experience data from the experience pool, and based on the experience data, use the ActorT network to generate target actions; the CriticT1 network and the CriticT2 network use the Bellman expectation formula to generate target values based on the target actions;
[0037] S25: Calculate the loss functions for training the Critic1 network and the Critic2 network by combining the target values and target actions generated in step S24;
[0038] S26: Based on the loss functions obtained in step S25, use a delayed update strategy to update the parameters of the TD3 architecture;
[0039] S27: Determine whether the current training round satisfies the sampling period: if it does, use the entropy-weighted reward adjustment structure to adjust the reward weights corresponding to each optimization target, and execute step S28; if it does not, directly proceed to step S28;
[0040] S28: Repeat steps S21 - S27 with the updated TD3 architecture to complete the training of the TD3 architecture.
[0041] Furthermore, in step S24, the target value is obtained through the following formula:
[0042] ;
[0043] where represents the target value, represents the discount factor, represents the CriticT1 network, represents the CriticT2 network, represents the network parameters of the CriticT1 network, represents the network parameters of the CriticT2 network, represents the target action, represents the next environmental state, and r represents the current reward.
[0044] Furthermore, in step S25, the loss functions for training the Critic1 network and the Critic2 network are:
[0045] ;
[0046] ;
[0047] where Represents the loss function of the Critic1 network, Represents the loss function of the Critic2 network, Represents the action and the current environmental state The result obtained by inputting into the Critic1 network, Represents the network parameters of the Critic1 network, Represents the action and the current environmental state The result obtained by inputting into the Critic2 network, Represents the network parameters of the Critic2 network, and MSE represents the mean squared error loss function.
[0048] A readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps of the suspension system reward weighting method combined with the entropy weight method provided by the present invention.
[0049] An electronic device includes:
[0050] A memory for storing a computer program;
[0051] A processor for implementing the steps of the suspension system reward weighting method combined with the entropy weight method provided by the present invention when executing the computer program.
[0052] Compared with the prior art, the present invention can achieve the following beneficial effects:
[0053] (1) In the suspension system reward weighting method combined with the entropy weight method of the present invention, the entropy weight method can well collect the fluctuation information of each state quantity in the actual training process, and can naturally assign large weights to the targets with large fluctuations and unstable training, and assign small weights to those with small fluctuations and stable training, so as to achieve the purpose of balancing the convergence degree and convergence speed difference between conflicting targets.
[0054] (2) In the suspension system reward weighting method combined with the entropy weight method of the present invention, a data flow matrix is used to construct a reward matrix, providing a new perspective for data analysis in deep reinforcement learning;
[0055] (3) In the suspension system reward weighting method combined with the entropy weight method of the present invention, the data flow matrix and the entropy weight method are combined to capture the fluctuation information in the characteristics of reward data, realizing low-cost and high-efficiency reward weighting, which can significantly enhance the coordination between performance indicators such as comfort, stability, and response speed of the agent controlling the suspension system, making the algorithm easier to be transplanted into actual engineering scenarios, and at the same time can greatly improve the comprehensive control effect of the suspension system. Description of the Drawings
[0056] The accompanying drawings, which form a part of the present invention, are used to provide a further understanding of the present invention. The schematic embodiments and descriptions thereof of the present invention are used to explain the present invention and do not constitute an improper limitation on the present invention. In the drawings:
[0057] Figure 1 It is a schematic flow chart of the suspension system reward empowerment method combining the entropy weight method according to the embodiment of the present invention;
[0058] Figure 2 It is a schematic diagram of the TD3 architecture according to the embodiment of the present invention;
[0059] Figure 3 It is a schematic diagram of the entropy weight reward adjustment structure according to the embodiment of the present invention;
[0060] Figure 4 It is a schematic diagram of the electronic device according to the embodiment of the present invention.
[0061] Explanation of reference numerals:
[0062] 1, electronic device; 2, external device; 3, processing unit; 4, bus; 5, network adapter; 6, display; 7, (I / O) interface; 8, system memory; 9, random access memory; 10, cache memory; 11, storage system; 12, utility tool; 13, program module. Detailed implementation manners
[0063] In order to make the purpose, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and do not constitute a limitation on the present invention.
[0064] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.
[0065] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "center", "longitudinal", "lateral", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation to the present invention. In addition, the terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first", "second", etc. may explicitly or implicitly include one or more of such features. In the description of the present invention, unless otherwise specified, the meaning of "a plurality" is two or more.
[0066] In the description of the present invention, it should be noted that unless otherwise clearly specified and limited, the terms "mounted", "connected", "coupled" shall be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral connection; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood through specific circumstances.
[0067] The present invention will be described in detail below with reference to the drawings and in conjunction with embodiments.
[0068] As Figures 1 to 2 shown, the suspension system reward empowerment method combining the entropy weight method according to the embodiment of the present invention includes:
[0069] S1: Determine a plurality of optimization objectives for controlling the suspension system.
[0070] In one embodiment, the optimization objectives for controlling the suspension system are respectively the displacement of the sprung mass, the velocity of the sprung mass, the acceleration of the sprung mass; the displacement of the unsprung mass, the velocity of the unsprung mass, the acceleration of the unsprung mass.
[0071] S2: Construct a TD3 architecture with an entropy weight reward adjustment structure, and train the TD3 architecture according to the plurality of optimization objectives in step S1.
[0072] Specifically, the TD3 architecture includes an Actor network, an ActorT network, a Critic1 network, a Critic2 network, a CriticT1 network, and a CriticT2 network, as well as an entropy-weighted reward adjustment structure. The Actor network receives the current environmental state and applies the generated execution action to the environment to generate the next environmental state. The ActorT network receives the next environmental state and generates a target action. The CriticT1 network and the CriticT2 network simultaneously receive the target action and the next environmental state. The outputs of the CriticT1 network and the CriticT2 network are combined with the current reward to obtain a target value. The Critic1 network and the Critic2 network simultaneously receive and calculate the loss with respect to the target value based on the current environmental state. The entropy-weighted reward adjustment structure periodically collects the generated rewards during the training process and calculates the reward weights corresponding to each optimization objective using the entropy-weight method, thereby changing the rewards within the next collection period.
[0073] It can be understood that in the TD3 architecture provided by the present invention, the Actor network receives the current environmental state , and applies the generated execution action to the environment to generate the next environmental state . The ActorT network receives the next environmental state , and generates a target action . The CriticT1 network and the CriticT2 network simultaneously receive the target action and the next environmental state , and combine the output target evaluation value and with the current reward r to obtain the target value y, represents the network parameters of the CriticT1 network, represents the network parameters of the CriticT2 network. The Critic1 network and the Critic2 network simultaneously receive the current environmental state and the next environmental state , and the Critic1 network and the Critic2 network simultaneously receive and calculate the loss with respect to the target value y based on the current environmental state s. The entropy-weighted reward adjustment structure periodically collects the generated rewards during the training process and calculates the reward weights corresponding to each optimization objective using the entropy-weight method, thereby changing the rewards within the next collection period.
[0074] In a specific embodiment, the Actor network, Actor-T network, two CriticT networks, and two Critic networks generally adopt a structure of 3 to 4 layers. Among them, the structures of the Actor network and Actor-T network are the same, and the structures of the two CriticT networks and the two Critic networks correspond to each other, that is, the structure of the CriticT1 network is the same as that of the CriticT1 network, and the structure of the CriticT2 network is the same as that of the CriticT2 network. The input layer of the Actor network can use a node to receive the state. After passing through 2 to 3 hidden layers (for example, the number of nodes is 300 and 200, and the ReLU activation function is used), the output layer generates an action. Since the generated action needs to be restricted within the range of the action space (for example, [−1,1]), the tanh activation function is usually used. If the range of the action space is larger, the output value can be extended to the target range through a linear transformation. The input layers of the two Critic networks both receive the concatenation result of the current environmental state and the action. The input layer should have two nodes. The hidden layer settings are similar to those of the Actor network (also using the ReLU activation function), and the output layer is a linear output for calculating the evaluation value of the generated action. The soft-update TD3 model is adopted to balance performance and computational cost while avoiding the problems of gradient disappearance and overfitting.
[0075] In some embodiments, the entropy-weighted reward adjustment structure includes a weight assignment trigger module, a reward matrix module, a calculation trigger module, and an entropy-weight calculation module. The weight assignment trigger module determines the sampling period of the reward and monitors the number of training rounds of the TD3 architecture. When the number of training rounds meets the sampling period, it triggers the reward matrix module; the reward matrix module collects the rewards in the sampling period and forms a reward matrix for each optimization objective; the calculation trigger module performs a fast Fourier transform on the reward matrix of each optimization objective to obtain the main frequency and main frequency amplitude of the corresponding reward matrix; compares the differences between the main frequencies corresponding to each optimization objective and the differences between the main frequency amplitudes corresponding to each optimization objective, and triggers the entropy-weight calculation module according to the comparison results; the entropy-weight calculation module calculates the reward weights corresponding to each optimization objective using the entropy-weight method based on the reward matrix. The sampling period is determined according to the actual situation. In a specific embodiment, when the total number of training rounds is 1000, the sampling period can be set to sample once every 100 steps. The reward matrix module collects the rewards corresponding to each optimization objective in the sampling period to form a data stream matrix, and performs normalization processing on the data stream matrix. The normalized data stream matrix is normalized column by column to obtain the reward matrix. In a specific embodiment, the trigger condition for triggering the entropy-weight calculation module according to the differences between the main frequencies and the differences between the main frequency amplitudes is determined according to the actual situation.
[0076] It should be noted that in the category of deep reinforcement learning algorithms, not only the TD3 algorithm has an environment, but all algorithms belonging to this category have an environment. Therefore, in order to improve the generality of the method and also propose a new perspective on data analysis in deep reinforcement learning, a data stream will be defined in the category of deep reinforcement learning algorithms. A data stream refers to a column of data constructed based on a certain logical relationship, which can be abstracted into a 1-row n-column matrix. The elements in each row are determined by the same logical relationship, and the data structure of its elements is not limited. A data stream matrix is an m×n matrix, where each row must be a separate data stream, and the data structure of the elements is not limited. For example, in the multi-objective optimization problem of deep reinforcement learning, m optimization objectives can be used as row indicators, and the n values generated by each state quantity to be optimized in the process of interacting with the environment in the causal logical order are arranged as row vectors according to the row indicators to form a data stream matrix. Corresponding to the reward matrix module of the present invention, the n time-series rewards of each of the m optimization objectives in the sampling period are collected to form an m×n data stream matrix. This step can be completed using popular programming software such as python and matlab.
[0077] In addition, the fluctuations in the data will cause fluctuations in the loss function during the training process, making the gradient update unstable. High-fluctuation training data may lead to large changes in the gradient during the optimization process, thus slowing down the convergence speed of the model. In the multi-objective optimization algorithm based on TD3, the fluctuations of the rewards of each target to be optimized are generally concerned. When the amplitude fluctuations between the rewards corresponding to each target to be optimized are large and the fluctuation frequencies are too different, the network will have inconsistent convergence effects on different target quantities, and even the convergence degrees and convergence speeds will differ greatly. At this time, it is said that the multi-objective convergence performance is poor. Therefore, the calculation trigger module designed in the entropy weight reward adjustment structure of the present invention performs a fast Fourier transform on the reward matrices of different optimization targets to obtain the main frequency and main frequency amplitude of their reward data streams, and preliminarily compares the differences in the main frequencies and amplitudes of each item.
[0078] The Entropy Weight Method is an objective method for determining index weights in multi-index decision-making problems. It uses the concept of information entropy to reflect the degree of difference between indexes, thereby giving the weights of each index. The core idea of the Entropy Weight Method is that the more dispersed the values of an index are among different samples, the greater the amount of information contained in the index and the higher the weight; on the contrary, if the values of the index do not vary much among samples, its weight will be lower because its discrimination ability for the evaluation result is weaker. From the above concept of the Entropy Weight Method, it can be seen that the role of the Entropy Weight Method is to evaluate the distribution of a set of data. The data stream with large fluctuations has a high weight, and the stable data stream has a low weight. For a certain optimization goal, the more dispersed the distribution of its reward data stream, the lower the training efficiency, and the more stable the reward distribution, the higher the training efficiency. To correct the difference in the target convergence effect caused by this difference in data fluctuations through weight setting, the Entropy Weight Method is just right. Based on this, the present invention combines the Entropy Weight Method to determine the reward weight.
[0079] In a specific embodiment, in the entropy calculation module, the information entropy corresponding to each optimization goal is calculated by the following formula:
[0080] ;
[0081] where represents the information entropy corresponding to the jth optimization goal, represents the element in the reward matrix;
[0082] According to the information entropy, the information utility value corresponding to each optimization goal is calculated by the following formula:
[0083] ;
[0084] where represents the information utility value corresponding to the jth optimization goal;
[0085] The reward weight corresponding to each optimization goal is calculated by the following formula:
[0086] ;
[0087] where represents the reward weight corresponding to the jth optimization goal. At this time, the sum of the reward weights of the optimization goals is 1.
[0088] After determining the reward weights corresponding to each optimization goal, the reward in the next sampling period is determined by the following formula:
[0089] ;
[0090] where represents the current reward of the jth optimization, Represents the previous reward of the j-th optimization.
[0091] In some embodiments, the training process in step S2 includes:
[0092] S21: Initialize the TD3 architecture. In a specific embodiment, the following parameters are randomly initialized using the Xavier initialization method. The initialization includes:
[0093] Initialize the environmental state; the state of the suspension is the state corresponding to the optimization target, that is, including the displacement, velocity, and acceleration of the sprung mass, and the displacement, velocity, and acceleration of the unsprung mass;
[0094] Randomly initialize the Critic1 network and the Critic2 network. Specifically, randomly initialize the network parameters of the Critic1 network and the network parameters of the Critic2 network for random initialization;
[0095] Initialize the two target Critic networks, namely the CriticT1 network and the CriticT2 network. Specifically, initialize the network parameters of the CriticT1 network and the network parameters of the CriticT2 network for initialization. The initialized network parameters are equal to the network parameters , and the initialized network parameters are equal to the network parameters ;
[0096] Randomly initialize the Actor network. Specifically, randomly initialize the network parameters of the Actor network for random initialization;
[0097] Initialize the target Actor network, namely the ActorT network. Specifically, initialize the network parameters of the ActorT network for initialization. The initialized network parameters are equal to the network parameters ;
[0098] Initialize the hyperparameters for updating the TD3 architecture. The hyperparameters include the discount factor , the soft update frequency , the policy update frequency policy_delay, and the learning rates of the Critic1 network and the Critic2 network. In a certain embodiment, the learning rate of the Actor network is usually set to 10 -4 , and the learning rates of the two Critic networks are both 10 -3 , and the soft update frequency Set in [0.005, 0.01].
[0099] S22: In the current environmental state, the Actor network generates an action and controls the execution of the action in the suspension system to obtain the next environmental state, and calculates the current reward by combining the corresponding rewards generated by the action for each optimization objective with the corresponding reward weights.
[0100] In some embodiments, in the current environmental state s, the Actor network generates an action added with exploration noise , that is:
[0101] ;
[0102] wherein, represents the Actor network, and the exploration noise is Gaussian noise.
[0103] Controls the execution of the action in the suspension system , obtains the next environmental state, and calculates the current reward r by combining the corresponding rewards generated by the action for the optimization objective with the corresponding reward weights, that is:
[0104] .
[0105] S23: Repeat step S22 multiple times, and integrate the obtained actions, next environmental states, current environmental states, and current rewards into experience data and store them in the experience pool.
[0106] It can be understood that by repeating step S22 multiple times, the obtained actions , next environmental states , current environmental states and the total reward value r are integrated into experience data and stored in the experience pool.
[0107] S24: Randomly sample a batch of experience data from the experience pool, and use the ActorT network to generate target actions according to the experience data; the CriticT1 network and the CriticT2 network use the Bellman expectation formula to generate target values according to the target actions.
[0108] In some embodiments, randomly sample a batch of experience data from the experience pool , and according to the experience data , use the ActorT network to generate target actions , that is:
[0109] ;
[0110] wherein, denotes the ActorT network, denotes the noise, and Gaussian noise is adopted in a certain embodiment.
[0111] The target value is obtained through the Bellman expectation formula:
[0112] ;
[0113] where, denotes the target value, denotes the discount factor, denotes the CriticT1 network, denotes the CriticT2 network.
[0114] S25: Combine the target value and the target action generated in step S24 to calculate the loss functions for training the Critic1 network and the Critic2 network.
[0115] In a certain embodiment, the loss functions for training the Critic1 network and the Critic2 network are:
[0116] ;
[0117] ;
[0118] where, denotes the loss function of the Critic1 network, denotes the loss function of the Critic2 network, denotes the result obtained by inputting the action and the current environmental state into the Critic1 network, denotes the result obtained by inputting the action and the current environmental state into the Critic2 network, and MSE denotes the mean squared error loss function.
[0119] S26: According to the loss functions obtained in step S25, use the delayed update strategy to update the parameters of the TD3 architecture.
[0120] During each round of training, update the parameters and the parameter through the following formula:
[0121] ;
[0122] ;
[0123] where, denotes the derivative of the loss function with respect to Compute the gradient with respect to the loss function in Compute the gradient
[0124] Update the parameters and parameter respectively according to parameter and parameter :
[0125] ;
[0126] ;
[0127] Every policy_delay training rounds, update the parameters and parameter , including:
[0128] Compute the policy loss through the following formula :
[0129] ;
[0130] Update the parameter according to the policy loss through the following formula:
[0131] ;
[0132] where denotes the gradient with respect to the policy loss in Compute the gradient;
[0133] Then update the parameter through the following formula:
[0134] ;
[0135] Update the current environmental state to the next environmental state .
[0136] S27: Determine whether the current training round satisfies the sampling period: If it does, adjust the reward weights corresponding to each optimization objective using entropy-weighted reward adjustment and execute step S28; if not, directly proceed to step S28
[0137] S28: Repeat steps S21 - S27 with the updated TD3 architecture to complete the training of the TD3 architecture
[0138] S3: Input the current environmental state that meets the optimization objective of step S1 into the TD3 model obtained in step S2 to predict the output action of the suspension system.
[0139] Correspondingly, according to an embodiment of the present invention, the present invention further provides an electronic device, a readable storage medium, and a computer program product.
[0140] Figure 4 It is a schematic structural diagram of an electronic device 1 provided in an embodiment of the present invention. Figure 4 It shows a block diagram of an exemplary electronic device 1 suitable for implementing the embodiments of the present invention. Figure 4 The shown electronic device 1 is only an example and should not bring any limitation to the functions and usage scope of the embodiments of the present invention.
[0141] As Figure 4 shown, the electronic device 1 is presented in the form of a general-purpose computing device. The electronic device 1 is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are only examples and are not intended to limit the implementation of the present invention described and / or claimed herein.
[0142] The components of the electronic device 1 may include but are not limited to: one or more processors or processing units 3, a system memory 8, and a bus 4 connecting different system components (including the system memory 8 and the processing unit 3).
[0143] The bus 4 represents one or more of several types of bus structures, including a memory bus or a memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any of the multiple bus structures. For example, these architectures include but are not limited to the Industry Standard Architecture (ISA) bus, the Micro Channel Architecture (MAC) bus, the Enhanced ISA bus, the Video Electronics Standards Association (VESA) local bus, and the Peripheral Component Interconnect (PCI) bus.
[0144] The electronic device 1 typically includes a variety of computer system-readable media. These media can be any available media accessible by the electronic device 1, including volatile and non-volatile media, removable and non-removable media.
[0145] System memory 8 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 9 and / or cache memory 10. The electronic device 1 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, the storage system 11 may be used for reading and writing on a non-removable, non-volatile magnetic medium ( Figure 4 not shown, commonly referred to as a "hard disk drive"). Although Figure 4 not shown in the figure, a disk drive for reading and writing on a removable non-volatile disk (such as a "floppy disk") and an optical disk drive for reading and writing on a removable non-volatile optical disk (such as a CD-ROM, DVD-ROM or other optical medium) may be provided. In these cases, each drive may be connected to the bus 4 through one or more data medium interfaces. The system memory 8 may include at least one program product having a set (such as at least one) of program modules configured to perform the functions of the embodiments of the present invention.
[0146] A program / utility 12 having a set (at least one) of program modules 13 may be stored, for example, in the system memory 8. Such program modules 13 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment. The program modules 13 generally perform the functions and / or methods in the embodiments described in the present invention.
[0147] The electronic device 1 may also communicate with one or more external devices 2 (such as a keyboard, a pointing device, a display 6, etc.), and may also communicate with one or more devices that enable a user to interact with the electronic device 1, and / or communicate with any device that enables the electronic device 1 to communicate with one or more other computing devices (such as a network card, a modem, etc.). Such communication may be carried out through an input / output (I / O) interface 7. Moreover, the electronic device 1 may also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter 5. As Figure 4 shown, the network adapter 5 communicates with other modules of the electronic device 1 through the bus 4. It should be understood that although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device 1, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0148] The processing unit 3 executes various functional applications and data processing by running programs stored in the system memory 8, such as implementing the suspension system reward empowerment method combining the entropy weight method provided by the embodiments of the present invention.
[0149] In an embodiment of the present invention, a non-transitory computer-readable storage medium storing computer instructions is further provided, on which a computer program is stored. When the program is executed by a processor, it is the suspension system reward weight assignment method combining the entropy weight method provided by all the embodiments of the present application.
[0150] The computer storage medium of the embodiment of the present invention may adopt any combination of one or more computer-readable media. The computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. More specific examples (non-exhaustive list) of the computer-readable storage medium include: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program may be used by or in combination with an instruction execution system, apparatus, or device.
[0151] The computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries the computer-readable program code. Such a propagated data signal may take various forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above. The computer-readable signal medium may also be any computer-readable medium other than the computer-readable storage medium, and the computer-readable medium may send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device.
[0152] The program code contained on a computer-readable medium can be transmitted by any suitable medium, including but not limited to wireless, wire, optical fiber cable, RF, etc., or any suitable combination of the above. The computer program code for performing the operations of the present invention can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages - such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network including a local area network (LAN) or a wide area network (WAN), or, it can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).
[0153] An embodiment of the present invention also provides a computer program product, including a computer program, which when executed by a processor implements the suspension system reward empowerment method combined with the entropy weight method as described above.
[0154] It should be understood that various forms of the processes shown above can be used, reordering, adding or deleting steps. For example, the steps described in the disclosure of the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution disclosed in the present invention can be achieved, and no limitations are made herein.
[0155] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for awarding weights to a suspension system by combining the entropy weight method, characterized in that Including: S1: Determine multiple optimization objectives for controlling the suspension system; S2: Construct a TD3 architecture with an entropy-weighted reward adjustment structure and train the TD3 architecture according to the multiple optimization objectives in step S1; wherein, The entropy-weighted reward adjustment structure periodically collects the rewards generated during the training process and calculates the reward weights corresponding to each optimization objective using the entropy weight method; The entropy-weighted reward adjustment structure in step S2 includes a weight assignment trigger module, a reward matrix module, a calculation trigger module, and an entropy weight calculation module; wherein: The weight assignment trigger module determines the sampling period of the reward and monitors the number of training rounds of the TD3 architecture. When the number of training rounds meets the sampling period, it triggers the reward matrix module; The reward matrix module collects the rewards in the sampling period and forms a reward matrix for each optimization objective; The calculation trigger module performs a fast Fourier transform on the reward matrix of each optimization objective to obtain the main frequency and the main frequency amplitude of the corresponding reward matrix; compares the differences between the main frequencies corresponding to each optimization objective and the differences between the main frequency amplitudes corresponding to each optimization objective, and triggers the entropy weight calculation module according to the comparison results; The entropy weight calculation module calculates the reward weights corresponding to each optimization objective using the entropy weight method based on the reward matrix; The reward matrix module collects the rewards corresponding to each optimization objective in the sampling period to form a data stream matrix, and performs standardization processing on the data stream matrix; normalizes the standardized data stream matrix column by column to obtain the reward matrix; S3: Input the current environmental state that meets the optimization objectives in step S1 into the TD3 model obtained in step S2 to predict the output action of the suspension system.
2. The method for rewarding and assigning weights to a suspension system combining the entropy weight method according to claim 1, characterized in that In the entropy weight calculation module, the information entropy corresponding to each optimization objective is calculated by the following formula: ; Among them, represents the information entropy corresponding to the j-th optimization objective, represents the element in the said reward matrix; According to the information entropy, the information utility value corresponding to each optimization objective is calculated by the following formula: ; Among them, represents the information utility value corresponding to the j-th optimization objective; The reward weight corresponding to each optimization objective is calculated by the following formula: ; Among them, represents the reward weight corresponding to the j-th optimization objective.
3. The suspension system reward weight assignment method combining the entropy weight method according to claim 1, characterized in that The TD3 architecture in step S2 further includes an Actor network, an ActorT network, a Critic1 network, a Critic2 network, a CriticT1 network, and a CriticT2 network; wherein, The Actor network receives the current environmental state and applies the generated execution action to the environment to generate the next environmental state; The ActorT network receives the next environmental state and generates a target action; The CriticT1 network and the CriticT2 network simultaneously receive the target action and the next environmental state, and the outputs of the CriticT1 network and the CriticT2 network are combined with the current reward to obtain a target value; The Critic1 network and the Critic2 network simultaneously receive and calculate the loss with respect to the target value according to the current environmental state.
4. The suspension system reward weight assignment method combining the entropy weight method according to claim 3, characterized in that, The training process in step S2 includes: S21: Initialize the TD3 architecture; S22: In the current environmental state, the Actor network generates an action and controls the execution of the action in the suspension system to obtain the next environmental state, and calculates the current reward by combining the corresponding rewards generated by the action for each optimization objective with the corresponding reward weights; S23: Repeat step S22 multiple times, and integrate the obtained action, the next environmental state, the current environmental state, and the current reward into experience data and store them in the experience pool; S24: Randomly sample a batch of experience data from the experience pool, and use the ActorT network to generate a target action according to the experience data; the CriticT1 network and the CriticT2 network use the Bellman expectation formula to generate a target value according to the target action; S25: Calculate and train the loss functions of the Critic1 network and the Critic2 network by combining the target value and the target action generated in step S24; S26: According to the loss functions obtained in step S25, perform parameter updates on the TD3 architecture using a delayed update strategy; S27: Determine whether the current training round meets the sampling period: if it meets, use the entropy weight reward adjustment structure to adjust the reward weights corresponding to each optimization objective and execute step S28; if it does not meet, directly proceed to step S28; S28: Repeat steps S21 - S27 with the updated TD3 architecture to complete the training of the TD3 architecture.
5. The method for rewarding and weighting the suspension system combining the entropy weight method according to claim 3, wherein In step S24, the target value is obtained through the following formula: ; wherein, represents the target value, represents the discount factor, represents the CriticT1 network, represents the CriticT2 network, represents the network parameters of the CriticT1 network, represents the network parameters of the CriticT2 network, represents the target action, represents the next environmental state, and r represents the current reward.
6. The method for rewarding and weighting the suspension system combining the entropy weight method according to claim 3, characterized in that In step S25, the loss functions for training the Critic1 network and the Critic2 network are: ; ; Among them, represents the loss function of the Critic1 network, represents the loss function of the Critic2 network, represents the action and the current environmental state input into the Critic1 network to obtain the result, represents the network parameters of the Critic1 network, represents the action and the current environmental state input into the Critic2 network to obtain the result, represents the network parameters of the Critic2 network, and MSE represents the mean squared error loss function.
7. A readable storage medium, characterized in that, The readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps of the suspension system reward weight assignment method combined with the entropy weight method as described in any one of claims 1 to 6.
8. An electronic device, characterized in that, It includes: A memory for storing a computer program; A processor for implementing the steps of the suspension system reward weight assignment method combined with the entropy weight method as described in any one of claims 1 to 6 when executing the computer program.
Citation Information
Patent Citations
Method for controlling vertical vibration of electric vehicle driven by hub motor based on TD3 algorithm
CN118966033A
Method for optimizing PID control parameters of semi-active suspension of vehicle
WO2024125584A1