Improved PPO algorithm discrete parameter identification method

CN117744750BActive Publication Date: 2026-10-09SHANGHAI UNIVERSITY OF ELECTRIC POWER
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311608435.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-11-28
Publication Date
2026-10-09
Estimated Expiration
2043-11-28

AI Technical Summary

Technical Problem

但如果动作维度较大且每维动作个数较多时如:d为动作空间维度、n为每维动作空间动作个数,则输出层数节点个数为,则会出现神经网络的输出层和隐藏层维度爆炸的问题,会导致深度网络更新非常慢,甚至无法进行学习

Benefits of technology

[0053] This invention improves the PPO algorithm by pre-defining the parameter model to be identified and performing fine discrete partitioning in the action space. It is expected to provide more accurate parameter estimation, accelerate the parameter identification process, reduce the waste of computing resources, converge to appropriate parameter values ​​faster in a limited time, improve the training efficiency of the algorithm, and better adapt to and perform under different environmental conditions and problem settings; thus improving the overall performance of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117744750B_ABST
    Figure CN117744750B_ABST
Patent Text Reader

Abstract

The application belongs to the field of deep reinforcement learning and model parameter identification optimization, and discloses an improved PPO algorithm discrete parameter identification method, wherein the improved PPO algorithm comprises state, state transition strategy, action and reward; the method comprises the following steps: preset definition is performed on a to-be-identified parameter model based on the improved PPO algorithm, and state, action space and reward are defined respectively; the preset definition of the improved PPO algorithm is set in the action space range according to experience, and each parameter is discretely divided into multiple discrete values; the parameter identification process of the to-be-identified parameter model is stopped based on the fact that the maximum iteration number of the algorithm or the reward value meets the demand, so as to terminate the training of the Actor network parameter output value, sample the action value, and take the action value as the final identification parameter result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of deep reinforcement learning and model parameter identification and optimization, and more specifically, to an improved discrete parameter identification method for the PPO algorithm. Background Technology

[0002] Existing parameter identification methods mainly include traditional methods such as least squares method, maximum likelihood estimation, genetic algorithm, and particle swarm optimization. However, when faced with complex high-dimensional nonlinear, discontinuous objective functions and constraints, and when the problem cannot be expressed by a rigorous mathematical model, traditional algorithms often fail to obtain optimal results and may not be able to solve the problem when the amount of data is large.

[0003] Existing deep reinforcement learning algorithms (such as DQN and PPO) mainly target discrete action optimization problems with a single action dimension. If the agent's actions are in a multi-dimensional space, it is usually necessary to transform the multi-dimensional action space into a one-dimensional action space. However, if the action dimension is large and the number of actions per dimension is large (e.g., d is the action space dimension and n is the number of actions per dimension), then the number of output layers and nodes becomes... If this happens, the output and hidden layers of the neural network will explode, causing deep networks to update very slowly or even become unable to learn.

[0004] In view of this, the present invention provides an improved method for identifying discrete parameters of the PPO algorithm. Summary of the Invention

[0005] To overcome the problems in the existing technology, this invention proposes an improved discrete parameter identification method for the PPO algorithm. The purpose is to improve the shortcomings of traditional parameter identification methods in model parameter identification and existing PPO algorithms, and to better complete the parameter identification task of multi-discrete parameter models.

[0006] According to one aspect of the present invention, an improved method for identifying discrete parameters of a PPO algorithm is provided. The improved PPO algorithm includes a state, a state transition policy π, an action, and a reward; and includes the following steps:

[0007] Step S1: Based on the improved PPO algorithm, pre-define the parameter model to be identified, and define the state, action space, and reward respectively;

[0008] Step S2: Based on experience, set the action space range of the preset definition of the improved PPO algorithm, and discretize each parameter. A discrete value;

[0009] Step S3: Stop the parameter identification process of the model to be identified after the maximum number of iterations or the reward value meets the requirements. Obtain the action value action(d) by sampling the output value of the Actor network parameters obtained at the time of termination, and use the action value as the final identification parameter result.

[0010] In a preferred embodiment, the parameter model to be identified is a thermal network equivalent model, and the state is obtained based on the thermal network equivalent model. The state Including power loss and at least one thermal resistance parameter , , ;

[0011] ;

[0012] The action space Equivalent heat capacity parameter ;

[0013] ;

[0014] in: ; The number of corresponding equivalent heat capacity parameters corresponds to the number of discretized values ​​generated.

[0015] The reward : Calculate the mean square error between the measured temperature values ​​and the output temperature values ​​of the model of the parameter to be identified at all times, and use the reciprocal of the mean square error as the reward output;

[0016] ;

[0017] in The formula for mean squared error is... for Real-time measured temperature values; for The output temperature value of the parameter model to be identified at any time.

[0018] In a preferred embodiment, the improved PPO algorithm includes an Actor neural network layer and a Critic neural network layer. The data identified by the Actor neural network layer and the Critic neural network layer is sent to the equivalent model of the thermal network, and the equivalent model of the thermal network is updated and optimized.

[0019] In a preferred embodiment, the specific application logic of the Actor neural network layer is as follows:

[0020] Each equivalent heat capacity parameter is assigned an independent Actor neural network unit, Actor_net(d), based on the Actor neural network layer; where d is the number of equivalent heat capacity parameters.

[0021] Based on the number of actions for each equivalent heat capacity parameter, output the probability values ​​action(n)_logprob(d) and cross-entropy action(d)_entropy for each dimension of different equivalent heat capacity parameters.

[0022] In a preferred embodiment, the specific application logic of the Critic neural network layer is as follows:

[0023] Based on the output state value V (state) of the Critic neural network layer.

[0024] Training process for network parameters of Actor neural network layer and Critic neural network layer:

[0025] Initialize the parameters of the Actor neural network layer. After inputting the first set of states, output the probability values ​​of different discrete values ​​of the heat capacity parameter in each dimension, action(n)_logprob(d) and cross-entropy action(d_entropy).

[0026] The output probability value action(n)_logprob(d) is sampled to obtain the determined action value action(d);

[0027] The state and action value (d) are input into the equivalent hot network model, and the reward value is calculated. The next set of states is then output repeatedly.

[0028] Add [state, next_state, reward, action(d), action(d)_logprob] to the experience pool rpm to optimize the equivalent heat network model and output the equivalent heat capacity parameters;

[0029] The loss function is calculated using reward generalization advantage estimation (GAE) and the parameters in the Actor neural network layer and the Critic neural network layer are updated.

[0030] Repeat the above steps until the expected reward value is reached, then stop training.

[0031] In a preferred embodiment, the logic for updating the network loss function and parameters of the Actor neural network layer is as follows:

[0032] The generalization advantage estimate gae is used to avoid the problems of non-negative reward values ​​and insufficient representation of reward throughout the process.

[0033] The new action probability action(d)_logprobNew is obtained online through the Actor neural network layer, and the ratio(d) is obtained by comparing it with the action probability action(d)_logprobOld in the experience pool.

[0034] ;

[0035] And the ratioClip(d) is obtained by clipping with a parameter of 0.2;

[0036] ;

[0037] in: For the clipping function, Will Limited to the range [1-0.2, 1+0.2];

[0038] Surr1 and Surr2 are calculated using the following formulas.

[0039] ;

[0040] ;

[0041] in: The generalization advantage value without pruning; The generalization advantage value after clipping;

[0042] The average cross-entropy of the output values ​​of the Actor neural network layers is calculated, and combined with the surr value, the common loss value of the n Actor neural network layers is obtained. Then, the parameters of Actor1-Actor5 are updated respectively.

[0043] .

[0044] In a preferred embodiment, the logic for updating the Critic network loss function and parameters is as follows:

[0045] Based on the GAE value corresponding to the extracted experience pool data and the Critic network output value, the target state value Vtaget is obtained as follows:

[0046] ;

[0047] The loss value of the Critic network is calculated and updated by using the target state value Vtaget and the output value of the Critic network.

[0048] .

[0049] In a preferred embodiment, the parameter model to be identified is an inverter IGBT thermal network model. The inverter IGBT generates heat due to its own losses, which is ultimately dissipated through the cooling water channel. The heat dissipation path between the inverter IGBT and the water channel can be represented by an equivalent thermal network, where R1-R5 are equivalent thermal resistances and C1-C5 are equivalent thermal melting parameters. After applying the IGBT loss value to the inverter IGBT thermal network model for a certain period of time, the IGBT temperature tends to stabilize due to the balance between heat generation and heat dissipation.

[0050] According to another aspect of the present invention, an electronic device includes: a processor and a memory, wherein the memory stores a computer program that can be called by the processor;

[0051] The processor executes the aforementioned improved PPO algorithm discrete parameter identification method by calling the computer program stored in the memory.

[0052] The technical effects and advantages of the improved discrete parameter identification method of the PPO algorithm of this invention are as follows:

[0053] This invention improves the PPO algorithm by pre-defining the parameter model to be identified and performing fine discrete partitioning in the action space. It is expected to provide more accurate parameter estimation, accelerate the parameter identification process, reduce the waste of computing resources, converge to appropriate parameter values ​​faster in a limited time, improve the training efficiency of the algorithm, and better adapt to and perform under different environmental conditions and problem settings; thus improving the overall performance of the system. Attached Figure Description

[0054] Figure 1 This is a schematic diagram of the electronic component connections in the IGBT thermal network model of the present invention;

[0055] Figure 2 This is a framework diagram of the improved PPO algorithm for IGBT thermal network parameter identification in this invention;

[0056] Figure 3 This is a flowchart of the IGBT thermal network model parameter identification method of the present invention;

[0057] Figure 4 This is a simplified diagram of the Actor neural network model of the present invention;

[0058] Figure 5 This is a simplified diagram of the Critic neural network model of the present invention. Detailed Implementation

[0059] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0060] The improvements to the existing deep reinforcement learning PPO algorithm are as follows:

[0061] (1) The algorithm independently designs Actor neural network models for different actions. It independently calculates the ratio of the probability of new and old network actions for each Actor neural network model, and calculates the average after clipping as part of the loss function.

[0062] (2) To increase the stability of the action, the average cross-entropy of the Actor neural network model output is used as another part of the loss function;

[0063] (3) Add the two parts together to get a unified loss value ActorLoss, and update all Actor neural network models. By restricting the actions of different dimensions, the collaboration between actions of different dimensions can be better promoted.

[0064] Example 1

[0065] Please see Figure 1-5 As shown in this embodiment, an improved PPO algorithm discrete parameter identification method is described. The improved PPO algorithm includes state, state transition policy π, action, and reward. The steps for identifying parameters of a hot network model using the improved PPO algorithm are as follows:

[0066] Step S1: Based on the improved PPO algorithm, pre-define the parameter model to be identified, and define the state, action space, and reward respectively;

[0067] The parameter model to be identified is a thermal network equivalent model, and the state is obtained based on the thermal network equivalent model. The state Including power loss and at least one thermal resistance parameter , , ;

[0068] ;

[0069] The action space Equivalent heat capacity parameter ;

[0070] ;

[0071] in: ; The number of corresponding equivalent heat capacity parameters corresponds to the number of discretized values ​​generated.

[0072] The reward : Calculate the mean square error between the measured temperature values ​​and the output temperature values ​​of the model of the parameter to be identified at all times, and use the reciprocal of the mean square error as the reward output;

[0073] ;

[0074] in The formula for mean squared error is... for Real-time measured temperature values; for The output temperature value of the parameter model is to be identified at any time.

[0075] Step S2: Based on experience, set the action space range of the preset definition of the improved PPO algorithm, and discretize each parameter. A discrete value;

[0076] Specifically, as can be seen from step S1, the parameter to be identified in the parameter model is the equivalent heat capacity parameter. The parameters to be identified are: One, for multi-order thermal resistance parameters Each thermal resistance parameter corresponds to a different frequency response, so each equivalent heat capacity parameter is different. Based on experience, the action space range of each heat capacity parameter is given first, and each parameter is discretized.

[0077] Step S3: Stop the parameter identification process of the model to be identified after the maximum number of iterations or the reward value meets the requirements. Obtain the action value action(d) by sampling the output value of the Actor network parameters obtained at the time of termination, and use the action value as the final identification parameter result.

[0078] Each parameter is equally spaced. In deep reinforcement learning, these discrete values ​​are equivalent to the agent having [a number of discrete values]. There are 1 action dimension, and each action dimension has 1 action dimension. Each action space corresponds to an improved PPO algorithm architecture, such as Figure 2 .

[0079] The improved PPO algorithm includes an Actor neural network layer and a Critic neural network layer;

[0080] The specific application logic of the Actor neural network layer is as follows:

[0081] Each equivalent heat capacity parameter is assigned an independent Actor neural network unit, Actor_net(d), based on the Actor neural network layer; a simplified diagram of the Actor neural network layer is shown below. Figure 4 As shown;

[0082] Where d is the number of equivalent heat capacity parameters, and n is the number of actions for each equivalent heat capacity parameter.

[0083] Output the probability values ​​action(n)_logprob(d) and cross-entropy action(d)_entropy for each dimension of different equivalent heat capacity parameters based on the number of actions for each equivalent heat capacity parameter;

[0084] The specific application logic of the Critic neural network layer is as follows:

[0085] Based on the output state value V (state) of the Critic neural network layer, a simplified diagram is shown below. Figure 5 As shown.

[0086] Training process for network parameters of Actor neural network layer and Critic neural network layer:

[0087] Initialize the parameters of the Actor neural network layer. After inputting the first set of states, output the probability values ​​of different discrete values ​​of the heat capacity parameter in each dimension, action(n)_logprob(d) and cross-entropy action(d_entropy).

[0088] The output probability value action(n)_logprob(d) is sampled to obtain the determined action value action(d);

[0089] The state and action value (d) are input into the equivalent hot network model, and the reward value is calculated. The next set of states is then output repeatedly.

[0090] Add [state, next_state, reward, action(d), action(d)_logprob] to the experience pool rpm to optimize the equivalent heat network model and output the equivalent heat capacity parameters;

[0091] The loss function is calculated using reward generalization advantage estimation (GAE) and the parameters in the Actor neural network layer and the Critic neural network layer are updated.

[0092] Repeat the above steps until the expected reward value is reached, then stop training.

[0093] The logic for the network loss function and parameter update of the Actor neural network layer is as follows:

[0094] The generalization advantage estimate gae is used to avoid the problems of non-negative reward values ​​and insufficient representation of reward throughout the process.

[0095] The new action probability action(d)_logprobNew is obtained online through the Actor neural network layer, and the ratio(d) is obtained by comparing it with the action probability action(d)_logprobOld in the experience pool.

[0096] ;

[0097] And the ratioClip(d) is obtained by clipping with a parameter of 0.2;

[0098] ;

[0099] in: For the clipping function, Will Limited to the range [1-0.2, 1+0.2];

[0100] Surr1 and Surr2 are calculated using the following formulas.

[0101] ;

[0102] ;

[0103] in: The generalization advantage value without pruning; The generalization advantage value after clipping;

[0104] The average cross-entropy of the output values ​​of the Actor neural network layers is calculated, and combined with the surr value, the common loss value of the n Actor neural network layers is obtained. Then, the parameters of Actor1-Actor5 are updated respectively.

[0105] .

[0106] The logic for the loss function and parameter updates of the Critic network is as follows:

[0107] Based on the GAE value corresponding to the extracted experience pool data and the Critic network output value, the target state value Vtaget is obtained as follows:

[0108] ;

[0109] The loss value of the Critic network is calculated and updated by using the target state value Vtaget and the output value of the Critic network.

[0110] .

[0111] Example 2

[0112] Further explanation of the parameter model to be identified: The parameter model to be identified includes the inverter IGBT thermal network model; taking the parameter identification of the inverter IGBT thermal network model as an example, the inverter IGBT generates heat due to its own losses, which is ultimately dissipated through the cooling water channel. The heat dissipation path between the inverter IGBT and the water channel can be represented by an equivalent thermal network, where R1-R5 are the equivalent thermal resistances and C1-C5 are the equivalent thermal parameters. After applying the IGBT loss value to the inverter IGBT thermal network model for a certain period of time, the IGBT temperature tends to stabilize due to the balance between heat generation and heat dissipation. Figure 1 .

[0113] An implementation scheme for parameter identification of IGBT thermal network models by improving the PPO algorithm is as follows: Figure 3 The steps are as follows:

[0114] Step S1: Conduct IGBT temperature tests by setting up an experiment to obtain the measured IGBT temperature value Ttest_t;

[0115] The specific preset definition method is as follows:

[0116] Multiple sets of data were tested, and the temperature change process of the inverter's IGBTs was divided into intervals of Δt, defining states, action spaces, and rewards:

[0117] Specifically, the preset time interval is Δt = 10ms.

[0118] The state : Obtain power loss at all times and thermal resistance parameters ;

[0119]

[0120] The action space Equivalent heat capacity parameter ;

[0121] ;

[0122] in: , , , and These represent the number of discretized values ​​corresponding to the equivalent heat capacity parameter;

[0123] The reward : Calculate the mean square error between the measured temperature value of IGBT and the output temperature value of the inverter IGBT thermal network model at all times, and use the reciprocal of the mean square error as the reward output;

[0124] ;

[0125] in The formula for mean squared error is... for Real-time measured temperature values ​​of the inverter's IGBTs; for The output temperature value of the inverter IGBT thermal network model at any time.

[0126] Step S2: The IGBT thermal network model performs iterative calculations of IGBT temperature based on IGBT loss value, equivalent heat capacity parameter, and equivalent thermal resistance parameter output by the improved PPO algorithm, to obtain the estimated IGBT temperature output value Tout.

[0127] The improved PPO algorithm includes an Actor neural network layer and a Critic neural network layer;

[0128] The specific application logic of the Actor neural network layer is as follows:

[0129] Each equivalent heat capacity parameter is assigned an independent Actor neural network unit, Actor_net(d), where d = 1, 2, 3, 4, 5. A simplified diagram of the Actor neural network layer is shown below. Figure 4 As shown;

[0130] Where d is the number of equivalent heat capacity parameters, and n is the number of actions for each equivalent heat capacity parameter.

[0131] Output the probability values ​​action(n)_logprob(d) and cross-entropy action(d)_entropy for each dimension of different equivalent heat capacity parameters based on the number of actions for each equivalent heat capacity parameter;

[0132] The specific application logic of the Critic neural network layer is as follows:

[0133] Based on the output state value V (state) of the Critic neural network layer, a simplified diagram is shown below. Figure 5 As shown.

[0134] Training process for network parameters of Actor neural network layer and Critic neural network layer:

[0135] Initialize the parameters of the Actor neural network layer. After inputting the first set of states, output the discrete probability values ​​of each dimension of heat capacity parameter, action(n)_logprob(d) and cross-entropy action(d_entropy).

[0136] The output probability value action(n)_logprob(d) is sampled to obtain the determined action value action(d) (heat capacity parameter). ).

[0137] The state and action(d) are input into the inverter IGBT thermal network model, and the reward value is calculated. The next state is output as next_state.

[0138] Add [state, next_state, reward, action(d), action(d)_logprob] to the experience pool rpm.

[0139] The loss function is calculated using reward generalization advantage estimation (GAE) and the parameters in the Actor neural network layer and the Critic neural network layer are updated.

[0140] Repeat the above steps until the expected reward value is reached, then stop training.

[0141] Step S3: Generate an IGBT temperature estimation curve based on the IGBT temperature estimation output value Tout; generate an IGBT temperature test curve value based on the measured IGBT temperature value Ttest_t; calculate the mean square error between the IGBT temperature estimation curve and the IGBT temperature test curve value, and use the reciprocal of the mean square error as the reward output.

[0142] Step S4: The improved PPO algorithm is optimized based on state and action reward feedback, and outputs equivalent heat capacity parameters.

[0143] Step S5: Repeat S1-S4 until the maximum number of iterations of the algorithm or the reward value meets the requirements, then stop parameter identification.

[0144] Specifically: An IGBT thermal network model is constructed based on IGBT temperature test experimental data. The IGBT thermal network model includes IGBT loss values ​​and at least one equivalent heat capacity parameter. The IGBT loss values ​​and equivalent heat capacity parameter are used to output an equivalent thermal resistance parameter through the PPO algorithm. Based on the adjustment of the above parameters, the IGBT temperature is iteratively calculated to obtain the IGBT temperature estimation output value Tout. In addition, the measured IGBT temperature value Ttest_t is obtained through this IGBT temperature test experiment. Using the measured IGBT temperature value Ttest_t as the prediction target, the IGBT temperature estimation output value Tout continuously approximates the measured IGBT temperature value Ttest_t through training; thus, the IGBT thermal network model is trained.

[0145] The above formulas are all dimensionless calculations. The formulas are derived from software simulations based on a large amount of collected data to obtain the most recent real-world results. The preset parameters and thresholds in the formulas are set by those skilled in the art according to the actual situation.

[0146] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired or wireless network. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. A semiconductor medium can be a solid-state drive.

[0147] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed in this invention can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.

[0148] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0149] In the several embodiments provided by this invention, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only one method, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0150] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0151] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0152] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

[0153] In conclusion, the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. An improved method for identifying discrete parameters in the PPO algorithm, characterized in that, The improved PPO algorithm includes state, state transition policy π, action, and reward; it includes the following steps: Step S1: Based on the improved PPO algorithm, pre-define the parameter model to be identified, defining the state, action space, and reward respectively; the parameter model to be identified is a heat network equivalent model, and the state is obtained based on the heat network equivalent model. The state Including power loss and at least one thermal resistance parameter , , ; ; The action space Equivalent heat capacity parameter ; ; in: ; The number of corresponding equivalent heat capacity parameters corresponds to the number of discretized values ​​generated. award : Calculate the mean square error between the measured temperature values ​​and the output temperature values ​​of the model of the parameter to be identified at all times, and use the reciprocal of the mean square error as the reward output; ; in The formula for mean squared error is... for Real-time measured temperature values; for The output temperature value of the parameter model to be identified at any time; Step S2: Based on experience, set the action space range of the preset definition of the improved PPO algorithm, and discretize each parameter. A discrete value; Step S3: Stop the parameter identification process of the model to be identified after the maximum number of iterations or the reward value meets the requirements. Obtain the action value action(d) by sampling the output value of the Actor network parameters obtained at the time of termination, and use the action value as the final identification parameter result. The improved PPO algorithm includes an Actor neural network layer and a Critic neural network layer. The data identified by the Actor neural network layer and the Critic neural network layer is sent to the equivalent model of the thermal network, and the equivalent model of the thermal network is updated and optimized. The specific application logic of the Actor neural network layer is as follows: Each equivalent heat capacity parameter is assigned an independent Actor neural network unit, Actor_net(d), based on the Actor neural network layer; where d is the number of equivalent heat capacity parameters. Based on the number of actions for each equivalent heat capacity parameter, output the probability values ​​action(n)_logprob(d) and cross-entropy action(d)_entropy for each dimension of different equivalent heat capacity parameters; The logic for the network loss function and parameter update of the Actor neural network layer is as follows: The generalization advantage estimate gae is used to avoid the problems of non-negative reward value and insufficient full-process reward representation; the new action probability action(d)_logprobNew is obtained online through the Actor neural network layer, and the ratio(d) is obtained by comparing it with the action probability action(d)_logprobOld in the experience pool. ; And the ratioClip(d) is obtained by clipping with a parameter of 0.2; ; in: For the clipping function, Will Limited to the range [1-0.2, 1+0.2]; Surr1 and Surr2 are calculated using the following formulas. ; ; in: The generalization advantage value without pruning; The generalization advantage value after clipping; The average cross-entropy of the output values ​​of the Actor neural network layers is calculated, and combined with the surr value, the common loss value of the n Actor neural network layers is obtained. Then, the parameters of Actor1-Actor5 are updated respectively. 。 2. The improved PPO algorithm discrete parameter identification method according to claim 1, characterized in that, The specific application logic of the Critic neural network layer is as follows: Based on the output state value V (state) of the Critic neural network layer. Training process for network parameters of Actor neural network layer and Critic neural network layer: Initialize the parameters of the Actor neural network layer. After inputting the first set of states, output the probability values ​​of different discrete values ​​of the heat capacity parameter in each dimension, action(n)_logprob(d) and cross-entropy action(d_entropy). The output probability value action(n)_logprob(d) is sampled to obtain the determined action value action(d); The state and action value (d) are input into the equivalent hot network model, and the reward value is calculated. The next set of states is then output repeatedly. Add [state, next_state, reward, action(d), action(d)_logprob] to the experience pool rpm to optimize the equivalent heat network model and output the equivalent heat capacity parameters; The loss function is calculated by estimating the generalization advantage through reward and updating the parameters in the Actor neural network layer and the Critic neural network layer. Repeat the above steps until the expected reward value is reached, then stop training.

3. The improved PPO algorithm discrete parameter identification method according to claim 2, characterized in that, The logic for the loss function and parameter updates of the Critic network is as follows: Based on the GAE value corresponding to the extracted experience pool data and the Critic network output value, the target state value Vtaget is obtained as follows: ; The loss value of the Critic network is calculated and updated by using the target state value Vtaget and the output value of the Critic network. 。 4. The improved PPO algorithm discrete parameter identification method according to claim 3, characterized in that, The parameter model to be identified is the inverter IGBT thermal network model. The inverter IGBT generates heat due to its own losses, which is eventually dissipated through the cooling water channel. The heat dissipation path between the inverter IGBT and the water channel is represented by an equivalent thermal network, where R1-R5 are the equivalent thermal resistance and C1-C5 are the equivalent thermal melting parameters. After applying the IGBT loss value to the inverter IGBT thermal network model for a certain period of time, the IGBT temperature tends to stabilize due to the balance between heat generation and heat dissipation.

5. An electronic device, characterized in that, include: A processor and a memory, wherein the memory stores a computer program that can be called by the processor; The processor executes an improved PPO algorithm discrete parameter identification method according to any one of claims 1 to 4 by calling a computer program stored in the memory.