Power interaction model training method and device, power interaction method and device, equipment and medium

By constructing a model structure that integrates policy networks and value networks, and combining state data and value feedback at different times, the problem of low model training efficiency in power systems is solved, and efficient power resource interaction operation decision-making is achieved.

CN122020180APending Publication Date: 2026-05-12XIAN JIAOTONG LIVERPOOL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
XIAN JIAOTONG LIVERPOOL UNIV
Filing Date
2026-02-10
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing technologies employ centralized model training methods in power systems, resulting in high computational complexity, low model training efficiency, and an inability to output timely interactive operations between various entities.

Method used

We construct a model structure that integrates policy networks and value networks. By combining the state data, interactive operations, and value feedback of unit systems at different time points, we reduce the dimensionality of state data and improve model training efficiency through value difference calculation and parameter adjustment.

Benefits of technology

By reducing the dimensionality of state data, the problem of high computational complexity in traditional model training is solved, thereby improving model training efficiency and decision-making efficiency in power resource interaction operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122020180A_ABST
    Figure CN122020180A_ABST
Patent Text Reader

Abstract

The invention discloses an electric power resource interactive processing scene model training method and device, an electric power resource interactive processing method and device, equipment and a medium. The method comprises the following steps: applying unit systems in the electric power resource interaction system, establishing a training model for each unit system in the electric power resource interaction system, constructing a model structure in which a strategy network and a value network are coordinated, and performing value difference calculation and parameter adjustment by combining state data, interaction operation and value feedback of the unit systems at different moments. Therefore, the interactive operation decision of the unit system in the power resource interactive system is based on the own state data. According to the embodiment of the invention, the model training efficiency and the decision-making efficiency of power resource interaction operation can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of Internet technology, and in particular to a training method, apparatus, device, and medium for a scenario model of power resource interaction processing. Background Technology

[0002] The current power system is facing trends such as a high proportion of renewable energy integration, a more complex power structure, and rapid growth in distributed energy resources. These trends are driving the power market to evolve from a centralized dispatch model to a distributed model with multi-stakeholder cooperation and interaction.

[0003] Power resource interaction refers to the interaction of power resources among multiple energy entities in a power system, so as to make full and efficient use of electrical energy and reduce power generation costs.

[0004] To address the challenge of minimizing power generation costs through effective interaction among multiple stakeholders, current technologies often employ centralized model training. This method treats the entire power system as the training object, using the combined state data of all stakeholders to construct a high-dimensional input space for end-to-end modeling. However, this approach suffers from high computational complexity and low training efficiency. Furthermore, it cannot provide timely feedback on the interactions between stakeholders during application. Summary of the Invention

[0006] This invention provides a power resource interaction processing scenario model training, power resource interaction processing method, device, equipment, and medium, which can improve model training efficiency and decision-making efficiency of power resource interaction operations.

[0007] According to one aspect of the present invention, an embodiment of the present invention provides a method for training a power resource interaction processing scenario model, the method comprising:

[0008] The state data of the unit system at the first moment is input into the policy network to obtain the interactive operation of the unit system at the first moment.

[0009] The state data and interaction operations of the power resource interaction system at the first moment are input into the first value network to obtain the initial value;

[0010] Based on the state data and interactive operations of the unit system at the second time moment, the initial value is adjusted to obtain the target value of the unit system at the first time moment; the second time moment precedes the first time moment.

[0011] The value difference between the compensation value of the unit system at the second time moment and the target value is calculated; the compensation value at the second time moment is obtained by the unit system through processing in the second value network based on the state data and interaction operation of the power resource interaction system at the second time moment.

[0012] Based on the value difference, adjust the parameters of the policy network and the value network corresponding to the unit system.

[0013] According to another aspect of the present invention, embodiments of the present invention also provide a training device for a power resource interaction processing scenario model, the device comprising:

[0014] The operation determination module is used to input the state data of the unit system at the first moment into the policy network to obtain the interactive operation of the unit system at the first moment.

[0015] The value assessment module is used to input the state data and interaction operations of the power resource interaction system at the first moment into the first value network to obtain the initial value;

[0016] The target value determination module is used to adjust the initial value based on the state data and interactive operations of the unit system at a second time moment to obtain the target value of the unit system at a first time moment; the second time moment precedes the first time moment.

[0017] The difference calculation module is used to calculate the value difference between the compensation value of the unit system at the second time and the target value; the compensation value at the second time is obtained by the unit system through processing the input into the second value network based on the state data and interaction operation of the power resource interaction system at the second time.

[0018] The parameter adjustment module is used to adjust the parameters of the strategy network and the value network corresponding to the unit system according to the value difference.

[0019] According to one aspect of the present invention, an embodiment of the present invention provides a power resource interaction processing method, the method comprising:

[0020] Get the current state data;

[0021] The current state data is input into a pre-trained policy network to obtain the current interaction operation; the policy network is trained using a power resource interaction processing scenario model training method.

[0022] According to another aspect of the present invention, embodiments of the present invention also provide a power resource interaction processing device, the device comprising:

[0023] The data acquisition module is used to acquire the current state data.

[0024] The operation output module is used to input the current state data into a pre-trained policy network to obtain the current interactive operation; the policy network is trained using a power resource interaction processing scenario model training method.

[0025] According to another aspect of the present invention, embodiments of the present invention also provide a power resource interaction processing scenario model training or a power resource interaction processing device, the power resource interaction processing scenario model training or power resource interaction processing device comprising:

[0026] At least one processor; and

[0027] A memory that is communicatively connected to at least one processor; wherein,

[0028] The memory stores a computer program that can be executed by at least one processor, such that the at least one processor is able to perform the power resource interaction processing scenario model training or the power resource interaction processing method according to any embodiment of the present invention.

[0029] According to another aspect of the present invention, a computer-readable storage medium is provided, which stores computer instructions for causing a processor to execute and implement the power resource interaction processing scenario model training or power resource interaction processing method of any embodiment of the present invention.

[0030] According to another aspect of the present invention, a computer program product is provided, the computer program product comprising a computer program, which, when executed by a processor, implements the power resource interaction processing scenario model training or power resource interaction processing method according to any embodiment of the present invention.

[0031] The technical solution of this invention establishes training models for each unit system in the power resource interaction system, constructs a model structure that coordinates the policy network and the value network, and calculates value differences and adjusts parameters by combining the state data, interactive operations and value feedback of the unit system at different times. This enables the interactive operation decisions of the unit system in the power resource interaction system to be based on its own state data, reduces the dimensionality of the state data, solves the problem of high computational complexity in traditional model training, and improves the efficiency of model training.

[0032] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0033] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0034] Figure 1A This is a flowchart of a power resource interaction processing scenario model training method provided by an embodiment of the present invention;

[0035] Figure 1B This is a flowchart of a power resource interaction processing scenario model training method provided by an embodiment of the present invention;

[0036] Figure 2A This is a flowchart of a power resource interaction processing method provided by an embodiment of the present invention;

[0037] Figure 2B This is a schematic diagram of power resource interaction operation of a power resource interaction system according to an embodiment of the present invention;

[0038] Figure 3 This is a structural diagram of a power resource interaction processing scenario model training device provided in an embodiment of the present invention;

[0039] Figure 4 This is a structural diagram of a power resource interaction processing device provided according to an embodiment of the present invention;

[0040] Figure 5 This is a schematic diagram of the structure of a power resource interaction processing scenario model training or power resource interaction processing device provided in the embodiments of the present invention. Detailed Implementation

[0041] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0042] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0043] The acquisition, storage, and application of state data involved in the technical solutions of this invention comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0044] Figure 1A This is a flowchart illustrating a method for training a scenario model for power resource interaction processing, provided by an embodiment of the present invention. This embodiment is applicable to the training of unit system interaction models in a power resource interaction system. The method can be executed by a power resource interaction processing scenario model training device, which can be implemented in hardware and / or software. This power resource interaction processing scenario model training device can be configured in a server.

[0045] See Figure 1A The training method for the power resource interaction processing scenario model shown includes:

[0046] S101. Input the state data of the unit system at the first moment into the policy network to obtain the interactive operation of the unit system at the first moment.

[0047] The power resource exchange system can be a system capable of exchanging power resources. A unit system can be an independent system participating in power resource exchange within the power resource exchange system. Unit systems include: power output systems or power storage systems. A power output system can be a system with power generation capacity. A power storage system can be a system with energy storage and release functions. At least one power output system and one power storage system exchange power resources within the power resource exchange system.

[0048] In one specific embodiment, the power output system includes: a residential power output system, a commercial power output system, and an industrial power output system. The energy storage system includes: an energy storage platform. The residential power output system can refer to a photovoltaic power generation system in a residential area. The commercial power output system can refer to a photovoltaic power generation system in a commercial area. The industrial power output system can refer to an industrial thermal power generation system. The residential power output system primarily generates electricity through photovoltaics to supply the daily electricity needs of residents within the area. The commercial power output system primarily generates electricity through photovoltaics to supply the daily electricity needs of businesses within the area. The industrial power output system primarily uses traditional thermal power generation to power the corresponding industrial equipment.

[0049] A power resource exchange system can be constructed based on residential, commercial, and industrial power output systems and energy storage platforms, allowing for the exchange of electricity among these systems. For example, when a residential or commercial power output system has surplus power, this surplus power can be transported to the industrial power output system for use in industrial equipment operation, or it can be transported to an energy storage platform for storage.

[0050] The first moment can be the start time of the current training epoch in model training. State data can be data describing the power state of the unit system. State data includes at least one of the following: load state, energy storage state, and power generation. The policy network can be a network that outputs interactive operations based on the input state. Interactive operations can be adjustments made by the unit system to power resources. Interactive operations include at least one of the following: output power, input power, or load adjustment. Load adjustment refers to the unit system's adjustment of shifting a portion of peak-hour load to off-peak hours to rationally schedule electricity consumption. Load adjustment can refer to the amount of load transferred. The output power interactive operation can be an operation that causes the unit system to output a portion of its power based on the unit system's power supply and demand and grid constraints. The input power interactive operation can be an operation that inputs a portion of its power to the unit system based on the unit system's power supply and demand and grid constraints.

[0051] S102. Input the state data and interaction operations of the power resource interaction system at the first moment into the first value network to obtain the initial value.

[0052] The value network can be a network used to evaluate the quality of interactive operations. The first value network can be a network that conservatively evaluates the quality of interactive operations. The first value network adjusts its parameters by a small margin. That is, in the first value network, the adjusted parameters change little compared to the original parameters. The parameters of the first value network are adjusted at a slow rate of change, so that the first value network will not overestimate a certain interactive operation due to factors such as noise, effectively suppressing the surge in value. In other words, the evaluation of the interactive operation given by the first value network is a conservative evaluation, and the value output by the first value network is the initial value. The initial value can be the value of the interactive operation evaluated by the first value network. The input of the first value network is the state data and the interactive operation at the first moment, and the output is the initial value. The value is the degree of impact of the interactive operation of the evaluation unit system on the overall resource utilization efficiency of the power resource interaction system.

[0053] S103. Based on the state data and interactive operations of the unit system at the second moment, the initial value is adjusted to obtain the target value of the unit system at the first moment; the second moment precedes the first moment.

[0054] The second time step can be the start time of the previous training epoch in model training. The second time step is earlier than the first time step. The target value can refer to the target value that the interaction operation should achieve in terms of its impact on the overall resource utilization efficiency of the power resource interaction system. The target value includes the initial value of the interaction operation at the first time step and the reward of the interaction operation at the second time step.

[0055] In an optional embodiment, adjusting the initial value based on the state data and interaction operations of the unit system at the second time moment to obtain the target value of the unit system at the first time moment includes: determining the resource exchange cost of the unit system at the second time moment based on the state data and interaction operations of the unit system at the second time moment; calculating the Shapley value of the power resource interaction system at the first time moment based on the interaction operations of the power resource interaction system at the first time moment; and adjusting the initial value based on the resource exchange cost and the Shapley value to obtain the target value of the unit system at the first time moment.

[0056] Resource exchange cost refers to the cost incurred by a unit system in exchanging electrical resources to meet its own electricity needs. Based on the power generation and load demand of each unit system in the status data, and the amount of electricity exchanged between unit systems during interactive operations, the resource exchange cost can be determined.

[0057] In one specific embodiment, the power output system includes an industrial power output system. In addition to fixed costs such as equipment costs and transportation costs, the industrial power output system also includes energy consumption costs.

[0058] Specifically, the power generation of an industrial power output system consists of two parts: power generation within the carbon emission limit and additional power generation obtained through the purchase of green certificates. To meet environmental protection requirements, the power generation of traditional thermal power generation is limited. That is, the upper limit of power generation from an industrial power output system is the sum of the power generation meeting the carbon emission limit for environmental protection requirements and the additional power generation obtained through the purchase of green certificates. The power generation of an industrial power output system can be expressed by the formula:

[0059] E1=E C +E REC

[0060] Among them, E C Electricity generation with carbon emissions. E REC The amount of electricity generated for green certificates.

[0061] The power generation cost of an industrial power output system is related to the amount of coal consumed, F. C The relationship with power generation is as follows:

[0062] F C =θ×E1

[0063] Where θ is the energy consumption coefficient of the thermal power unit.

[0064] Furthermore, the energy consumption cost C1 of the industrial power output system is calculated as follows:

[0065] C1=F C ×P C

[0066] Among them, P C Cost of coal.

[0067] In one specific embodiment, the tradable electricity volume is determined based on the power generation and load demand of each unit system in the status data. Then, the resource exchange cost is determined based on the tradable electricity volume, interactive operations, and the costs of each unit system.

[0068] Specifically, the power output system includes: residential power output system, commercial power output system, and industrial power output system. The tradable electricity volume is determined as follows:

[0069] The tradable electricity volume of an industrial power output system can be expressed as:

[0070] E tr1 =D1-E1

[0071] Among them, E tr1D1 represents the tradable electricity volume of the industrial power output system, D1 represents the load demand of the industrial power output system, and E1 represents the power generation of the industrial power output system. A positive tradable electricity volume indicates the amount of electricity that can be sold. A negative tradable electricity volume indicates the amount of electricity that should be purchased.

[0072] The tradable electricity volume of a residential power output system can be expressed as:

[0073] E tr2 =D2-E2

[0074] Among them, E tr2 D2 represents the tradable electricity volume of the residential power output system, D2 represents the load demand of the residential power output system, and E2 represents the power generation of the residential power output system.

[0075] The tradable electricity volume of a commercial power output system can be expressed as:

[0076] E tr3 =D3-E3

[0077] Among them, E tr3 D3 represents the tradable electricity volume of the commercial power output system, D3 represents the load demand of the commercial power output system, and E3 represents the power generation of the commercial power output system.

[0078] The Shapley value describes the contribution of each unit system in a power resource interaction system to the power resource interaction operation. In a power resource interaction system, multiple unit systems form a cooperative alliance through resource interaction, which typically saves costs compared to their individual operations. The Shapley value is used to measure the contribution of each unit system to cost reduction in the power resource interaction system. In a power resource interaction system, the Shapley value of each unit system is calculated based on the interaction operation of each unit system at the first moment. The Shapley value is used as a weight value and multiplied by the resource exchange cost of the corresponding unit system to obtain the reward for the interaction operation of that unit system at the second moment. Therefore, the mathematical expression for the target value can be:

[0079] Y=φ t ×R 成本 +Q 初始

[0080] Where Y is the target value, φ t It is the Shapley value, R 成本 Q represents the cost of resource exchange. 初始 This is the initial value.

[0081] It is evident that by calculating resource exchange costs and Shapley values ​​and adjusting the initial value, we can ensure that the rewards of a unit system are proportional to its contribution to the power resource interaction system. This incentivizes the unit system to optimize its interaction operations with the goal of reducing the cost of the power resource interaction system. As a result, the interaction operations not only reduce their own costs but also reduce the overall cost of the power resource interaction system, effectively improving the utilization rate of power resources.

[0082] In an optional embodiment, adjusting the initial value based on the resource exchange cost and the Shapley value to obtain the target value of the unit system at the first moment includes: determining the number of times the state data is accessed based on the state data of the unit system at the first moment and the state data during the historical training process; calculating the exploration reward of the unit system at the first moment based on the number of accesses; calculating the corrected value based on the exploration reward, the resource exchange cost, and the Shapley value; and adjusting the initial value based on the corrected value to obtain the target value of the unit system at the first moment.

[0083] The access count of state data can refer to the number of times a certain type of state data appears during model training. The access count is obtained by querying the state data from the historical training state data to determine the number of times the state data appeared at the first moment.

[0084] The exploration reward is calculated based on the number of visits, using the following formula:

[0085] R exp =η / (N+1)

[0086] Where N represents the number of times the state data is accessed at the first moment, and η is the exploration adjustment coefficient; R exp The exploration reward is designed to incentivize the model to prioritize exploring less common state scenarios. As the expression for exploration reward suggests, the more frequently a certain state appears during model training, the lower the exploration reward; conversely, the less frequently a certain state appears, the higher the exploration reward. Exploration rewards encourage the model to prioritize exploring less common state scenarios, effectively compensating for cognitive blind spots and improving the completeness and comprehensiveness of the scenarios covered by the model.

[0087] The revised value can be a value used to correct the initial value.

[0088] The expression for the modified value can be:

[0089] Q 修正 =φ t ×R 成本 +λ t ×R exp

[0090] Where, λ t λ is the dynamic decay factor.t Starting from the preset value, the value decreases exponentially with the number of training iterations.

[0091] Understandably, in the early stages of model training, the focus should be on scene coverage, while in the later stages, the emphasis should be on fine-tuning the interactive operations. Therefore, a dynamic decay factor λ is introduced. t This results in exploration rewards accounting for a higher proportion in the early stages of model training than in the later stages.

[0092] The decay law of the dynamic decay factor can be described as follows:

[0093] λ t= λ0exp(-βt)

[0094] Where λ0 is the initial value and β is the exponential decay rate.

[0095] Therefore, the expression for the target value can be:

[0096] Y=φ t ×R 成本 +λ t ×R exp +Q 初始

[0097] It is evident that by calculating exploration rewards based on the number of visits, the model can be incentivized to prioritize exploring uncommon state scenarios, effectively compensating for the model's cognitive blind spots and improving the completeness and comprehensiveness of the scenarios covered by the model.

[0098] S104. Calculate the value difference between the compensation value of the unit system at the second time moment and the target value; the compensation value at the second time moment is obtained by the unit system inputting the state data and interaction operation of the power resource interaction system at the second time moment into the second value network for processing.

[0099] The second value network can be a more aggressive network for evaluating the merits of interactive operations. The parameters of the second value network are adjusted more significantly than those of the first value network. When adjusting its parameters, the second value network directly corrects itself to approximate the target value. This adjustment method makes the second value network more sensitive to changes in state data and interactive operations, enabling it to actively respond to new changes and provide new value evaluations, thus improving the model's learning efficiency.

[0100] The compensation value can be considered as a compensation for the impact of interactive actions on the overall resource utilization efficiency of the power resource interaction system. The second time step is earlier than the first time step. During the training rounds of the first time step, the state data and interactive operations of the second time step are input into the second value network to reconstruct the value of the interactive operations at the second time step, thus obtaining the compensation value for the second time step. The second time step is a historical time step relative to the first time step, and the compensation value of the second time step can be approximately considered as an actual evaluation of the interactive operations at the second time step. The value of the second time step is retrospectively corrected based on the target value of the first time step to obtain the value difference. The value difference can be the difference between the compensation value and the target value.

[0101] Understandably, at the second time step, the model can only determine the value of the interaction based on the state data at that time. However, at the first time step, the model has a more complete global perspective. In the training epochs at the first time step, the second value network re-evaluates the value of the interaction behavior at the second time step, and the resulting compensation value is a correction result for the value determination at the second time step. That is, the compensation value can be approximately regarded as the true value of the interaction behavior at the second time step.

[0102] S105. Adjust the parameters of the strategy network and the value network corresponding to the unit system according to the value difference.

[0103] Specifically, the parameters of the second value network are adjusted with the optimization objective of reducing value discrepancies. The adjusted parameters of the second value network are then merged with those of the first value network to adjust the parameters of the first value network. Finally, the parameters of the strategy network are adjusted with the optimization objective of increasing compensation value.

[0104] Optionally, the parameters of the second value network and the parameters of the first value network can be fused using the following formula:

[0105] a1=τa2+(1-τ)a1

[0106] Where a1 is the parameter of the first value network, a2 is the parameter of the second value network, and τ is the adjustment coefficient.

[0107] In one specific embodiment, each unit system has an independent experience replay pool. The experience replay pool is used to store sample data from the unit system's training. Besides acquiring real-time data for model training, model training can also be performed by sampling data from the experience replay pool. During model training using data sampled from the experience replay pool, the parameters of the policy network are adjusted as follows:

[0108] Construct the following objective function:

[0109] L(π) = -E st~D ,a t~π {Qπ (s) t ,a t )-αlogπ(a t |s t )}

[0110] Where L(π) is the loss function. E st~D The state data is represented as s t Sampling from the experience replay pool D. Q π (s) t ,a t ) represents the state data as s t The interactive operation is a t The compensation value output by the second value network. π (a) t |s t ) represents the state data as s t At that time, take interactive operation a t The probability distribution.

[0111] The parameters of the policy network are adjusted with the goal of reducing the loss value of the loss function.

[0112] Optionally, the parameters of the policy network and value network of each unit system are periodically uploaded to the global server; the global server aggregates the parameters based on a weighted average method, outputs global aggregate parameters, and broadcasts them to each unit system; after receiving the global aggregate parameters, each unit system merges them with its current parameters and adjusts its own parameters.

[0113] The technical solution of this invention establishes training models for each unit system in the power resource interaction system, constructs a model structure that coordinates the policy network and the value network, and calculates value differences and adjusts parameters by combining the state data, interactive operations and value feedback of the unit system at different times. This enables the interactive operation decisions of the unit system in the power resource interaction system to be based on its own state data, reduces the dimensionality of the state data, solves the problem of high computational complexity in traditional model training, and improves the efficiency of model training.

[0114] In a specific embodiment, the training process of the power resource interaction processing scenario model is as follows: Figure 1B As shown.

[0115] First, the status data of each unit system is collected and preprocessed. The preprocessed status data is then assessed to determine if it is within a valid range; that is, whether it can be used as input to the policy network. If the preprocessed status data can be used as input to the policy network, it is input into the policy network to construct the state; if the preprocessed status data cannot be used as input to the policy network, the status data is corrected.

[0116] The unit system generates interaction strategies and outputs interaction operations based on the state of the constructed unit system. It then determines whether the interaction operations of the unit system are feasible. The interaction operations of a unit system cannot exceed its own physical limitations. For example, if the unit system is an industrial power output system, it is subject to physical constraints such as the structure and heat dissipation capacity of industrial equipment; load adjustment operations cannot exceed the industrial load limit. If the interaction operation is not feasible, the interaction strategy is regenerated based on the action constraints. If the interaction operation is feasible, the interaction operations of each unit system are simulated. It then determines whether supply and demand balance can be achieved after the interaction operation. If supply and demand balance can be achieved, the cost of each unit system is calculated. If supply and demand balance cannot be achieved, the amount of electricity stored in the power storage system is adjusted. Finally, it determines whether the energy storage state of the power storage system exceeds physical limitations after adjustment. If the energy storage state of the power storage system does not exceed physical limitations, the cost of each unit system is calculated; if the energy storage state of the power storage system exceeds physical limitations, the amount of electricity stored in the power storage system is readjusted.

[0117] Finally, the interaction operations and state data are input into the value network of each unit system, and the value network evaluates the interaction operations. Based on the calculated cost of the unit system and the value evaluation output by the value network, the parameters of the value network and policy network are adjusted.

[0118] Figure 2A This is a flowchart illustrating a power resource interaction processing method provided in an embodiment of the present invention. This embodiment is applicable to power resource interaction within a unit system of a power resource interaction system. The method can be executed by a power resource interaction processing device, which can be implemented in hardware and / or software. This power resource interaction processing device can be configured in a server.

[0119] It should be noted that for parts not described in detail in the embodiments of the present invention, please refer to the descriptions in other embodiments.

[0120] See Figure 2A The power resource interaction method shown includes:

[0121] S201. Obtain the current state data.

[0122] The system collects current-time status data through smart meters in each unit system. These smart meters are used to monitor real-time information such as energy consumption, power generation, and energy storage.

[0123] S202. Input the current state data into a pre-trained policy network to obtain the current interactive operation; the policy network is trained using a power resource interaction processing scenario model training method.

[0124] In this process, the state data of the unit system is input into the corresponding policy network to obtain the interaction operation of the unit system at the current moment.

[0125] In practical applications, the focus is primarily on the interaction of unit systems. After the interaction of unit systems, the power resource interaction system can determine the actual cost reduction based on the actual situation. Therefore, in application, it is only necessary to input the state data into the pre-trained policy network to obtain the interaction operations. There is no need to input the state data and interaction operations into the value network to determine the quality of the interaction operations. Optionally, in application, the quality of the interaction operations can be evaluated based on the actual cost reduction.

[0126] The method of inputting state data into the policy network to obtain interactive operations can reduce the computation process and output interactive operations in a timely manner.

[0127] In an optional embodiment, the step of inputting the current state data into a pre-trained policy network to obtain the current interaction operation includes: when the unit system is the power output system, inputting the load level, energy storage status, output and output cost of the power output system in the current time period into the pre-trained policy network to obtain the external grid interaction quantity, system interaction quantity and load adjustment quantity at the current time.

[0128] Here, load level can be the level of total electricity demand it can handle. State of Charge (SOC) can be the current state of charge of an energy storage unit. SOC can be expressed as the percentage of remaining available electricity relative to rated capacity. Output can be the active power output by the power output system. Output cost can be the cost incurred by the power output system in generating and outputting a unit of electrical energy.

[0129] External grid interaction can refer to the exchange of electrical energy between the power resource interaction system and the external public grid. Understandably, if the total power generation of the power output system has a surplus after meeting the load demands of each unit system, this surplus can be transmitted to the external grid. If the sum of the total power generation of the power output system and the storage capacity of the power storage system is insufficient to meet the total load demand of the power output system, then electricity needs to be input from the external grid. System interaction can refer to the exchange of electrical energy between the unit systems within the power resource interaction system. Load adjustment can refer to the load amount used to increase or decrease the original electricity demand.

[0130] It is evident that by customizing the input dimensions of the strategy network for the power output system and using core operating parameters such as load level, energy storage status, output, and output cost as the basis for decision-making, the external grid interaction, system interaction, and load adjustment quantities output by the strategy network can accurately match the real-time operating status of the power output system. This improves the scientificity and adaptability of the interactive operation decision-making of the power output system and ensures the efficient allocation of power resources inside and outside the system.

[0131] In an optional embodiment, inputting the current state data into a pre-trained policy network to obtain the current interaction operation includes:

[0132] When the unit system is the power energy storage system, the energy storage status and output cost of the power energy storage system in the current time period are input into the pre-trained policy network to obtain the system interaction quantity and output cost adjustment quantity at the current moment.

[0133] The output cost adjustment amount can be an adjustment amount that increases or decreases the cost benchmark value generated per unit of electrical energy output.

[0134] It is evident that by selecting energy storage status and output cost as the core input parameters of the strategy network based on the operating characteristics of the power energy storage system, the system interaction quantity and output cost adjustment quantity output by the strategy network can be matched with the charging and discharging capacity of the energy storage system, thereby improving the accuracy of the energy storage system's interactive decision-making.

[0135] The technical solution of this invention uses a strategy network execution unit system based on a pre-trained scenario model to make interactive operation decisions. This enables each unit system to output an appropriate interactive operation based on its current state data, reducing computational complexity, solving the problem of not being able to output interactive operations in a timely manner, and improving the decision-making efficiency of power resource interactive operations.

[0136] In one specific embodiment, the power resource interaction operation of the power resource interaction system can be as follows: Figure 2B As shown.

[0137] The power resource interaction system includes: residential power output system, commercial power output system, industrial power output system, and energy storage platform.

[0138] For residential power output systems, they can supply power to commercial power output systems, industrial power output systems, energy storage platforms, and external power grids, and can also receive power from these systems. Residential power output systems can also issue renewable energy certificates to industrial power output systems.

[0139] Commercial power generation systems are similar to residential power generation systems. Industrial power generation systems can receive renewable energy certificates from both commercial and residential power generation systems. These renewable energy certificates are used to increase the upper limit of power generation for industrial power generation systems.

[0140] Figure 3 This is a schematic diagram of a power resource interaction processing scenario model training device provided in an embodiment of the present invention. This embodiment of the present invention is applicable to the training of unit system interaction models in a power resource interaction system. The device can execute a power resource interaction processing scenario model training method and can be implemented in hardware and / or software.

[0141] See Figure 3 The power resource interaction processing scenario model training device shown includes:

[0142] The operation determination module 301 is used to input the state data of the unit system at the first moment into the policy network to obtain the interactive operation of the unit system at the first moment.

[0143] The value assessment module 302 is used to input the state data and interaction operations of the power resource interaction system at the first moment into the first value network to obtain the initial value;

[0144] The target value determination module 303 is used to adjust the initial value based on the state data and interactive operations of the unit system at the second time moment to obtain the target value of the unit system at the first time moment; the second time moment precedes the first time moment.

[0145] The difference calculation module 304 is used to calculate the value difference between the compensation value of the unit system at the second time and the target value; the compensation value at the second time is obtained by the unit system inputting the state data and interaction operation of the power resource interaction system at the second time into the second value network for processing;

[0146] The parameter adjustment module 305 is used to adjust the parameters of the strategy network and the value network corresponding to the unit system according to the value difference.

[0147] The technical solution of this invention establishes training models for each unit system in the power resource interaction system, constructs a model structure that coordinates the policy network and the value network, and calculates value differences and adjusts parameters by combining the state data, interactive operations and value feedback of the unit system at different times. This enables the interactive operation decisions of the unit system in the power resource interaction system to be based on its own state data, reduces the dimensionality of the state data, solves the problem of high computational complexity in traditional model training, and improves the efficiency of model training.

[0148] In an optional embodiment, the target value determination module 303 includes:

[0149] The resource exchange cost determination unit is used to determine the resource exchange cost of the unit system at the second time moment based on the state data and interactive operations of the unit system at the second time moment.

[0150] The Shapley value determination unit is used to calculate the Shapley value of the power resource interaction system at the first moment based on the interaction operation of the power resource interaction system at the first moment.

[0151] The target value determination unit is used to adjust the initial value based on the resource exchange cost and the Shapley value to obtain the target value of the unit system at the first moment.

[0152] In an optional embodiment, the target value determination unit includes:

[0153] The number of times the state data is accessed is determined based on the state data of the unit system at the first moment and the state data during the historical training process;

[0154] An exploration reward determination unit is used to calculate the exploration reward of the unit system at the first moment based on the number of visits;

[0155] The modified value determination unit is used to calculate the modified value based on the exploration reward, the resource exchange cost, and the Shapley value.

[0156] The target value determination unit is used to adjust the initial value according to the modified value to obtain the target value of the unit system at the first moment.

[0157] The power resource interaction processing scenario model training device provided in this embodiment of the invention can execute the power resource interaction processing scenario model training method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of executing the power resource interaction processing scenario model training method.

[0158] Figure 4 This is a schematic diagram of a power resource interaction processing device provided in an embodiment of the present invention. The embodiment of the present invention is applicable to the interaction processing of unit systems in a power resource interaction system. This device can execute power resource interaction processing methods and can be implemented in hardware and / or software.

[0159] See Figure 4 The power resource interaction processing device shown includes:

[0160] Data acquisition module 401 is used to acquire the status data at the current moment;

[0161] The operation output module 402 is used to input the current state data into a pre-trained policy network to obtain the current interactive operation; the policy network is trained by the power resource interaction processing scenario model training method.

[0162] The technical solution of this invention uses a strategy network execution unit system based on a pre-trained scenario model to make interactive operation decisions. This enables each unit system to output an appropriate interactive operation based on its current state data, reducing computational complexity, solving the problem of not being able to output interactive operations in a timely manner, and improving the decision-making efficiency of power resource interactive operations.

[0163] In an optional embodiment, the operation output module 402 includes:

[0164] The power output system unit is used to input the load level, energy storage status, output and output cost of the power output system in the current time period into a pre-trained strategy network when the unit system is the power output system, so as to obtain the external power grid interaction quantity, system interaction quantity and load adjustment quantity at the current time.

[0165] In an optional embodiment, the operation output module 402 includes:

[0166] The power energy storage system unit is used to input the energy storage status and output cost of the power energy storage system in the current time period into a pre-trained policy network when the unit system is the power energy storage system, so as to obtain the system interaction quantity and output cost adjustment quantity at the current moment.

[0167] The power resource interaction processing device provided in this embodiment of the invention can execute the power resource interaction processing method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects for executing the power resource interaction processing method.

[0168] Figure 5 A schematic diagram of the structure of a power resource interaction processing scenario model training or power resource interaction processing device 500, which can be used to implement embodiments of the present invention, is shown.

[0169] like Figure 5As shown, the power resource interaction processing scenario model training or power resource interaction processing device 500 includes at least one processor 501 and a memory, such as a read-only memory 502 or a random access memory 503, communicatively connected to the at least one processor 501. The memory stores computer programs executable by the at least one processor. The processor 501 can perform various appropriate actions and processes based on the computer program stored in the read-only memory 502 or loaded from the storage unit 508 into the random access memory 503. The random access memory 503 can also store various programs and data required for the operation of the power resource interaction processing scenario model training or power resource interaction processing device 500. The processor 501, read-only memory 502, and random access memory 503 are interconnected via a bus 504. An input / output interface 505 is also connected to the bus 504.

[0170] Multiple components in the power resource interaction processing scenario model training or power resource interaction processing device 500 are connected to the input / output interface 505, including: an input unit 506, such as a keyboard, mouse, etc.; an output unit 507, such as various types of displays, speakers, etc.; a storage unit 508, such as a disk, optical disk, etc.; and a communication unit 509, such as a network card, modem, wireless transceiver, etc. The communication unit 509 allows the power resource interaction processing scenario model training or power resource interaction processing device 500 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0171] Processor 501 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 501 include, but are not limited to, central processing units, graphics processing units, various special-purpose artificial intelligence computing chips, various processors running machine learning model algorithms, digital signal processors, and any suitable processor, controller, microcontroller, etc. Processor 501 performs the various methods and processes described above, such as power resource interaction processing scenario model training or power resource interaction processing methods.

[0172] In some embodiments, the power resource interaction processing scenario model training or power resource interaction processing method can be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed on the power resource interaction processing scenario model training or power resource interaction processing device 500 via read-only memory 502 and / or communication unit 509. When the computer program is loaded into random access memory 503 and executed by processor 501, one or more steps of the power resource interaction processing scenario model training method described above can be performed. Alternatively, in other embodiments, processor 501 can be configured to perform the power resource interaction processing scenario model training or power resource interaction processing method by any other suitable means (e.g., by means of firmware).

[0173] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays, application-specific integrated circuits (ASICs), application-specific standard products (ASICs), systems-on-a-chip (SoCs), complex programmable logic devices, computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0174] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0175] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory, read-only memory, erasable programmable read-only memory, flash memory, optical fiber, portable compact disk read-only memory, optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0176] To provide user interaction, the systems and techniques described herein can be implemented on an operational detection device. This power resource interaction processing scenario model training or power resource interaction processing device includes: a display device (e.g., a cathode ray tube or liquid crystal display monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the power resource interaction processing scenario model training or power resource interaction processing device. Other types of devices can also be used to provide user interaction; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0177] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0178] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a host product within the cloud computing service system. This addresses the shortcomings of traditional physical hosts and virtual private servers, such as high management difficulty and weak business scalability.

[0179] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0180] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A method for training a scenario model for interactive processing of power resources, characterized in that, A unit system applied in a power resource interaction system, the power resource interaction system comprising: at least one power output system and a power energy storage system, the unit system comprising: the power output system or the power energy storage system; the method comprising: The state data of the unit system at the first moment is input into the policy network to obtain the interactive operation of the unit system at the first moment. The state data and interaction operations of the power resource interaction system at the first moment are input into the first value network to obtain the initial value; Based on the state data and interactive operations of the unit system at the second time moment, the initial value is adjusted to obtain the target value of the unit system at the first time moment; the second time moment precedes the first time moment. The value difference between the compensation value of the unit system at the second time moment and the target value is calculated; the compensation value at the second time moment is obtained by the unit system through processing the input into the second value network based on the state data and interaction operation of the power resource interaction system at the second time moment. Based on the value difference, adjust the parameters of the policy network and the value network corresponding to the unit system.

2. The method according to claim 1, characterized in that, The step of adjusting the initial value based on the state data and interaction operations of the unit system at the second time moment to obtain the target value of the unit system at the first time moment includes: Based on the state data and interactive operations of the unit system at the second time, determine the resource exchange cost of the unit system at the second time. Based on the interaction operation of the power resource interaction system at the first moment, calculate the Shapley value of the power resource interaction system at the first moment; The initial value is adjusted based on the resource exchange cost and the Shapley value to obtain the target value of the unit system at the first moment.

3. The method according to claim 2, characterized in that, The step of adjusting the initial value based on the resource exchange cost and the Shapley value to obtain the target value of the unit system at the first moment includes: The number of times the state data is accessed is determined based on the state data of the unit system at the first moment and the state data during the historical training process; Based on the number of visits, calculate the exploration reward of the unit system at the first moment; Calculate the adjusted value based on the exploration reward, the resource exchange cost, and the Shapley value; The initial value is adjusted based on the corrected value to obtain the target value of the unit system at the first moment.

4. A method for interactive processing of power resources, characterized in that, A unit system applied in a power resource interaction system, the power resource interaction system comprising: at least one power output system and a power energy storage system, the unit system comprising: the power output system or the power energy storage system; the method comprising: Get the current state data; The current state data is input into a pre-trained policy network to obtain the current interaction operation; the policy network is trained using the power resource interaction processing scenario model training method as described in any one of claims 1-4.

5. The method according to claim 4, characterized in that, The step of inputting the current state data into a pre-trained policy network to obtain the current interaction operation includes: When the unit system is the power output system, the load level, energy storage status, output and output cost of the power output system in the current time period are input into the pre-trained strategy network to obtain the external power grid interaction quantity, system interaction quantity and load adjustment quantity at the current moment.

6. The method according to claim 4, characterized in that, The step of inputting the current state data into a pre-trained policy network to obtain the current interaction operation includes: When the unit system is the power energy storage system, the energy storage status and output cost of the power energy storage system in the current time period are input into the pre-trained policy network to obtain the system interaction quantity and output cost adjustment quantity at the current moment.

7. A training device for a power resource interaction processing scenario model, characterized in that, A unit system applied in a power resource interaction system, the power resource interaction system comprising: at least one power output system and a power energy storage system, the unit system comprising: the power output system or the power energy storage system; the device comprising: The operation determination module is used to input the state data of the unit system at the first moment into the policy network to obtain the interactive operation of the unit system at the first moment. The value assessment module is used to input the state data and interaction operations of the power resource interaction system at the first moment into the first value network to obtain the initial value; The target value determination module is used to adjust the initial value based on the state data and interactive operations of the unit system at a second time moment to obtain the target value of the unit system at a first time moment; the second time moment precedes the first time moment. The difference calculation module is used to calculate the value difference between the compensation value of the unit system at the second time and the target value; the compensation value at the second time is obtained by the unit system through processing the input into the second value network based on the state data and interaction operation of the power resource interaction system at the second time. The parameter adjustment module is used to adjust the parameters of the strategy network and the value network corresponding to the unit system according to the value difference.

8. A power resource interaction processing device, characterized in that, A unit system applied in a power resource interaction system, the power resource interaction system comprising: at least one power output system and a power energy storage system, the unit system comprising: the power output system or the power energy storage system; the device comprising: The data acquisition module is used to acquire the current state data. An operation output module is used to input the current state data into a pre-trained policy network to obtain the current interactive operation; the policy network is trained by the power resource interaction processing scenario model training method as described in any one of claims 1-4.

9. A power resource interaction processing scenario model training or power resource interaction processing device, characterized in that, The power resource interaction processing scenario model training or power resource interaction processing device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the power resource interaction processing scenario model training or the power resource interaction processing method as described in any one of claims 1-6.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the power resource interaction processing scenario model training or power resource interaction processing method as described in any one of claims 1-6.