A generator set power dispatching method, device and system

By using a neural network-based generator power scheduling method, which leverages the collaborative work of master and slave neural networks, the problem of PID controllers struggling to approach target power in dynamic environments is solved. This method achieves rapid convergence and stable control, thereby improving the regulation efficiency and adaptive capability of generator sets.

CN121091695BActive Publication Date: 2026-01-30SHENZHEN HOPEWIND ELECTRIC CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511634594.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-10
Publication Date
2026-01-30
Estimated Expiration
2045-11-10

AI Technical Summary

Technical Problem

Existing PID controllers are difficult to adapt to complex and dynamically changing operating conditions in solar and wind power generation systems, making it difficult to bring the total power generation close to the target value and maintain convergence in a timely manner.

Method used

A generator power scheduling method based on neural networks is adopted. The main neural network attempts strategies within the same scheduling cycle, calculates scheduling state data and candidate action sequences, and uses an experience buffer to train and update the main neural network from the secondary neural network, thereby realizing real-time power scheduling of generators.

Benefits of technology

Under dynamic operating conditions, it can quickly converge to the target power and maintain stable control, significantly improving regulation efficiency and adaptive capability, and ensuring that the power dispatch of the generator set can approach the target value in a timely manner and maintain convergence in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121091695B_ABST
    Figure CN121091695B_ABST
Patent Text Reader

Abstract

This invention discloses a generator power dispatching method, apparatus, and system. The method includes using a main neural network to perform at least one strategy attempt within the same dispatching cycle, and calculating and reasoning based on the target power and real-time power to obtain a candidate action sequence; calculating the total score sequence of the main neural network within the same dispatching cycle using the candidate action sequence; recording the total score sequence and candidate action sequence in an experience buffer, and training a secondary neural network to synchronize with the main neural network. This invention generates candidate actions and calculates the total score through multiple inferences by the main neural network within the same dispatching cycle; writing the score and candidate actions into the experience buffer to train the secondary neural network and synchronize it with the main neural network. Within the same dispatching cycle, the total generating power can be promptly approximated to the target value and maintained convergence. Through the collaborative work of the main neural network and the secondary neural network, the regulation efficiency and adaptive capability are significantly improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of power systems, in particular to a generator set power scheduling method, device and system. BACKGROUND

[0002] PID controller is a classical control algorithm widely used in industrial control systems. The basic idea of PID controller is to adjust the output of the system as close as possible to the target value through three control parameters of proportion (P), integral (I) and derivative (D). Specifically, the proportion control adjusts the output directly through the current error, the integral control adjusts the output through the cumulative error, and the derivative control adjusts the output through the rate of change of error.

[0003] The existing solar (and wind) power scheduling mostly uses PID controller to close-loop adjust the power output of each unit to meet the target power issued by the total control room; however, PID control is based on fixed parameters, which is difficult to adapt to the nonlinear and time-varying characteristics caused by irradiance, wind speed, temperature and load fluctuations in power generation scenarios, resulting in lack of adaptive ability under complex and dynamic operating conditions, and it is difficult to approximate the target value and maintain convergence of the total power generation power in the same scheduling period. SUMMARY

[0004] The embodiments of the present application provide a generator set power scheduling method, device and system, aiming at solving the technical problem that the PID control algorithm in the prior art lacks the ability to approximate the target value and maintain convergence of the total power generation power in the same scheduling period of the power generation system.

[0005] In a first aspect, the embodiments of the present application provide a neural network-based generator set power scheduling method, comprising:

[0006] Receiving the target power issued by the total control room;

[0007] Using the main neural network to perform at least one strategy attempt in the same scheduling period, outputting a power scheduling instruction for controlling the generator set controller each time a strategy attempt is performed, and performing difference calculation and action reasoning according to the target power and the real-time power of the generator set, to obtain scheduling state data and candidate action sequence respectively;

[0008] When the scheduling state data meets the preset condition, the power scheduling in the current scheduling period is completed;

[0009] Using the candidate action sequence to calculate the total score sequence of the main neural network in the same scheduling period;

[0010] record the total score sequence and the candidate action sequence to an experience buffer, and train a slave neural network based on the experience buffer to synchronize to the master neural network, to obtain an updated master neural network.

[0011] In a second aspect, an embodiment of the present application provides a generator unit power scheduling device based on a neural network, comprising:

[0012] a data receiving unit configured to receive a target power issued by a master control room;

[0013] a network scheduling unit configured to perform at least one strategy attempt in a same scheduling period by using a master neural network, output a power scheduling instruction for controlling a generator unit controller each time a strategy attempt is performed, and perform difference calculation and action inference according to the target power and real-time power of the generator unit, to obtain scheduling state data and a candidate action sequence respectively;

[0014] a data judging unit configured to complete power scheduling in a current scheduling period when the scheduling state data meets a preset condition;

[0015] a sequence calculating unit configured to calculate the total score sequence of the master neural network in the same scheduling period by using the candidate action sequence;

[0016] a network updating unit configured to record the total score sequence and the candidate action sequence to an experience buffer, and train a slave neural network based on the experience buffer to synchronize to the master neural network, to obtain an updated master neural network.

[0017] In a third aspect, an embodiment of the present application provides a generator unit power scheduling system based on a neural network, comprising the generator unit power scheduling device based on a neural network, a master control module and a plurality of generator units according to the second aspect, the master control module sends a target instruction to the generator unit power scheduling device based on a neural network to perform data scheduling, and sends the target instruction to the corresponding generator units through the generator unit power scheduling device based on a neural network.

[0018] The embodiment of the present application provides a generator set power scheduling method based on a neural network, which comprises the following steps: receiving a target power issued by a central control room; performing at least one strategy attempt in the same scheduling period by using a master neural network, outputting a power scheduling instruction for controlling a generator set controller each time a strategy attempt is performed, and respectively obtaining scheduling state data and a candidate action sequence by performing difference calculation and action reasoning on the target power and real-time power of the generator set; completing power scheduling in the current scheduling period when the scheduling state data meets a preset condition; calculating a total score sequence of the master neural network in the same scheduling period by using the candidate action sequence; recording the total score sequence and the candidate action sequence to an experience buffer area, and training a slave neural network based on the experience buffer area to synchronize to the master neural network, so as to obtain an updated master neural network. The present application forms a scheduling state by using the difference between the target power of the central control room and the real-time power, and the master neural network reasons and generates a candidate action multiple times in the same scheduling period, so that the total power of the generator set can be approximated to the target value and maintained to converge in the same scheduling period. The total score of each reasoning is calculated, and the score and the candidate action are written to the experience buffer area to train the slave neural network and synchronize to the master neural network. Through the cooperative work of the master neural network and the slave neural network, the real-time power scheduling of the generator set is realized, the target power can be quickly converged and stably controlled under dynamic working conditions, and the adjustment efficiency and the adaptive ability are significantly improved.

[0019] The embodiment of the present application also provides a generator set power scheduling device and system based on a neural network, which also has the beneficial effects described above. BRIEF DESCRIPTION OF DRAWINGS

[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.

[0021] Figure 1 A flowchart of a generator set power scheduling method based on a neural network provided by the embodiment of the present application is shown in the figure.

[0022] Figure 2 Another flowchart of a generator set power scheduling method based on a neural network provided by the embodiment of the present application is shown in the figure.

[0023] Figure 3 A schematic block diagram of a generator set power scheduling device based on a neural network provided by the embodiment of the present application is shown in the figure.

[0024] Figure 4A schematic block diagram of a generator set power scheduling system based on a neural network is provided for an embodiment of the present application. DETAILED DESCRIPTION

[0025] The technical solutions in the embodiments of the present application will be clearly and completely described with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative work fall within the protection scope of the present application.

[0026] It should be understood that when used in the specification and the appended claims, the terms "comprise" and "include" indicate the presence of described features, integers, steps, operations, elements, and / or components, but do not exclude one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0027] It should also be understood that the terms used in the present application specification are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in the present application specification and the appended claims, the singular forms "a", "an" and "the" are intended to include the plural forms unless the context clearly indicates otherwise.

[0028] It should be further understood that the term "and / or" used in the present application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes these combinations.

[0029] Please see the following Figure 1 , Figure 1 A flowchart of a neural network-based generator set power scheduling method is provided for an embodiment of the present application, specifically comprising steps S101-S105.

[0030] S101, receiving a target power issued by a total control room;

[0031] S102, using a main neural network to perform at least one strategy attempt in the same scheduling period, outputting a power scheduling instruction for controlling a generator set controller each time a strategy attempt is performed, and performing difference calculation and action reasoning according to the target power and real-time power of the generator set, to obtain scheduling state data and a candidate action sequence, respectively;

[0032] S103, when the scheduling state data meets a preset condition, completing power scheduling in the current scheduling period;

[0033] S104, calculating a total score sequence of the main neural network in the same scheduling period by using the candidate action sequence;

[0034] S105, record the total score sequence and the candidate action sequence to an experience buffer, and train a slave neural network based on the experience buffer to synchronize to the master neural network, to obtain an updated master neural network.

[0035] In an embodiment, before the step S101, the method comprises:

[0036] obtaining a historical data set of the generator set, and training a model based on the historical data set to obtain a basic neural network;

[0037] performing deployment initialization on the basic neural network to obtain the master neural network and the slave neural network respectively.

[0038] In the embodiment, a historical data set of the generator set in a given operation interval is obtained, and the historical data set includes samples representing the power output of the generator set and the association relationship between the power output and the target instruction. A preset neural network model is trained offline based on the historical data set to obtain a basic neural network. The parameters of the basic neural network include weights representing the mapping relationship between the target power and the power set value of the generator set. The basic neural network is initialized in a control device with computing and communication capabilities. The master neural network and the slave neural network are instantiated in the same operation environment, and the initial parameters of the two are derived from the basic neural network to ensure that the initial reasoning and the subsequent learning process have consistent parameter basis.

[0039] In combination with Figure 2 As shown in FIG. 1, in step S101, the intermediate controller acts as a scheduling execution subject, receives the target power instruction issued by the central control room at the beginning of the current scheduling period, and establishes a time identifier corresponding to the instruction, which is used for subsequent association of real-time measurement data and strategy reasoning input. The target power is used to constrain the power distribution and instruction output of multiple generator sets in the current period.

[0040] In step S102, the intermediate controller calls the deployed master neural network to perform at least one strategy attempt in the same scheduling period. In each attempt, the target power and the collected real-time power of each generator set are calculated by difference to obtain scheduling state data depicting the current deviation, which is encoded and input into the master neural network for action reasoning to obtain the power set value of each generator set. The set value is packaged as a control instruction and used to drive the generator set controller to execute. At the same time, the system records the power distribution actions output by the master neural network in each attempt in sequence to obtain a candidate action sequence.

[0041] In an embodiment, the step S102 comprises:

[0042] performing time alignment on the target power and the real-time power to obtain a power sample set;

[0043] adding the power sample set by the index of the generator set one by one to obtain the actual available power of all generator sets;

[0044] performing absolute difference operation on the actual available power and the target power to obtain absolute error;

[0045] structuring and packaging the absolute error to obtain scheduling state data.

[0046] In the embodiment, the control device receives the target power issued by the central control room and the real-time power data of each generator set collected through the communication link in the current scheduling period, and pairs and time-aligns the two according to the time tag to construct a power sample set at the same time. After time alignment, the real-time power of each unit at the same time is added one by one according to the index of the generator set to obtain the actual available power of all generator sets at that time. Then, the absolute difference operation is performed on the actual available power and the target power to obtain the absolute error at that time. Finally, the target power, the actual available power, the absolute error, and the total number of associated devices and time tag information are structured and packaged to form scheduling state data and used for subsequent strategy reasoning and score calculation. For ease of measurement and reuse, the above actual available power and absolute error can be uniformly defined as follows:

[0047] wherein, is the target power issued by the central control room (the sum of the power that each unit should generate), n is the total number of devices, is the absolute error in a certain dispatching, is the power that all units can actually generate, is the real-time output power of the i-th unit.

[0048] In an embodiment, the step S102 further comprises:

[0049] performing relative deviation calculation on the target power and the actual available power to obtain a relative deviation parameter;

[0050] when the relative deviation parameter is greater than a preset threshold, triggering a gradient update algorithm to adjust the weight of the main neural network to obtain an updated main neural network;

[0051] performing reasoning calculation on the updated main neural network and the scheduling state data to obtain a power setting value set of each unit at the next time, and packaging the power setting value set as a broadcast instruction set;

[0052] issuing the broadcast instruction set and recalculating the relative deviation to obtain a new relative deviation parameter;

[0053] ​The comparison with the preset threshold is repeatedly performed using the new relative deviation parameter until the relative deviation parameter is less than or equal to the preset threshold, and a scheduling result is output.

[0054] In the embodiment, the control device first obtains a relative deviation parameter of the target power and the actual available power according to the aforementioned time alignment and summation result; when the relative deviation parameter is greater than a preset threshold (in the embodiment, the preset threshold can be 2%), a gradient update is triggered immediately to update the weights of the main neural network to obtain an updated main neural network. Subsequently, the updated main neural network is called to combine the current scheduling state data to perform inference, to generate a power setting value set of each unit at the next time, and the set is encapsulated as a broadcast instruction set and transmitted to the corresponding generator set through a communication link; after receiving the real-time power returned, the system recalculates the relative deviation parameter and compares it with the preset threshold, and if the relative deviation parameter is still out of limit, the aforementioned update-inference-issuing-retest cycle is continued until the relative deviation parameter is less than or equal to the preset threshold, and a scheduling result is output and the current scheduling period is ended.

[0055] For unified measurement and update, the relative deviation and the main neural network weight update in the embodiment can be calculated according to the following formula: ;

[0056] wherein, and respectively represent the main neural network weight vector at the current time and the last time, is a learning rate, is the gradient of the current loss to the weight.

[0057] In an embodiment, the step S102 further comprises:

[0058] performing feature encoding on the scheduling state data to obtain a state input vector;

[0059] inputting the state input vector into the main neural network to perform one-time strategy inference to obtain an action output and an actual available power corresponding to the action output;

[0060] performing difference calculation on the actual available power and the target power in the scheduling state data to obtain a power error parameter;

[0061] performing segmented reward mapping on the power error parameter based on a preset threshold to obtain an instant reward label;

[0062] combining and recording the instant reward label and the action output, and repeatedly performing strategy inference according to a preset upper limit of attempts within the same scheduling period to obtain a candidate action sequence arranged in an attempt order.

[0063] In this embodiment, the obtained scheduling status data is feature-encoded, and the target power, actual power deviation information and necessary context related to the current scheduling cycle are quantized into a state input vector of the same dimension and scale. Then, the state input vector is input into the main neural network to perform a policy reasoning to obtain the action output for power allocation of each unit and the actual power that can be generated corresponding to the action, which is used to characterize the response of each unit in the current state when the policy is adopted.

[0064] Secondly, the difference between the actual available power and the target power in the scheduling status data is calculated to obtain the power error parameter. This error is then segmented and mapped according to a power error threshold to obtain an immediate reward label, which characterizes the positive / negative feedback of this policy attempt. Specifically, the following segmented reward function can be used to evaluate the scheduling behavior of each policy attempt and provide an immediate reward or penalty: ;

[0065] in, This is the power error threshold; when hour, A positive value indicates a reward for the scheduling behavior of the main neural network; when... hour, A negative value indicates a penalty for the scheduling behavior of the main neural network.

[0066] Furthermore, the instant reward labels and corresponding action outputs are combined and recorded one by one to generate "action-reward" pairs. Within the same scheduling cycle, the process of "encoding-reasoning-error calculation-reward mapping-combination recording" is repeatedly executed according to the preset attempt limit to obtain candidate action sequences organized by time or attempt order, providing input for subsequent sequence scoring and experience playback. To avoid excessive attempts and improve scheduling efficiency, the number of attempts can be explicitly counted and constrained.

[0067] In step S103, the system detects the scheduling status data according to a preset convergence criterion. When the constraint with power deviation as the core meets the preset condition (e.g., the relative deviation does not exceed a specified threshold), it is determined that the power scheduling in this cycle is completed; otherwise, the system continues to execute subsequent strategy attempts and instruction issuance in this cycle until the convergence condition is met and then the current cycle ends.

[0068] In step S104, the recorded candidate action sequence is quantitatively evaluated. First, each attempt is associated with its scheduling state at the time of its occurrence by time index, and the immediate reward molecule of the attempt is calculated. Then, according to the configured hyperparameters, the immediate reward molecule is weighted and synthesized with the future reward molecule used to measure long-term efficiency, and the total score sequence arranged in the order of attempts is output to reflect the comprehensive benefit of each candidate action in the same scheduling period.

[0069] In an embodiment, the step S104 comprises:

[0070] indexing the candidate action sequence with the scheduling state data based on the attempt order to obtain an action-state pair sequence;

[0071] calculating an immediate reward score for each attempt in the action-state pair sequence to obtain an immediate reward score sequence;

[0072] weighting and synthesizing the immediate reward score sequence with a preset future reward score sequence based on a hyperparameter to obtain a total score sequence.

[0073] In the embodiment, the scheduling system indexes the candidate scheduling actions with the scheduling state data at the time of generation in a one-to-one manner based on the attempt order to obtain an action-state pair sequence arranged according to the attempt order; then for each attempt in the action-state pair sequence, the error-reward mapping determined in the step S104 is called to calculate the immediate reward score of the attempt according to the power error corresponding to the action, thereby obtaining an immediate reward score sequence arranged according to the attempt order; on this basis, the system weights and synthesizes the immediate reward score sequence with a preset future reward score (used to quantify the impact on subsequent attempts / overall efficiency in the current scheduling period) according to the configured hyperparameter for long-term return to obtain a total score sequence output according to the attempt order. To unify the calculation caliber, the synthesis relationship can be expressed as:

[0074] wherein, is the immediate reward score of the attempt, is the future reward score corresponding to the attempt, is a hyperparameter (for example, 0.9 can be taken) used to balance between short-term and long-term returns, is the total score of the corresponding attempt.

[0075] In an embodiment, the step S104 further comprises:

[0076] counting the attempts of the candidate action sequence in the same scheduling period to obtain the number of attempts;

[0077] performing reciprocal transformation and constant multiplication on the number of attempts to obtain a future reward score sequence.

[0078] In the embodiment, the control device counts the policy attempts experienced by the main neural network in performing power scheduling while the candidate action sequence is generated to obtain the number of attempts corresponding to the period ​(the number of times is used to measure the cost of attempts required to reach convergence in the current period). Since the number of attempts can only be determined after the main neural network completes the current power scheduling, it is considered zero during the scheduling process. When the current period record is written to the experience buffer, it is filled in according to the actual situation and the delay term is calculated accordingly; wherein the scheduling record with higher total score can be used as a better historical sample for subsequent learning. In determining , the system performs a reciprocal transformation on the number of times and multiplies a constant coefficient to obtain the future reward numerator sequence of the period, and aligns it with the candidate action sequence in the order of attempts (multiplexed to each attempt site as needed), thereby obtaining the future reward numerator sequence; the future reward numerator sequence is used to participate in the calculation of the total score sequence of the current period and the experience storage together with the immediate reward numerator sequence. In this embodiment, the future reward numerator sequence is defined as: ;

[0079] , wherein represents the total number of attempts of the main neural network in the power scheduling, and the constant coefficient can be 10 to realize scale normalization and weight stabilization.

[0080] In step S105, the total score sequence and the candidate action sequence are written into the experience buffer in a one-to-one correspondence; when the experience accumulation reaches a preset size (for example, 1000), the training process of the neural network is triggered, and the optimization strategy of random gradient descent combined with gradient accumulation descent is used to update the parameters of the slave neural network in the training stage, and the weights are synchronized to the main neural network after the training is completed, to obtain an updated model for subsequent scheduling period reasoning.

[0081] In an embodiment, the step S105 comprises:

[0082] Sampling the experience buffer to obtain a micro-batch sample set organized in time sequence;

[0083] Calculating the corresponding loss term for each micro-batch sample set to obtain a loss sequence;

[0084] Calculating the gradient of the parameter vector of the slave neural network before updating according to the loss sequence to obtain a gradient sequence;

[0085] Cumulatively adding each item in the gradient sequence in a cumulative step and performing mean value processing with the number of micro-batches as the divisor to obtain an average gradient;

[0086] Performing product operation on the average gradient and the learning rate parameter, and subtracting the product operation result from the parameter vector of the slave neural network before updating to obtain the parameter vector of the slave neural network after updating.

[0087] In this embodiment, the experience buffer is first sampled sequentially over time to obtain a micro-batch sample set consisting of several consecutive samples; then, the loss term is calculated for each micro-batch to obtain a loss sequence arranged in chronological order. The gradient sequence is obtained by calculating the gradient from the current parameter vector of the neural network based on this loss sequence. Within a cumulative step, the gradients of each micro-batch are accumulated and averaged using the batch size to obtain the average gradient (equivalent to a one-time update for a larger batch, if the loss of each micro-batch is already batch-mean). Finally, the average gradient is weighted by the learning rate, and this weighted result is subtracted from the parameter vector before the update to obtain the updated parameter vector. The above update can be calculated using the following formula: ;

[0088] in, and These represent the parameter vectors of the neural network before and after the update, respectively. Where is the learning rate, and N is the number of micro-batches for gradient accumulation. For the loss term of the i-th micro-batch, For the gradient of the loss with respect to the parameters under the current parameters, in the above equation... This is the cumulative gradient after mean normalization, used for a single parameter update.

[0089] In summary, this invention improves upon the insufficient adaptability of existing PID control in renewable energy scenarios such as solar power. By cooperating with a main neural network and a slave neural network, it achieves real-time power scheduling of photovoltaic generator sets and / or wind turbine generator sets. It can quickly converge to the target power and maintain stable control under dynamic operating conditions, significantly improving regulation efficiency and adaptability. At the same time, a simulation evaluation model is constructed for the verification of the method's effectiveness and the basis for parameter configuration, thereby ensuring the feasibility and consistency of the scheduling strategy.

[0090] Combination Figure 3 As shown, Figure 3 A schematic block diagram of a generator power dispatching device based on a neural network provided in this embodiment of the invention. The generator power dispatching device 300 based on a neural network includes:

[0091] Data receiving unit 301 is used to receive the target power sent by the central control room;

[0092] The network scheduling unit 302 is used to perform at least one strategy attempt within the same scheduling cycle using the main neural network. Each time a strategy attempt is performed, a power scheduling command for controlling the generator set controller is output. The difference between the target power and the real-time power of the generator set is calculated and action reasoning is performed to obtain scheduling status data and candidate action sequences, respectively.

[0093] The data judgment unit 303 is configured to complete power scheduling in the current scheduling period when the scheduling state data meets a preset condition.

[0094] The sequence calculation unit 304 is configured to calculate a total score sequence of the main neural network in the same scheduling period by using the candidate action sequence.

[0095] The network updating unit 305 is configured to record the total score sequence and the candidate action sequence to an experience buffer, and train a slave neural network based on the experience buffer to synchronize to the main neural network, to obtain an updated main neural network.

[0096] In the embodiment, the data receiving unit 301 receives a target power issued by a control room; the network scheduling unit 302 performs at least one strategy attempt in the same scheduling period by using the main neural network, outputs a power scheduling instruction for controlling a generator set controller each time a strategy attempt is performed, and respectively obtains scheduling state data and a candidate action sequence by performing difference calculation and action reasoning according to the target power and real-time power of the generator set; the data judgment unit 303 completes power scheduling in the current scheduling period when the scheduling state data meets a preset condition; the sequence calculation unit 304 calculates a total score sequence of the main neural network in the same scheduling period by using the candidate action sequence; and the network updating unit 305 records the total score sequence and the candidate action sequence to an experience buffer, and trains a slave neural network based on the experience buffer to synchronize to the main neural network, to obtain an updated main neural network.

[0097] In an embodiment, the neural network-based generator set power scheduling apparatus 300 is further configured to:

[0098] obtain a historical data set of the generator set, and perform model training on the historical data set to obtain a basic neural network;

[0099] perform deployment initialization on the basic neural network to respectively obtain the main neural network and the slave neural network.

[0100] In an embodiment, the network scheduling unit 302 is specifically configured to:

[0101] perform time alignment on the target power and the real-time power to obtain a power sample set;

[0102] add the power sample set according to the index of the generator set one by one to obtain actual available power of all generator sets;

[0103] perform absolute difference operation on the actual available power and the target power to obtain an absolute error;

[0104] The absolute error is structured and packaged to obtain scheduling state data.

[0105] In an embodiment, the network scheduling unit 302 is further specific for:

[0106] A relative deviation calculation is performed on the target power and the actual available power to obtain a relative deviation parameter.

[0107] When the relative deviation parameter is greater than a preset threshold, a gradient update algorithm is triggered to adjust the weights of the main neural network to obtain an updated main neural network.

[0108] An inference calculation is performed on the updated main neural network and the scheduling state data to obtain a power setting value set of each unit at the next time, and the power setting value set is packaged as a broadcast instruction set.

[0109] The broadcast instruction set is issued and the relative deviation is recalculated to obtain a new relative deviation parameter.

[0110] The comparison with the preset threshold is repeatedly performed using the new relative deviation parameter until it is less than or equal to the preset threshold, and a scheduling result is output.

[0111] In an embodiment, the network scheduling unit 302 is specific for:

[0112] Feature encoding is performed on the scheduling state data to obtain a state input vector.

[0113] The state input vector is input into the main neural network for one-time strategy inference to obtain an action output and an actual available power corresponding thereto.

[0114] A difference calculation is performed on the actual available power and the target power in the scheduling state data to obtain a power error parameter.

[0115] The power error parameter is segmented and reward-mapped based on a preset threshold to obtain an instant reward label.

[0116] The instant reward label and the action output are combined and recorded, and the strategy inference is repeated within a preset upper limit of attempts in the same scheduling period to obtain a candidate action sequence arranged in an attempt order.

[0117] In an embodiment, the sequence calculation unit 304 is specific for:

[0118] The candidate action sequence and the scheduling state data are index-associated based on the attempt order to obtain an action-state pair sequence.

[0119] calculate an immediate reward molecule for each attempt in the sequence of action-state pairs, to obtain a sequence of immediate reward molecules;

[0120] weight the sequence of immediate reward molecules and a preset sequence of future reward molecules based on hyperparameters, to obtain a sequence of total scores.

[0121] In an embodiment, the sequence calculation unit 304 is further specifically configured to:

[0122] count the number of attempts of the candidate action sequence in the same scheduling period, to obtain a number of attempts;

[0123] perform inverse transformation on the number of attempts and multiply the inverse transformation result by a constant, to obtain a sequence of future reward molecules.

[0124] In an embodiment, the network updating unit 305 is specifically configured to:

[0125] sample the experience buffer to obtain a set of micro-batch samples organized in time sequence;

[0126] calculate a loss term for each micro-batch sample in the set of micro-batch samples, to obtain a sequence of losses;

[0127] calculate a gradient of the parameter vector before the neural network is updated according to the sequence of losses, to obtain a sequence of gradients;

[0128] perform item-by-item accumulation on the sequence of gradients in one accumulation step and average the sequence of gradients by dividing the sequence of gradients by the number of micro-batches, to obtain an average gradient;

[0129] perform product operation on the average gradient and a learning rate parameter, and subtract the product operation result from the parameter vector before the neural network is updated, to obtain a parameter vector after the neural network is updated.

[0130] Since the embodiments of the device part correspond to the embodiments of the method part, the embodiments of the device part are described in the description of the embodiments of the method part, and are not described here.

[0131] The embodiments of the present application also provide an intermediate controller, which comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the method for power scheduling of a generator set based on a neural network when executing the computer program.

[0132] As shown in Figure 4 The embodiments of the present application also provide a system for power scheduling of a generator set based on a neural network, which comprises the intermediate controller (as shown in Figure 4 ), a total control module (as shown in Figure 4central control room) and multiple generator sets (such as photovoltaic and / or wind power generator sets) in the power plant, the central control module sends target instructions to the neural network-based generator set power scheduling device for data scheduling, and the neural network-based generator set power scheduling device sends the target instructions to the corresponding generator sets respectively. Figure 4

[0133] In this embodiment, the neural network-based generator set power scheduling device is integrated in the intermediate controller, the intermediate controller is provided with a memory and a processor, and a computer program executable on the processor is stored in the memory, and the program instantiates a master neural network and a slave neural network respectively when running; multiple generator sets are connected with the intermediate controller through wired or wireless industrial communication links for data interaction, and the central control module and the intermediate controller are connected through an upper monitoring network for reciprocal transmission of instructions and states, forming a hierarchical architecture of "central control module-intermediate controller-multiple generator sets".

[0134] The working process of the system is as follows: first, the central control module issues a target power instruction to the intermediate controller at the beginning of a scheduling period; the intermediate controller performs inference calculation on the target power and known constraints based on the weight parameters of the current master neural network, obtains a set of power set values corresponding to each generator set, and encapsulates the set as a broadcast instruction set, which is then issued to the corresponding generator set for execution through the communication link. Subsequently, the intermediate controller collects real-time output power information from each generator set, calculates the difference between the target power and the output power, forms scheduling state data for describing the current deviation, and generates an instant reward label for this strategy attempt based on the scheduling state data, which is used to evaluate the goodness of the power allocation action.

[0135] In the same scheduling period, the intermediate controller explicitly counts the strategy attempts of the master neural network, records the total number of attempts experienced to reach the current convergence condition; when a power scheduling is completed, the corresponding delay reward is calculated based on the recorded number of attempts, and the instant rewards are combined to form the total score metric of this scheduling. The intermediate controller writes each "action-state-score" generated by the master neural network in the current scheduling period into an experience buffer to obtain experience data organized in chronological order, which is used for subsequent learning.

[0136] ​To ensure the stability and computational efficiency of learning, when the amount of data accumulated in the experience buffer reaches a preset threshold (e.g., 1000), the intermediate controller triggers an offline / quasi-online training process of the from neural network; in the training stage, an optimization strategy combining gradient accumulation descent and stochastic gradient descent is adopted to iteratively update the parameters of the from neural network to fit the behavior-reward relationship in the experience data. After the training is completed, the updated weights of the from neural network are synchronized to the master neural network to obtain the updated master neural network for performing inference and instruction generation in the subsequent scheduling period, so as to form a closed-loop operation mechanism of "target issuance-strategy inference-instruction broadcast-deviation evaluation-experience accumulation-model update-weight synchronization" at the system level, realizing real-time power scheduling and continuous adaptive optimization of the photovoltaic generator set and / or the wind power generator set.

[0137] The various embodiments described in the specification are progressive in nature, and each embodiment focuses on the differences from other embodiments. The same or similar parts between embodiments can be referred to each other. For the system disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and the relevant part can be referred to the method part. It should be pointed out that for ordinary skilled in the art, without departing from the principles of the present application, the present application can be improved and modified, and these improvements and modifications also fall within the scope of protection of the claims of the present application.

[0138] It should also be noted that in the specification, the relationship terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between the entities or operations. Moreover, the terms "include", "contain" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or equipment including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or equipment. Without more limitations, the element defined by the statement "including a" does not exclude the presence of another identical element in the process, method, article or equipment including the element.

Claims

1. A method for power dispatch of a generator set based on neural network, characterized in that, The method comprises the following steps: receiving a target power issued by a general control room; performing at least one strategy attempt in the same scheduling period by using a master neural network, outputting a power scheduling instruction for controlling a generator set controller each time a strategy attempt is performed, and performing difference calculation and action reasoning on the target power and real-time power of the generator set to obtain scheduling state data and a candidate action sequence, respectively; when the scheduling state data meets a preset condition, completing power scheduling in the current scheduling period; calculating a total score sequence of the master neural network in the same scheduling period by using the candidate action sequence; recording the total score sequence and the candidate action sequence to an experience buffer, and training a slave neural network based on the experience buffer to synchronize to the master neural network to obtain an updated master neural network.

2. The neural network-based genset power dispatching method of claim 1, wherein, Before the step of receiving the target power issued by the general control room, the method comprises the following steps: obtaining a historical data set of the generator set, and performing model training on the historical data set to obtain a basic neural network; performing deployment initialization on the basic neural network to obtain the master neural network and the slave neural network, respectively.

3. The neural network-based genset power dispatching method of claim 1, wherein, The step of performing difference calculation and action reasoning on the target power and real-time power of the generator set to obtain scheduling state data and a candidate action sequence comprises the following steps: performing time alignment on the target power and the real-time power to obtain a power sample set; adding the power sample set one by one according to the index of the generator set to obtain actual available power of all generator sets; performing absolute difference operation on the actual available power and the target power to obtain an absolute error; performing structured packaging on the absolute error to obtain scheduling state data.

4. The neural network-based genset power dispatching method of claim 3, wherein, The step of performing difference calculation and action reasoning on the target power and real-time power of the generator set to obtain scheduling state data and a candidate action sequence further comprises the following steps: performing relative deviation calculation on the target power and the actual available power to obtain a relative deviation parameter; when the relative deviation parameter is greater than a preset threshold, triggering a gradient update algorithm to adjust the weight of the master neural network to obtain an updated master neural network; performing reasoning calculation on the updated master neural network and the scheduling state data to obtain a power setting value set of each unit at the next moment, and packaging the power setting value set as a broadcast instruction set; issuing the broadcast instruction set and recalculating the relative deviation to obtain a new relative deviation parameter; repeating the comparison with the preset threshold by using the new relative deviation parameter until the new relative deviation parameter is less than or equal to the preset threshold, and outputting a scheduling result.

5. The neural network-based genset power dispatching method of claim 1, wherein, The step of performing difference calculation and action reasoning on the target power and real-time power of the generator set to obtain scheduling state data and a candidate action sequence further comprises the following steps: performing feature encoding on the scheduling state data to obtain a state input vector; inputting the state input vector into the master neural network to perform one-time strategy reasoning to obtain an action output and an actual available power corresponding to the action output; performing difference calculation on the actual available power and the target power in the scheduling state data to obtain a power error parameter; Segmenting and reward mapping the power error parameter based on a preset threshold to obtain an instant reward label; Combining the instant reward label and the action output for record, and repeatedly reasoning in a preset upper limit of attempts in the same scheduling period to obtain a candidate action sequence arranged in an attempt order.

6. The neural network-based genset power dispatching method of claim 1, wherein, The total score sequence of the main neural network in the same scheduling period is calculated using the candidate action sequence, including: Indexing and associating the candidate action sequence with the scheduling state data based on the attempt order to obtain an action-state pair sequence; Calculating the instant reward numerator for each attempt in the action-state pair sequence to obtain an instant reward numerator sequence; Based on the hyperparameters, the instant reward numerator sequence is weighted and synthesized with a preset future reward numerator sequence to obtain a total score sequence.

7. The neural network-based genset power dispatching method of claim 6, wherein, The total score sequence of the main neural network in the same scheduling period is calculated using the candidate action sequence, further including: Statistical record of attempts of the candidate action sequence in the same scheduling period to obtain the number of attempts; Inverse transformation and constant multiplication of the number of attempts to obtain a future reward numerator sequence.

8. The neural network-based genset power dispatching method of claim 1, wherein, The slave neural network is trained based on the experience buffer, including: Sampling the experience buffer to obtain a micro-batch sample set organized in time sequence; Calculating the corresponding loss term for each micro-batch sample set to obtain a loss sequence; According to the loss sequence, the gradient of the parameter vector before updating the slave neural network is calculated to obtain a gradient sequence; The gradient sequence is added up item by item in one accumulation step and averaged by dividing by the number of micro-batches to obtain an average gradient; The average gradient is multiplied by the learning rate parameter, and the parameter vector before updating the slave neural network is subtracted from the product operation result to obtain the parameter vector after updating the slave neural network.

9. A neural network-based power plant power dispatching device, characterized by, Including: A data receiving unit for receiving a target power issued by a total control room; A network scheduling unit for performing at least one strategy attempt in the same scheduling period using the main neural network, outputting a power scheduling instruction for controlling the generator set controller each time a strategy attempt is performed, and calculating the difference between the target power and the real-time power of the generator set and reasoning the action to obtain scheduling state data and a candidate action sequence, respectively; A data judgment unit for completing power scheduling in the current scheduling period when the scheduling state data meets the preset conditions; A sequence calculation unit for calculating the total score sequence of the main neural network in the same scheduling period using the candidate action sequence; A network updating unit for recording the total score sequence and the candidate action sequence to the experience buffer, and training the slave neural network based on the experience buffer to synchronize to the main neural network to obtain the updated main neural network.

10. A neural network-based power plant power dispatch system, characterized by, The neural network-based generator set power scheduling device, the total control module and the plurality of generator sets as claimed in claim 9 are included, the total control module sends target instructions to the neural network-based generator set power scheduling device for data scheduling, and sends to the corresponding generator set through the neural network-based generator set power scheduling device respectively.

Citation Information

Patent Citations

  • Trajectory tracking control method and device, equipment and storage medium

    CN111665861A

  • Systems and methods for next-best action using a multi-objective reward based sequential framework

    US20250245478A1