Data-driven non-zero-sum game multi-energy system scheduling method and device

By employing a data-driven non-zero-sum game-based multi-energy system scheduling method, and utilizing internal model dynamics and output feedback to construct a cost function and iterative learning algorithm, the problem of unknown parameters and disturbances in multi-energy systems is solved. This method achieves Nash equilibrium and disturbance suppression, thereby improving the transient response performance of the system.

CN120975521BActive Publication Date: 2026-02-06BEIJING INST OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511492492.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-20
Publication Date
2026-02-06
Estimated Expiration
2045-10-20

AI Technical Summary

Technical Problem

In multi-energy systems, when parameters and disturbances are unknown, traditional non-zero-sum game output adjustment methods based on exact models are difficult to effectively achieve Nash equilibrium among various energy systems.

Method used

This paper designs a data-driven non-zero-sum game-based multi-energy system scheduling method. By measuring system data, a feedback scheduling strategy model is constructed. The cost function is established by utilizing internal model dynamics and output feedback. Stable control strategies are obtained through iterative learning to achieve control and scheduling of each energy supplier.

Benefits of technology

Under conditions of unknown parameters and disturbances, this method achieves Nash equilibrium in multi-energy systems, suppresses dynamically changing external disturbances, improves transient response performance, reduces computational complexity, and adapts to non-initial stable system scheduling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120975521B_ABST
    Figure CN120975521B_ABST
Patent Text Reader

Abstract

The application provides a data-driven non-zero-sum game multi-energy system scheduling method and device, and belongs to the technical field of game control and decision-making. The method designs a scheduling strategy expression form based on an internal model system, and the control strategy gain in the strategy and the internal model parameter of the internal model system are unknown but can be solved. The internal model parameter is solved based on external system state data, and the control strategy gain estimation value is determined based on two-stage learning, so that the balanced scheduling strategy of each energy supplier in the multi-energy system is solved. By using the application, Nash equilibrium between each energy system can be realized only through measurable system data under the condition that the parameters and disturbances of the multi-energy system are unknown, and the limitation of traditional accurate model technology is made up.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of game control and decision-making, in particular to a data-driven non-zero-sum game multi-energy system scheduling method and device. BACKGROUND

[0002] The output control scheduling method of multi-energy system non-zero-sum game is a strategic adjustment technology in the game theory framework for complex interactive environment. Unlike zero-sum game, non-zero-sum game emphasizes that the relationship between participants is not purely competitive, but can seek common interests to achieve a win-win or multi-win situation. In multi-energy system non-zero-sum game, the income or loss of each participant depends not only on its own decision, but also on the decision of other participants. Therefore, the scheduling method aims to maximize the overall benefit or achieve a certain cooperation goal by adjusting the strategies of each participant.

[0003] However, in actual multi-energy systems, there are often problems such as unknown parameter information and external disturbances, making it difficult to directly use traditional non-zero-sum game output regulation methods based on accurate models.

[0004] Therefore, a data-driven non-zero-sum game output control scheduling algorithm is designed to achieve Nash equilibrium between each energy system only through measurable system data under the condition of unknown multi-energy system parameters and disturbances. SUMMARY

[0005] Therefore, the present application provides a data-driven non-zero-sum game multi-energy system scheduling method and device, which can achieve Nash equilibrium between each energy system only through measurable system data under the condition of unknown multi-energy system parameters and disturbances.

[0006] To solve the above technical problems, the present application is implemented as follows.

[0007] A data-driven non-zero-sum game multi-energy system scheduling method, comprising:

[0008] Step 1: based on the multi-energy system model and the external system model with unknown system parameters, a feedback scheduling strategy model based on internal model dynamics and output feedback is designed, wherein the control strategy gain and the internal model parameters of the internal model system are unknown;

[0009] Step 2: by measuring the state data of the external system, the characteristic polynomial of the external system matrix is constructed, and the internal model parameter value is determined by using the coefficients of the characteristic polynomial;

[0010] Step 3: a cost function is established for the multi-energy system model and the feedback scheduling strategy model; the cost function adds a cost term to guide the steady state of the internal model;

[0011] Step 4: based on the multi-energy system model, taking the inner model system as an augmented term, a first augmented system is obtained; then, by using the solvability of the regulator equation, the first augmented system is converted into a second augmented system with the same control strategy gain by combining coordinate transformation;

[0012] Step 5: based on the cost function, the second augmented system is subjected to first-stage iterative learning using measurable multi-energy system data, and a stable control strategy is obtained;

[0013] Step 6: based on the cost function and the stable control strategy, second-stage iterative learning is performed to obtain the control strategy gain estimation value of each energy supplier in the multi-energy system; the control strategy gain estimation value and the inner model parameter value determined in step 2 are substituted into the feedback scheduling strategy model to obtain the control scheduling strategy of each energy supplier.

[0014] Preferably, in step 1, the feedback scheduling strategy model is:

[0015]

[0016] wherein,

[0017]

[0018] wherein, denotes a discrete time; is the control scheduling strategy of the i-th energy supplier; i is the system state of the multi-energy system; is the system tracking output of the multi-energy system; denotes the inner model system; and is the control strategy gain of the i-th energy supplier; and i is the inner model parameter of the inner model system; , , and are parameters to be determined. Preferably, in step 2, the characteristic polynomial of the external system matrix is constructed by measuring the state data of the external system, when the external system state data satisfies the set initial excitation condition, the parameter identification update of the external system matrix is performed based on the filter equation, and the characteristic polynomial is constructed according to the parameter-identified external system matrix.

[0019] Preferably, in step 2, the inner model parameter value is determined by using the coefficients of the characteristic polynomial:

[0020] Preferably, in step 2, the inner model parameter value is determined by using the coefficients of the characteristic polynomial:

[0021] ​The characteristic polynomial of the external system system matrix is expressed as ;

[0022] wherein, is an eigenvalue of the external system system matrix; is an external system state dimension; is a characteristic polynomial coefficient;

[0023] The internal model system is designed using the characteristic polynomial coefficient as follows:

[0024]

[0025] wherein, , , and

[0026] ,

[0027] wherein, denotes a discrete time instant; denotes an internal model system; is a system tracking output of the multi-energy system; and are internal model parameters; denotes a Kronecker product operation; is an n dimensional unit vector; n is a system state dimension;

[0028] The characteristic polynomial coefficient obtained based on the external system state data is substituted into the internal model system to obtain values of the internal model parameters and .

[0029] Preferably, in step 3, the cost function is designed as a sum of a cooperation cost term and a competition cost term, wherein the cooperation cost term adds a cost term for guiding the internal model to a steady state; the cooperation cost term is minimized to achieve the goal of closed-loop stability and output tracking of the multi-energy system in cooperation, and the competition cost term is minimized to achieve the goal of Nash equilibrium among the multiple energy supply systems; by solving the minimized cost function for each energy system, the non-zero-sum game multi-energy system scheduling is achieved.

[0030] Preferably, in step 5, the first-stage iterative learning adopts a value iteration algorithm; the second-stage iterative learning adopts a policy iteration algorithm; a switching iteration criterion is used to determine whether a stable control strategy is obtained; if so, the second-stage iterative learning is switched to;

[0031] The switching iteration criterion is that eigenvalues of a closed-loop system matrix of the multi-energy system are calculated according to a current iteration result; when the eigenvalues are less than 0, it is determined that a stable control strategy is obtained.

[0032] Preferably, the multi-energy system model is:

[0033]

[0034] wherein, denotes a discrete time, is a system state of the multi-energy system, is a control scheduling strategy of an i-th energy supplier, i is an external system state, is a system tracking output of the multi-energy system; 、 、 、 、 、 is an unknown system parameter matrix, and N is a number of energy suppliers included in the multi-energy system;

[0035] Step 4 converts the obtained second augmented system into:

[0036]

[0037] wherein, ; , , is a stable system state, is a stable internal model state. is a vector merging operator; , is a stable control scheduling strategy; , .

[0038] The application also provides a data-driven non-zero-sum game multi-energy system scheduling device, comprising:

[0039] a model construction module configured to design a feedback scheduling strategy model based on an internal model dynamic and output feedback based on a multi-energy system model and an external system model, wherein a control strategy gain and an internal model parameter of the internal model system are unknown;

[0040] an internal model parameter determination module configured to construct a characteristic polynomial of a system matrix of the external system by measuring state data of the external system, and determine an internal model parameter value by using coefficients of the characteristic polynomial;

[0041] a model conversion module configured to convert the multi-energy system model into an augmented system based on the multi-energy system model, take the internal model system as an augmented term, and combine a regulator equation to be solvable;

[0042] The first optimization module is configured to perform first-stage iterative learning on the augmented system based on a cost function and measurable multi-energy system data to obtain a stable control strategy, wherein the cost function adds a cost term for guiding the internal model to a steady state.

[0043] The second optimization module is configured to perform second-stage iterative learning based on the cost function and the stable control strategy to obtain control strategy gain estimation values of the energy suppliers in the multi-energy system.

[0044] The strategy generation module is configured to substitute the control strategy gain estimation values and the internal model parameter values determined by the internal model parameter determination module into a feedback scheduling strategy model to obtain the control scheduling strategies of the energy suppliers.

[0045] Preferably, the feedback scheduling strategy model constructed by the model construction module is as follows:

[0046]

[0047] wherein,

[0048]

[0049] wherein, denotes a discrete time; is a control scheduling strategy of an i-th energy supplier; i is a system state of the multi-energy system; is a system tracking output of the multi-energy system; denotes an internal model system; and is a control strategy gain of the i-th energy supplier; and i is an internal model parameter of the internal model system; , , and are unknown design parameters; The internal model parameter determination module constructs a characteristic polynomial of an external system matrix as follows:

[0050]

[0051] wherein,

[0052] is an eigenvalue of the external system matrix; is a dimension of the external system state; is a characteristic polynomial coefficient; The internal model system is designed using the characteristic polynomial coefficient as follows:

[0053]

[0054] ​​

[0055] wherein, , , and

[0056] ,

[0057] wherein, denotes the Kronecker product operation; is a unit vector of dimension n ; n is the system state dimension;

[0058] substitute the characteristic polynomial coefficients obtained based on the external system state data into the internal model system to obtain the values of the internal model parameters and .

[0059] Preferably, the first optimization module is further used to determine whether a stable control strategy is obtained by using a switching iteration criterion; if yes, switching to the second optimization module for second-stage iteration learning.

[0060] The switching iteration criterion is that eigenvalues of a closed-loop system matrix of the multi-energy system are calculated according to a current iteration result; when the eigenvalues are less than 0, it is determined that a stable control strategy is obtained.

[0061] Advantages:

[0062] (1) In the case that parameters and disturbances of the multi-energy system are unknown, the present application only uses measurable system data to update energy supply scheduling strategies of each energy supplier online, so that the multi-energy system can achieve Nash equilibrium among each energy supply system while achieving disturbance suppression and tracking a reference output trajectory.

[0063] (2) In the case that an external system dynamic is unknown, the present application only uses measurable external system state data to update internal model parameters online, and the method can self-adaptively adjust the internal model parameters to suppress dynamically changing external disturbances.

[0064] (3) The cost function designed in the present application adds an internal model steady-state term, and by minimizing the cost term, the internal model can be dynamically quickly converged to an internal model steady-state trajectory, so that the transient response performance of the multi-energy system is improved.

[0065] (4) The present application constructs two types of augmented systems which can be converted to each other, and by using the characteristic that the two types of augmented systems have the same system state matrix and input matrix, only the solvability of the regulator equation is used to establish a data-driven iteration learning algorithm; compared with existing algorithms, the present application only needs to satisfy the solvability of the regulator equation, without solving the regulator equation, so that the calculation complexity of the learning algorithm is reduced.

[0066] (5) The application establishes a two-stage iterative method in the optimization process, and designs a switching criterion; the first-stage iterative learning established does not need a stable initial feedback gain, and in the iterative process, a switching iterative criterion is used to determine whether a stable control strategy is obtained, if yes, switching to the second-stage strategy iterative learning; compared with existing methods which all need a stable initial feedback gain, the application can be applied to realize non-initial stable multi-energy system non-zero-sum game scheduling. BRIEF DESCRIPTION OF DRAWINGS

[0067] Figure 1 It is a schematic diagram of the multi-energy system of the application.

[0068] Figure 2 It is an implementation flowchart of the data-driven non-zero-sum game multi-energy system scheduling method of the embodiment of the application.

[0069] Figure 3 It is a structure block diagram of the data-driven non-zero-sum game multi-energy system scheduling device of the embodiment of the application.

[0070] Figure 4 It is a learning algorithm iterative convergence diagram of the specific example of the application;

[0071] Figure 5 It is a scheduling strategy diagram of the energy supplier 1, 2, 3 and 4 of the specific example of the application;

[0072] Figure 6 It is a multi-energy system output tracking effect diagram of the specific example of the application;

[0073] Figure 7 It is a multi-energy system state trajectory diagram of the specific example of the application. DETAILED DESCRIPTION

[0074] The application will be described in detail below in combination with the drawings and examples.

[0075] Figure 1The schematic diagram of a multi-energy system to which the scheduling scheme of the application is directed, the system comprising N multiple energy suppliers, the energy supplier output forming the total energy output of the system after being processed by the comprehensive processing module. The comprehensive processing module in the system is complex in structure, not fixed, and the external disturbance is also unknown, so it is difficult to accurately model. The application is designed for this situation, a data-driven non-zero-sum game multi-energy system scheduling scheme, in the case of unknown multi-energy system parameters and disturbances, only through measurable system data, Nash equilibrium between each energy system is realized. And, through the modeling of external disturbance and target trajectory, using the measurable external system state, an external system identification method based on initial excitation is established, the minimum characteristic polynomial system of the identified external system matrix is used to construct the internal mode system parameter, the core function of the internal mode system is to provide the compensation term corresponding to the disturbance and the target trajectory, finally, a scheduling strategy based on the internal mode system is designed, the multi-energy system output can track the set category of reference trajectory, and can also suppress the set category of interference signal.

[0076] Figure 2 The implementation flowchart of the data-driven non-zero-sum game multi-energy system scheduling method of the embodiment of the application is shown. Referring to the drawings, the steps of the method include:

[0077] Step S1: a discrete-time multi-energy system model with unknown system parameters is established, as follows:

[0078]

[0079] Wherein, k represents the discrete time, is the system state of the multi-energy system, is the control scheduling output of the energy supplier, is the external system state, is the system tracking output, is the external disturbance signal, is the expected tracking reference output signal, , , , and are unknown system parameter matrices. represents the real number field, n , m i , p , q、 、 are the multi-energy system state, the ithe dimension of the control scheduling output of the energy supplier, the tracking output of the multi-energy system, the external system state, the external disturbance signal, and the reference output signal; define a set , N is the number of energy suppliers in the multi-energy system; the augmented matrix .

[0080] In the present application, it is assumed that the system matrix satisfies is stabilizable and This assumption makes the regulator equation constructed in step S5 solvable.

[0081] Step S2: Establish an external system model of a class of unknown system parameters. Here, the external system model includes an external disturbance model and a target trajectory model. The selection of the "class" is the class of the multi-energy system output capable of tracking the reference trajectory, and the class capable of suppressing the disturbance signal.

[0082] In the multi-energy system established in step S1, the external disturbance is , and the expected reference output trajectory to be tracked is , which is generated by the following autonomous system:

[0083]

[0084] wherein and are the parameter-unknown autonomous system matrices. Define , wherein is a vector merging operator, and the external system can be obtained as follows:

[0085]

[0086] wherein is the parameter-unknown external system matrix. This class of external systems can generate a series of signals, such as a combination of step functions with arbitrary amplitudes, sine functions with arbitrary amplitudes and initial phases, and ramp functions with arbitrary slopes. In combination with the external system, the original system can be described as:

[0087]

[0088] wherein , .

[0089] Step S3: For the multi-energy system established in step S1 and the external system established in step S2, a feedback scheduling strategy model based on the internal model dynamics and output feedback is designed, which has the following form:

[0090]

[0091] wherein the reference internal model is:

[0092]

[0093] wherein, is an inner model system, ; and is the control strategy gain of the first energy supplier; i and is the inner model parameter of the inner model system; and is the parameter to be determined in the subsequent step. Step S4: By measuring the state data of the external system, an external system system matrix identification method based on initial excitation is established, a characteristic polynomial of the external system system matrix is constructed, and the inner model parameter value is determined by using the coefficients of the characteristic polynomial.

[0094] In this step, according to the external system model established in step S2, the core filter equation is designed by using the measurable external system data as follows:

[0095]

[0096]

[0097]

[0098] wherein, is an adjustable identification gain, the smaller the value, the faster the convergence speed; is the measured external system state, and the dimension is , is a data regression matrix, is a unit matrix with a dimension of , is the dimension of the external system state, denotes the Kronecker product operation, and are filter data regression terms. When the collected external system state data meets the initial excitation condition, the identification update equation is established based on the core filter equation as follows:

[0099]

[0100] Based on the above identification update equation, the parameter identification update of the external system system matrix is performed. is the parameter estimation value of the external system system matrix. The eigenvalue is calculated by using the parameter-identified external system system matrix, and the characteristic polynomial is constructed.

[0101] The initial excitation condition may be, for example, that for each time , the external system state data is collected​ , the composition data regression matrix , if there is a time satisfying

[0102]

[0103] The collected external system state data is said to satisfy the initial excitation condition, wherein, is an arbitrarily small positive real number defined by the user, is a unit matrix of dimension .

[0104] The above characteristic polynomial is expressed as , wherein is an eigenvalue of the external system matrix; the characteristic polynomial coefficient is a calculable constant. Using the coefficients of the characteristic polynomial, the internal model system is designed as follows:

[0105]

[0106] , wherein, , , and

[0107] , .

[0108] In the formula, denotes the Kronecker product operation; is a unit vector of dimension n .

[0109] Substitute the characteristic polynomial coefficients obtained based on the external system state data into the above formula to obtain the values of the internal model parameters and .

[0110] Step S5: Establishing the cost function of the multi-energy system for the multi-energy system established in step S1 and the feedback scheduling strategy model designed in step S3.

[0111] The application adds a cost term for guiding the internal model steady state in the cost function, and by minimizing the cost term, the internal model dynamic can quickly converge to the internal model steady state trajectory; when the internal model dynamic is in a steady state, the internal model system can generate a compensation signal to offset external disturbances and track target trajectories, thereby achieving disturbance suppression and target tracking of the multi-energy system.

[0112] In a preferred embodiment, the cost function is designed to be the sum of a cooperation cost term and a competition cost term, wherein the cooperation cost term adds a cost term for guiding the internal model to a steady state; the cooperation cost term is minimized to achieve the closed-loop stability of the multi-energy system and the output tracking, and the competition cost term is minimized to achieve the Nash equilibrium among the multiple energy supply systems; by solving the cost function for each energy system, the non-zero-sum game multi-energy system scheduling is finally achieved.

[0113] Based on the above design idea, a specific design of a cost function is given as follows:

[0114]

[0115] wherein, is the cost function of the j th energy supply system; , , when , when is a user-defined weight matrix, is defined as , and , and , i.e. , , , are the stable system state, the stable internal model state and the stable control scheduling strategy, respectively, and the specific values are:

[0116]

[0117] wherein, and are the solutions of the following regulator equation:

[0118]

[0119] It is worth noting that in the entire solving process, the above regulator equation is not required to be solved, only the solvability of the regulator equation is required, so that an augmented system with the same system matrix can be constructed.

[0120] Step S6: According to the internal model system obtained in step S4, a new augmented system is constructed, and a new cost function is converted according to the multi-energy system cost function designed in step S5.

[0121] In this step, first, based on the multi-energy system model, the internal model system is taken as an augmented term to obtain a first augmented system.

[0122] Definition , and where is a vector merging operator.

[0123] Based on the multi-energy system model, a first augmented system is obtained by taking the internal model system as an augmented term:

[0124]

[0125] Then, by using the solvability of the regulator equation and combining the coordinate transformation, the first augmented system is converted into a second augmented system with the same state system matrix and input system matrix:

[0126]

[0127] wherein , , and .

[0128] Compared with the first augmented system, the second augmented system no longer explicitly contains the disturbance term and the target tracking term The reason is that, by combining the solution of the regulator equation, the internal model dynamics can generate a compensation signal to offset the external disturbance and the tracking target trajectory, and by changing the coordinates representing the compensation signal, the first augmented system can be written as the second augmented system.

[0129] At this time, the multi-energy system cost function designed in step S5 can be converted into:

[0130]

[0131] wherein , the matrix , , is a user-defined weight matrix in step S5.

[0132] The classical method to minimize the above multi-energy system cost function is to solve the following coupled algebraic Riccati equation:

[0133]

[0134] wherein is a matrix to be solved, is an optimal feedback gain to be solved, and the solution form is

[0135]

[0136] It can be easily found that the solution of the above equation depends on the unknown system matrix and Therefore, it is not directly solvable. In particular, it is noted that the first and second augmented systems have the same system matrix and Therefore, a data-driven algorithm can be established to solve the feedback gain using the measurable state, input, and output data of the first augmented system. Meanwhile, since the state, input, and output data of the second augmented system are not used in the entire data-driven algorithm construction process, there is no need to solve the regulator equation to obtain the stable system state, stable internal model state, and stable control scheduling strategy , .

[0137] The following steps S7-S9 are an iterative optimization process, which includes first-stage iterative learning and second-stage iterative learning. The first-stage iterative learning adopts a value iteration algorithm (step S7), and the second-stage iterative learning adopts a policy iteration algorithm (step S9). Whether to switch from the first stage to the second stage is determined by using a switching iteration criterion (step S8).

[0138] Step S7: Based on the cost function, the second augmented system is subjected to first-stage iterative learning using measurable multi-energy system data, to obtain a stable control strategy.

[0139] First, using the measurable multi-energy system state data, input data, and output data, a value iteration-based iterative learning equation is designed in combination with a value iteration algorithm.

[0140] Let be a symmetric matrix vectorization operator, be an asymmetric matrix vectorization operator, be a vector-to-diagonal matrix operator. Using the state data and , the input data , and the system output data , for each k time, the value iteration-based iterative learning equation is designed as follows:

[0141]

[0142] wherein ;

[0143] ;

[0144] .

[0145] The matrix parameters to be estimated are: , , , , , . wherein, is the solution of the Lyapunov equation of the multi-energy system in the i-th iteration process.

[0146] The method for solving the iterative learning equation by using the least square method is as follows:

[0147]

[0148] wherein,

[0149]

[0150]

[0151] wherein, is a pseudo-inverse operation, is the discrete time of the s time data.

[0152] The control strategy updating equation is:

[0153]

[0154] wherein, is an arbitrary matrix. The control gain in the above equation is wherein, . By solving the above equation, the to-be-designed parameters and in step S3 can be obtained.

[0155] The cost function value estimation equation is:

[0156]

[0157] wherein, is an arbitrary matrix. This step can ensure that the iteration number satisfies is a stable control strategy.

[0158] Step S8: It is judged whether the switching standard is reached. If yes, step S9 is executed; otherwise, step S7 is continuously executed.

[0159] In a preferred solution, the design of the switching iteration criterion is as follows: the eigenvalue of the closed-loop system matrix of the multi-energy system is calculated according to the current iteration result; when the eigenvalue is less than 0, it is determined that a stable control strategy is obtained.

[0160] According to the solution of the value iteration learning equation established in step S7, the switching iteration criterion is expressed as follows: ​​

[0161]

[0162] wherein, , is the eigenvalue calculation operation. Take the eigenvalue of , determine whether it is less than 0, and if so, it is considered that the switching iteration criterion is met.

[0163] After the above switching iteration criterion is met, the iteration equation is switched from the value iteration learning equation designed in step S7 to the policy iteration learning equation designed in step S9.

[0164] Step S9: Based on the cost function, in combination with the stable control strategy obtained in step S7, a second-stage iteration learning is performed to obtain the control strategy gain estimation value of each energy provider in the multi-energy system.

[0165] In combination with the stable control strategy obtained in step S7, the measurable multi-energy system state data, input data and output data are used in combination with the policy iteration algorithm to design the iteration learning equation;

[0166] Define . In combination with the operator defined in step S7, the state data and , the input data and the output data , for each k time, for the first energy provider system, the iteration learning equation based on policy iteration is designed as follows:

[0167]

[0168] wherein, . The matrix parameters to be estimated are: , , , , . Wherein, is the solution of the multi-energy system Lyapunov equation for updating the input of the first energy provider in the first iteration process. Similar to step S7, s time data are selected, and the least square method is used to solve the iteration learning equation based on policy iteration in step S9.

[0169] In the first iteration, for the first energy provider, the control strategy update equation is:

[0170]

[0171] wherein, is an estimated value is the th sub-main diagonal matrix. A threshold value is defined, and when the iteration ends, and the control strategy obtained in this iteration is called the estimated optimal control strategy .

[0172]

[0173] Step S10: According to the dynamic feedback control strategy designed in step S3, the internal model system estimated in step S4, and the control gain estimated in step S9, a robust control strategy for the multi-energy system established in step S1 and the external system established in step S2 is obtained as follows:

[0174]

[0175] In the above formula, and are the data obtained in step S4; and are the results obtained in step S9.

[0176] The robust control strategy set of energy participants obtained above, i.e. , is the optimal control strategy in the sense of Nash equilibrium.

[0177] Based on the above method, the application also provides a data-driven non-zero-sum game multi-energy system scheduling device, as shown in Figure 3 , which comprises a model construction module, an internal model parameter determination module, a model conversion module, a first optimization module, a second optimization module, and a strategy generation module. Among them,

[0178] The model construction module is used to design a feedback scheduling strategy model based on internal model dynamics and output feedback based on a system parameter unknown multi-energy system model and an external system model, wherein the control strategy gain and the internal model parameter of the internal model system are unknown. In a preferred embodiment, the model construction module corresponds to steps S1-S3 above.

[0179] The internal model parameter determination module is used to construct the characteristic polynomial of the external system matrix by measuring the state data of the external system, and determine the internal model parameter value by using the coefficients of the characteristic polynomial. In a preferred embodiment, the internal model parameter determination module corresponds to step S4 above.

[0180] a model conversion module configured to convert the multi-energy system model into an augmented system based on the multi-energy system model, the inner model system as an augmented term, and the solvability of the regulator equation. In a preferred embodiment, the model conversion module corresponds to step S6 above.

[0181] a first optimization module configured to perform a first stage of iterative learning of the augmented system based on a cost function and measurable multi-energy system data to obtain a stable control policy, wherein the cost function incorporates a cost term for guiding the inner model to a steady state. In a preferred embodiment, the first optimization module corresponds to step S7 above.

[0182] a second optimization module configured to perform a second stage of iterative learning based on the cost function and the stable control policy to obtain control policy gain estimates for each energy supplier in the multi-energy system. In a preferred embodiment, the second optimization module corresponds to step S9 above.

[0183] a policy generation module configured to substitute the control policy gain estimates and the inner model parameter values determined by the inner model parameter determination module into a feedback scheduling policy model to obtain a control scheduling policy for each energy supplier.

[0184] wherein the feedback scheduling policy model constructed by the model construction module is:

[0185] wherein,

[0186]

[0187] wherein, denotes a discrete time instant; is a control scheduling policy for the i th energy supplier; is a system state of the multi-energy system; is a system tracking output of the multi-energy system; denotes an inner model system; and is a control policy gain for the i th energy supplier; and is an inner model parameter of the inner model system; , , and are unknown design parameters;

[0188] the inner model parameter determination module constructs a characteristic polynomial of a system matrix of the outer system as:

[0189]

[0190] wherein, Eigenvalues of the system matrix of the external system State dimension of the external system Characteristic polynomial coefficients

[0191] Design the internal model system using the characteristic polynomial coefficients as

[0192]

[0193] wherein, , , and

[0194] ,

[0195] wherein, denotes the Kronecker product operation; is a unit vector of dimension n ; n is the state dimension of the system

[0196] Substitute the characteristic polynomial coefficients obtained based on the state data of the external system into the internal model system to obtain the values of the internal model parameters and .

[0197] The first optimization module is further configured to, corresponding to step S8, determine whether a stable control strategy is obtained using a switching iteration criterion; if yes, switch to the second optimization module to perform second-stage iterative learning; wherein the switching iteration criterion is: calculating eigenvalues of a closed-loop system matrix according to a current iteration result; when the eigenvalues are less than 0, it is determined that a stable control strategy is obtained.

[0198] The following describes the implementation process and effects of the present application by taking specific cases.

[0199] Step S1: consider a discrete-time system model containing 4 energy suppliers as follows:

[0200]

[0201] It is worth mentioning that the specific parameters of the above system matrix are given to verify the effectiveness of the data-driven algorithm, and the above matrix parameters are not required in the algorithm.

[0202] Step S2: consider an external system model as follows:

[0203]

[0204] wherein, the external disturbance is , and the reference output trajectory is .

[0205] Step S3: design the form of dynamic feedback control policy as follows:

[0206]

[0207] where, is the internal model state, and are the parameters to be estimated.

[0208] Step S4: design the internal model parameters as follows:

[0209]

[0210] Step S5, Step S6: obtain the multi-energy system cost function as follows:

[0211]

[0212] where, the rest .

[0213] Step S7, Step S8, Step S9: obtain the control policy gain estimation values of the four energy supply participants by the proposed iterative learning algorithm:

[0214]

[0215]

[0216]

[0217]

[0218] Step S10: update the control policy of the four energy supply participants as follows:

[0219]

[0220] where, and are the estimation values obtained in Step S9, is the parameter value designed in Step S4.

[0221] Simulation is carried out in matlab, and the simulation results are as follows: Figure 4 、 Figure 5 、 Figure 6 and Figure 7 . It can be seen that the present method does not depend on accurate system parameter information, and only uses system state, input and output information, and At time 129, the estimated optimal control strategy of the four energy suppliers is obtained after 41 iterations. Simulation results show that the proposed data-driven method can handle the problem of unknown parameters and unknown external disturbances of the multi-energy system, and can achieve the effect of tracking the reference signal and disturbance suppression in the sense of Nash equilibrium.

[0222] The above specific embodiments only describe the design principles of the present application, and the shapes and names of the components in the description can be different and are not limited. Therefore, those skilled in the art of the present application can modify or equivalently replace the technical solutions described in the foregoing embodiments; and these modifications and replacements do not deviate from the purpose and technical solutions of the present application, and should all belong to the protection scope of the present application.

Claims

1. A data-driven non-zero-sum game-based multi-energy system scheduling method, characterized in that, include: Step 1: Based on the multi-energy system model and external system model with unknown system parameters, design a feedback scheduling strategy model based on internal model dynamics and output feedback, where the control strategy gain and the internal model parameters of the internal model system are unknown; Step 2: By measuring the state data of the external system, construct the characteristic polynomial of the external system matrix, and use the coefficients of the characteristic polynomial to determine the internal model parameter values; Step 3: Establish a cost function for the multi-energy system model and the feedback scheduling strategy model; add a cost term to the cost function to guide the steady state of the internal model; Step 4: Based on the multi-energy system model, the internal model system is used as the augmentation term to obtain the first augmented system; then, by utilizing the solvability of the regulator equation and combining coordinate transformation, the first augmented system is transformed into a second augmented system with the same control strategy gain. Step 5: Based on the cost function, using measurable multi-energy system data, perform the first stage of iterative learning on the second augmented system to obtain a stable control strategy; Step 6: Based on the cost function and stable control strategy, perform the second stage of iterative learning to obtain the control strategy gain estimate of each energy supplier in the multi-energy system; substitute the control strategy gain estimate and the internal model parameter value determined in Step 2 into the feedback scheduling strategy model to obtain the control scheduling strategy of each energy supplier.

2. The data-driven non-zero-sum game-based multi-energy system scheduling method as described in claim 1, characterized in that, In step 1, the feedback scheduling strategy model is as follows: in, In the formula, and Represents discrete time points; For the first i Control and scheduling strategies for individual energy suppliers; For multi-energy systems The system state at any given moment; For system tracking output of multi-energy systems; Indicates the internal model system; and For the first i The control strategy gain of each energy supplier; and For the internal mold parameters of the internal mold system; , , and These are parameters to be determined.

3. The data-driven non-zero-sum game-based multi-energy system scheduling method as described in claim 1, characterized in that, In step 2, the characteristic polynomial of the external system system matrix is ​​constructed by measuring the state data of the external system as follows: measuring the state data of the external system, when the state data of the external system meets the set initial excitation conditions, the external system system matrix is ​​updated by parameter identification based on the filtering equation, and the characteristic polynomial is constructed based on the external system system matrix after parameter identification.

4. The data-driven non-zero-sum game-based multi-energy system scheduling method as described in claim 1, characterized in that, In step 2, the internal model parameter value is determined using the coefficients of the characteristic polynomial: The characteristic polynomial of the external system matrix is ​​expressed as: ; in, These are the eigenvalues ​​of the external system's system matrix; For the external system state dimension; The coefficients of the characteristic polynomial; The internal model system is designed using the coefficients of the characteristic polynomial as follows: in, , ,and , In the formula, Represents discrete time points; Indicates the internal model system; For system tracking output of multi-energy systems; and These are the parameters of the internal mold; This represents the Kronecker product operation; for n A unit vector of dimension; n It is the system state dimension; Substituting the characteristic polynomial coefficients obtained from the external system state data into the internal model system yields the internal model parameters. and The value of .

5. The data-driven non-zero-sum game-based multi-energy system scheduling method as described in claim 1, characterized in that, In step 3, the cost function is designed as the sum of the cooperation cost term and the competition cost term, wherein the cooperation cost term is added to the cost term guiding the steady state of the internal model; minimizing the cooperation cost term aims to achieve the closed-loop stability and output tracking of the multi-energy system through cooperation, while minimizing the competition cost term aims to achieve the Nash equilibrium among multiple energy suppliers; by solving the minimization cost function for each energy supplier, a non-zero-sum game multi-energy system scheduling is achieved.

6. The data-driven non-zero-sum game-based multi-energy system scheduling method as described in claim 3, characterized in that, The first stage of iterative learning uses a value iteration algorithm; The second stage of iterative learning employs a policy iterative algorithm; a switching iterative criterion is used to determine whether a stable control policy has been obtained. If so, then switch to the second stage of iterative learning; The switching iteration criterion is: calculating the eigenvalues ​​of the multi-energy system closed-loop system matrix based on the current iteration result; When the eigenvalue is less than 0, a stable control strategy is determined.

7. The data-driven non-zero-sum game-based multi-energy system scheduling method as described in claim 2, characterized in that, The multi-energy system model is as follows: In the formula, For multi-energy systems The system state at any given moment. For the first i Control and scheduling strategies for individual energy suppliers For external system status, For system tracking output of multi-energy systems; , , , , Let N be the unknown system parameter matrix, and N be the number of energy providers included in the multi-energy system. The second augmented system obtained from step 4 is: In the formula, ; , , To stabilize the system state, To stabilize the internal mold state; For vector merging operators; , To stabilize the control and scheduling strategy; , .

8. A data-driven non-zero-sum game-based multi-energy system scheduling device, characterized in that, include: The model building module is used to design a feedback scheduling strategy model based on internal model dynamics and output feedback, based on a multi-energy system model and an external system model with unknown system parameters, where the control strategy gain and the internal model parameters of the internal model system are unknown. The internal model parameter determination module is used to construct the characteristic polynomial of the external system matrix by measuring the state data of the external system, and to determine the internal model parameter values ​​using the coefficients of the characteristic polynomial. The model conversion module is used to convert the multi-energy system model into an augmented system based on the multi-energy system model, using the internal model system as an augmented term, and combining the solvability of the regulator equation; The first optimization module is used to perform the first stage of iterative learning on the augmented system based on the cost function and using measurable multi-energy system data to obtain a stable control strategy. The cost function incorporates a cost term to guide the steady state of the internal model; The second optimization module is used to perform a second-stage iterative learning based on the cost function and the stable control strategy to obtain the estimated control strategy gain of each energy supplier in the multi-energy system. The strategy generation module is used to substitute the estimated value of the control strategy gain and the value of the internal model parameter determined by the internal model parameter determination module into the feedback scheduling strategy model to obtain the control scheduling strategy of each power supplier.

9. The data-driven non-zero-sum game-based multi-energy system scheduling device as described in claim 8, characterized in that, The feedback scheduling strategy model constructed by the model building module is as follows: in, In the formula, and Represents discrete time points; For the first i Control and scheduling strategies for individual energy suppliers; The system state of a multi-energy system; For system tracking output of multi-energy systems; Indicates the internal model system; and For the first i The control strategy gain of each energy supplier; and For the internal mold parameters of the internal mold system; , , and These are unknown parameters to be designed. The characteristic polynomial of the external system matrix constructed by the internal model parameter determination module is: in, These are the eigenvalues ​​of the external system's system matrix; For the external system state dimension; The coefficients of the characteristic polynomial; The internal model system is designed using the coefficients of the characteristic polynomial as follows: in, , ,and In the formula, This represents the Kronecker product operation; for n A unit vector of dimension; n It is the system state dimension; Substituting the characteristic polynomial coefficients obtained from the external system state data into the internal model system yields the internal model parameters. and The value of .

10. The data-driven non-zero-sum game-based multi-energy system scheduling device as described in claim 8, characterized in that, The first optimization module is further used to determine whether a stable control strategy has been obtained by using the switching iteration criterion; If so, switch to the second optimization module for the second stage of iterative learning; The switching iteration criterion is: calculating the eigenvalues ​​of the multi-energy system closed-loop system matrix based on the current iteration result; When the eigenvalue is less than 0, a stable control strategy is determined.

Citation Information

Patent Citations

  • Multi-individual optimization control method based on non-strategy Q learning

    CN110083063A

  • Integrated energy system autonomous scheduling method based on non-cooperative game

    CN115423348A