Distributed optical storage dynamic game optimization design method, system, equipment and medium
By constructing a node model of a distributed photovoltaic-storage system and introducing parameterized quantum circuits and a meta-learning framework, the problems of slow convergence and poor generalization ability of photovoltaic-storage systems in dynamic environments are solved, and fast-response and adaptive distributed photovoltaic-storage system optimization is achieved.
Patent Information
- Application Number
- CN202511422190.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-30
- Publication Date
- 2026-02-10
AI Technical Summary
Existing optimization strategies for optical storage systems are slow to converge and have poor generalization ability when facing task migration and new scenario changes. They are difficult to adapt to dynamic distributed game scenarios, and traditional methods have high communication pressure and slow response speed.
A node model of a distributed photovoltaic storage system is constructed, and a parameterized quantum circuit is introduced as a policy generator. Combined with a meta-learning framework, parameter updates driven by local game theory and payoff bias are used to achieve rapid optimization of the distributed photovoltaic storage system.
It improves the convergence speed and global coordination of strategy generation, enhances the system's adaptability and responsiveness in dynamic environments, and ensures that the system can quickly approach the optimal operating state within a limited scheduling cycle.
Smart Images

Figure CN121503196A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of photovoltaic energy storage system technology, and in particular to a distributed photovoltaic energy storage dynamic game optimization design method, system, equipment and medium. Background Technology
[0002] With the rapid increase in the penetration rate of new energy sources, distributed photovoltaic and energy storage systems are widely deployed at the end of the power grid, realizing a flexible operation mode of local power generation and local consumption. However, existing scheduling strategies mostly rely on centralized optimization, which has high computational pressure, slow response speed, and difficulty in dealing with incomplete information and system uncertainty in game-theoretic environments. In addition, traditional optimization strategies generally suffer from slow convergence and poor generalization when facing task migration and new scenario changes, making them difficult to adapt to dynamic distributed game scenarios. Quantum computing has natural parallelism and high-dimensional space representation capabilities, making it suitable for high-complexity optimization searches. Meta-learning, as a method of "learning how to learn," can quickly obtain near-optimal strategies in different scheduling scenarios. However, there is currently a lack of research on combining the two for game-theoretic scheduling of photovoltaic and energy storage systems. Summary of the Invention
[0003] In view of the aforementioned existing problems, the present invention is proposed.
[0004] Therefore, this invention provides a distributed optical storage dynamic game optimization design method, system, device, and medium to solve the problems of slow convergence and poor generalization ability of existing optimization strategies when facing task migration and new scenario changes.
[0005] To solve the above-mentioned technical problems, the present invention provides the following technical solution:
[0006] In a first aspect, the present invention provides a distributed optical storage dynamic game-theoretic optimization design method, comprising:
[0007] Obtain the photovoltaic module parameters, energy storage unit characteristic parameters, and local load data of each distributed node, and construct a node model of the distributed photovoltaic-storage system in combination with grid operation constraints and control objectives;
[0008] Based on the node model of the distributed optical storage system, a dynamic game model of incomplete information between nodes is constructed by defining the local state interaction, strategy space coupling and payoff function dependency between nodes.
[0009] Based on the incomplete information dynamic game model, a parameterized quantum circuit is introduced as a strategy generator, and the local game strategy is obtained by utilizing the properties of quantum superposition and entanglement.
[0010] Based on the local game strategy, by adopting the model-independent meta-learning framework, the inner loop adaptation and outer loop optimization are performed on multiple historical scheduling tasks to obtain the optimal initial parameter set;
[0011] Based on the optimal initial parameter set, each node initializes the parameterized quantum circuit and generates an initial strategy. Through parallel game and payoff bias-driven parameter updates, the system converges to equilibrium after multiple rounds of iteration, thus obtaining the optimal game strategy within the current scheduling period.
[0012] As a preferred embodiment of the distributed optical-storage dynamic game-theoretic optimization design method described in this invention, the step of constructing the distributed optical-storage system node model includes:
[0013] Each operating unit in the distributed photovoltaic-storage system is abstracted as an autonomous intelligent agent node consisting of photovoltaic power generation equipment, energy storage unit and game optimization controller, and a collaborative scheduling network is built through the interconnection of multiple nodes;
[0014] A photovoltaic power output model is established based on the photovoltaic equipment parameters and real-time irradiance conditions of the intelligent agent nodes, and the photovoltaic power output of each node at different times is calculated.
[0015] Based on the electrical characteristics of the energy storage unit, an energy storage dynamic model with the state of charge as the core is established to calculate the energy storage charging and discharging power, and set the upper and lower limits of the energy storage charging and discharging power and the range of the state of charge value.
[0016] Based on photovoltaic power output and energy storage charging and discharging power, combined with local load demand, nodal power balance equations are constructed to calculate grid interaction power;
[0017] Based on grid interaction power and energy storage charging and discharging power, an economic objective function is defined.
[0018] By integrating the photovoltaic output model, energy storage dynamic model, power balance equation and economic objective function, as well as the state input and control output definitions of the game-theoretic optimization controller, a node model of a distributed photovoltaic-energy storage system is constructed.
[0019] As a preferred embodiment of the distributed optical storage dynamic game optimization design method of the present invention, wherein: the construction of the incomplete information dynamic game model among nodes includes:
[0020] Map the economic objective function of each autonomous intelligent agent node to a game payoff function;
[0021] Based on the defined state input vector and control policy output vector, the strategy space of each node in the game is defined.
[0022] Based on the policy space, a continuous decision sequence of each node at multiple time steps is established to form a dynamic decision structure.
[0023] Based on the dynamic decision-making structure, each node is set to only be able to obtain local state and neighbor information, thus constructing a partially visible information structure.
[0024] Based on the game payoff function, strategy space, dynamic decision structure, and information structure, a dynamic game model with incomplete information between nodes is constructed.
[0025] As a preferred embodiment of the distributed optical storage dynamic game optimization design method of the present invention, wherein obtaining the local game strategy includes:
[0026] The state vector of the agent node is converted into angle parameters through a normalized mapping function, and the initial quantum state is constructed based on the angle parameters;
[0027] Based on the initial quantum state, a parameterized quantum circuit consisting of multiple rotating gates and entanglement gates is constructed;
[0028] Based on parameterized quantum circuits and initial quantum states, a strategy probability distribution is generated through quantum measurement;
[0029] Based on the policy probability distribution and the game payoff function, negative expected payoff is defined as the loss function, and the policy gradient algorithm is used to update the parameters of the parameterized quantum circuit.
[0030] By embedding optimized parameterized quantum circuits into a game optimization controller, a local game strategy is obtained.
[0031] As a preferred embodiment of the distributed optical-storage dynamic game-theoretic optimization design method of the present invention, wherein obtaining the optimal initial parameter set includes:
[0032] Construct a historical task set containing multiple typical scheduling scenarios;
[0033] Define a meta-learning optimization function;
[0034] Based on the meta-learning optimization function, a model-independent meta-learning algorithm is adopted. The initial parameters are iteratively optimized by alternately executing the inner loop task and the outer loop meta-update.
[0035] When the initial parameters converge, the optimal set of initial parameters is obtained;
[0036] As a preferred embodiment of the distributed optical storage dynamic game optimization design method of the present invention, wherein obtaining the optimal game strategy within the current scheduling period includes:
[0037] Define the joint objective function of the system;
[0038] The parameterized quantum circuits of each node are initialized based on the optimal initial parameter set obtained from meta-learning training.
[0039] Each node inputs its real-time sensed local state vector into the parameterized quantum circuit to generate an initial policy distribution;
[0040] All nodes participate in the dynamic game in parallel, and the game equilibrium mechanism is selected according to the type of interaction relationship between nodes to carry out strategy interaction;
[0041] Based on the profit deviation generated by the strategy execution result, a feedback-driven parameter optimization mechanism is executed to update the adjustable parameters of the parameterized quantum circuit in reverse.
[0042] By repeatedly executing strategy generation, game interaction, and parameter updates, the optimal game strategy for the current scheduling period is finally output.
[0043] The beneficial effects of this preferred technical solution are that by setting a joint objective function for the system and combining it with meta-learning to initialize parameterized quantum circuits, collaborative optimization based on local perception and distributed game theory is achieved. The parameter update is carried out using a feedback mechanism driven by payoff bias, which significantly improves the convergence speed and global coordination of strategy generation.
[0044] As a preferred embodiment of the distributed optical-storage dynamic game-theoretic optimization design method of the present invention, the feedback-driven parameter optimization mechanism includes:
[0045] In the distributed game process, each node executes an action based on the strategy output by the current parameterized quantum circuit and obtains the actual payoff value.
[0046] Calculate the deviation between the actual return and the target return to obtain the return deviation.
[0047] The aforementioned profit bias is used as a gradient signal for reinforcement learning, and the adjustable parameters in the parameterized quantum circuit are adjusted through the backpropagation algorithm.
[0048] The steps of calculating the profit deviation and adjusting the adjustable parameters constitute a parameter update. The parameter update is iterated alternately with other policy execution steps in each scheduling cycle until the policy converges.
[0049] The beneficial effects of this preferred technical solution are that by dynamically adjusting the parameterized quantum circuit using the payoff deviation as a gradient signal, a closed-loop iteration of strategy generation and parameter optimization is achieved, which significantly improves the adaptability and convergence speed of the strategies of each node in the distributed game environment, and ensures that the system quickly approaches the optimal operating state within a finite scheduling period.
[0050] Secondly, the present invention provides a distributed optical storage dynamic game optimization design system, comprising:
[0051] The node modeling module is used to obtain the photovoltaic module parameters, energy storage unit characteristic parameters and local load data of each distributed node, and to construct the node model of the distributed photovoltaic-storage system in combination with grid operation constraints and control objectives.
[0052] The game modeling module is used to construct a dynamic game model with incomplete information between nodes based on the node model of a distributed optical storage system by defining the local state interaction, strategy space coupling and payoff function dependency between nodes.
[0053] The quantum strategy generation module is used to introduce parameterized quantum circuits as a strategy generator based on the dynamic game model with incomplete information, and to obtain local game strategies by utilizing the properties of quantum superposition and entanglement.
[0054] The meta-learning parameter optimization module is used to obtain the optimal initial parameter set by performing inner loop adaptation and outer loop optimization on multiple historical scheduling tasks based on local game strategy and by adopting a model-independent meta-learning framework.
[0055] The real-time game execution module is used to initialize parameterized quantum circuits and generate initial strategies for each node based on the optimal initial parameter set. Through parallel game and parameter updates driven by payoff deviation, the system converges to equilibrium after multiple rounds of iteration, thus obtaining the optimal game strategy within the current scheduling period.
[0056] Thirdly, the present invention provides an electronic device, comprising:
[0057] Memory, used to store programs;
[0058] A processor is configured to execute the computer-executable instructions, which, when executed by the processor, implement the steps of the distributed optical storage dynamic game optimization design method.
[0059] Fourthly, the present invention provides a computer-readable storage medium, comprising: when the program is executed by a processor, the step of implementing the distributed optical storage dynamic game optimization design method.
[0060] The beneficial effects of this invention are as follows: By constructing a node model that integrates equipment characteristics and power grid constraints, this invention achieves a precise characterization and personalized optimization basis for the operating behavior of each node; by establishing a non-complete information game model with local state interaction and payoff dependence, it achieves low-communication collaborative decision-making under a decentralized architecture; by introducing parameterized quantum circuits as a policy generator, it utilizes quantum properties to efficiently express high-dimensional nonlinear policy spaces; by extracting common knowledge across multiple tasks through a meta-learning framework, it obtains optimal initial parameters with cross-scenario generalization capabilities, achieving rapid convergence in new environments; by using a local state-driven parallel game execution mechanism, it achieves distributed optimal policy generation under the condition of satisfying local Nash equilibrium; and by using local parameter gradient fine-tuning based on payoff deviation, it achieves online adaptive policy updates, improving the system's responsiveness and robustness to dynamic disturbances. Attached Figure Description
[0061] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Wherein:
[0062] Figure 1 This is a schematic diagram of the basic process of a distributed optical-storage dynamic game optimization design method provided in one embodiment of the present invention. Detailed Implementation
[0063] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of the present invention.
[0064] Example 1, referring to Figure 1 As an embodiment of the present invention, a distributed optical storage dynamic game optimization design method is provided, such as... Figure 1 As shown, it includes:
[0065] S100: Obtain the photovoltaic module parameters, energy storage unit characteristic parameters and local load data of each distributed node, and construct a distributed photovoltaic-storage system node model in combination with grid operation constraints and control objectives;
[0066] S200: Based on the node model of a distributed optical storage system, a dynamic game model with incomplete information between nodes is constructed by defining the local state interaction, strategy space coupling and payoff function dependency between nodes.
[0067] S300: Based on the incomplete information dynamic game model, a parameterized quantum circuit is introduced as a strategy generator, and the local game strategy is obtained by utilizing the properties of quantum superposition and entanglement.
[0068] S400: Based on a local game strategy, it uses a model-independent meta-learning framework to perform inner loop adaptation and outer loop optimization on multiple historical scheduling tasks to obtain the optimal initial parameter set.
[0069] S500: Based on the optimal initial parameter set, each node initializes the parameterized quantum circuit and generates an initial strategy. Through parallel game and payoff bias-driven parameter updates, it converges to equilibrium after multiple rounds of iteration, obtaining the optimal game strategy within the current scheduling period.
[0070] It should be noted that existing optimization strategies for distributed photovoltaic energy storage systems largely rely on centralized coordination or preset control rules based on static environmental assumptions, which face numerous challenges in actual operation. On the one hand, traditional centralized optimization methods depend on global information collection and centralized decision-making, resulting in heavy communication burdens and high response latency, making it difficult to adapt to the high penetration and highly random distributed energy access demands. On the other hand, existing distributed control strategies are mostly based on fixed models or empirical rules, lacking online adaptability to dynamic disturbances such as light fluctuations and load abrupt changes, and are prone to performance degradation due to model mismatch. Furthermore, traditional reinforcement learning methods suffer from low policy search efficiency, slow convergence speed, and initialization sensitivity in complex nonlinear environments, and are difficult to achieve rapid transfer in new tasks, resulting in high deployment costs and poor stability. Therefore, existing methods still have significant shortcomings in decentralized collaboration, rapid response, online adaptation, and cross-scenario generalization capabilities.
[0071] Therefore, to address the issues of slow convergence and poor generalization ability of existing optimization strategies when facing task migration and new scenario changes, the S100-S500 steps are used to construct an online adaptive game mechanism driven by localized modeling, quantum strategy generation, and meta-learning. This enables the distributed optical storage system to respond quickly, coordinately optimize, and continuously evolve in the face of dynamic environments with low communication dependence.
[0072] Example 2, this is an embodiment of the present invention, which provides a distributed optical storage dynamic game optimization design method based on the previous embodiment, including:
[0073] In this embodiment of the application, step S100, which involves constructing a distributed optical storage system node model, includes:
[0074] Each operating unit in the distributed photovoltaic-storage system is abstracted as an autonomous intelligent agent node consisting of photovoltaic power generation equipment, energy storage unit and game optimization controller, and a collaborative scheduling network is built through the interconnection of multiple nodes;
[0075] A photovoltaic power output model is established based on the photovoltaic equipment parameters and real-time irradiance conditions of the intelligent agent nodes, and the photovoltaic power output of each node at different times is calculated.
[0076] Based on the electrical characteristics of the energy storage unit, an energy storage dynamic model with the state of charge as the core is established to calculate the energy storage charging and discharging power, and set the upper and lower limits of the energy storage charging and discharging power and the range of the state of charge value.
[0077] Based on photovoltaic power output and energy storage charging and discharging power, combined with local load demand, nodal power balance equations are constructed to calculate grid interaction power;
[0078] Based on grid interaction power and energy storage charging and discharging power, an economic objective function is defined.
[0079] By integrating the photovoltaic output model, energy storage dynamic model, power balance equation and economic objective function, as well as the state input and control output definitions of the game-theoretic optimization controller, a node model of a distributed photovoltaic-energy storage system is constructed.
[0080] In this embodiment, the intelligent agent node consists of a photovoltaic power generation device, an energy storage unit, and a game optimization controller;
[0081] In this embodiment of the application, the photovoltaic output model is expressed as a product model based on effective irradiance:
[0082] P pv,i (t)=η pv ·A i ·G i (t)
[0083] In the formula, P pv,i (t) represents the photovoltaic output of the i-th node at time t; η pv A represents the photovoltaic efficiency coefficient. i G represents the area of the photovoltaic module. i (t) represents the solar irradiance per unit area.
[0084] In this embodiment of the application, the energy storage dynamic model with state of charge (SOC) as its core is represented as follows:
[0085]
[0086] In the formula, SOC i (t), SOC i (t+1) represents the energy storage state at times t and t+1, and its range is [0,1]. This refers to the charging power. E represents the discharge power. max,i η is the rated capacity of the energy storage. ch η dis Δt represents the energy storage charge / discharge efficiency, and Δt represents the time step.
[0087] And set energy storage operation constraints:
[0088] Charging power constraints:
[0089]
[0090] Discharge power constraint:
[0091]
[0092] SOC range constraints:
[0093] SOC min ≤SOC i (t)≤SOCmax
[0094] In the formula, P ess,max This represents the maximum energy storage capacity; SOC max SOC min These represent the maximum and minimum SOC values for energy storage, respectively.
[0095] In this embodiment, the node power balance equation is constructed as follows:
[0096]
[0097] In the formula, D i (t) represents the local load demand; P pv,i (t) represents the photovoltaic output of the i-th node at time t; This refers to the charging power. P is the discharge power; grid,i (t) represents the power purchased from or fed to the grid.
[0098] In this embodiment of the application, the economic objective function is expressed as:
[0099]
[0100] In the formula, λ i (t) represents the electricity price; α represents the weight of energy storage operating costs; P grid,i (t) represents the power purchased from or fed into the grid; This refers to the charging power. This represents the discharge power.
[0101] In this embodiment of the application, the state input of the game optimization controller is represented as follows:
[0102] x i (t)=[P pv,i (t),D i (t),SOC i (t),λ i (t)]
[0103] In this embodiment of the application, the control output of the game optimization controller is expressed as:
[0104]
[0105] In the formula, P pv,i (t) represents photovoltaic output; D i (t) represents local load demand; SOC i (t) represents the energy storage state; λ i (t) represents the local electricity price signal; P grid,i (t) represents the power purchased from or fed into the grid; This refers to the charging power. This represents the discharge power.
[0106] In this embodiment of the application, step S200, which involves constructing a dynamic game model with incomplete information between nodes, includes:
[0107] Map the economic objective function of each autonomous intelligent agent node to a game payoff function;
[0108] Based on the defined state input vector and control policy output vector, the strategy space of each node in the game is defined.
[0109] Based on the policy space, a continuous decision sequence of each node at multiple time steps is established to form a dynamic decision structure.
[0110] Based on the dynamic decision-making structure, each node is set to only be able to obtain local state and neighbor information, thus constructing a partially visible information structure.
[0111] Based on the game payoff function, strategy space, dynamic decision structure, and information structure, a dynamic game model with incomplete information between nodes is constructed.
[0112] In this embodiment of the application, in order to achieve efficient collaborative optimization of the distributed photovoltaic energy storage system under decentralized coordination, a decision-making framework based on incomplete information dynamic game is constructed, which specifically includes system structure modeling, node behavior modeling, game mechanism modeling and power-price coupling modeling.
[0113] In this embodiment of the application, system structure modeling includes assuming the system contains N distributed intelligent nodes, and the node set is denoted as:
[0114] n = {1, 2, ..., N}
[0115] Each node i∈N consists of a local state, a policy space, and a payoff function, and participates in the game and makes independent decisions.
[0116] In this embodiment of the application, node behavior modeling includes:
[0117] a. State space
[0118] Each node has the following local state at any time t:
[0119] x i (t)=[P pv,i (t),D i (t),SOC i (t),λ i (t)]
[0120] In the formula, P pv,i (t) represents photovoltaic output; D i (t) represents the load demand; SOCi (t) represents the energy storage state; λ i (t) represents the local electricity price signal.
[0121] b. Strategy Space
[0122] The game strategy of a node is defined as follows:
[0123]
[0124] In this embodiment of the application, the game mechanism modeling includes:
[0125] a. The game type is a dynamic non-cooperative game.
[0126] The goal of each node is:
[0127] Given other node strategies -i In such cases, minimize local operating costs or maximize local benefits.
[0128] b. Node payment function
[0129] The revenue function for the i-th node is defined as follows:
[0130] u i (s i ,s -i ) = R i (s i )-C i (s i )-ψ·φ i (s i ,s -i )
[0131] In the formula, R i (s i C represents local electricity sales revenue; i (s i ) refers to the energy consumption cost of energy storage, electricity price cost, etc.; φ i (S i ,S -i ) represents the penalty cost for system stability or power deviation; ψ represents the penalty weighting factor.
[0132] c. Modeling of incomplete information structures
[0133] Each node can only perceive its own information and cannot obtain the complete state of other nodes. -i Or strategy S -i .
[0134] Its observable information is as follows:
[0135]
[0136] In the formula, x i (t) represents the state vector of other nodes; This is an estimated price signal.
[0137] d. Dynamic game theory
[0138] The game is played in discrete time intervals t = 1, ..., T, with the objective of maximizing the long-run cumulative payoff.
[0139]
[0140] In the formula, E is the expected value under the strategy distribution; γ is the discount factor, ranging from [0,1]; u i Let s be the profit function for the i-th node; i (t) represents the game strategy of the node; S -i (t) represents the game strategy of other nodes.
[0141] In this embodiment of the application, power-price coupling modeling includes:
[0142] a. Power complementarity constraint
[0143] In some cases, power complementary channels can be constructed between nodes, and their complementary relationships are as follows:
[0144]
[0145] In the formula, P grid,i (t) represents the power purchased from or fed into the grid; P pv,i (t) represents photovoltaic output; D i (t) represents the local load demand; This refers to the charging power. N represents the discharge power. i P is the set of neighboring nodes that are connected to node i. ij (t) represents the energy output from node j to node i.
[0146] b. Electricity Price Game
[0147] If nodes are allowed to quote prices independently, then each node's strategy includes a pricing decision:
[0148] λ i (t)∈[λ min ,λ max ]
[0149] In the formula, λ i (t) represents the local electricity price signal; λ max , λ min These represent the maximum and minimum electricity prices, respectively.
[0150] The game will form a Stackelberg game framework, with the system's market liquidation price and energy allocation derived from the game solution. The strategy evolution and optimization mechanism is as follows: In each round of the game, a node updates its strategy s based on historical payoffs and state. i (t) This invention uses quantum search and meta-learning to explore and migrate the policy space (see S300 and S400), enabling nodes to rapidly approach Nash equilibrium in a finite number of game rounds.
[0151] In this embodiment of the application, the local game strategy obtained in step S300 includes:
[0152] The state vector of the agent node is converted into angle parameters through a normalized mapping function, and the initial quantum state is constructed based on the angle parameters;
[0153] Based on the initial quantum state, a parameterized quantum circuit consisting of multiple rotating gates and entanglement gates is constructed;
[0154] Based on parameterized quantum circuits and initial quantum states, a strategy probability distribution is generated through quantum measurement;
[0155] Based on the policy probability distribution and the game payoff function, negative expected payoff is defined as the loss function, and the policy gradient algorithm is used to update the parameters of the parameterized quantum circuit.
[0156] By embedding optimized parameterized quantum circuits into a game optimization controller, a local game strategy is obtained.
[0157] In this embodiment, using a parameterized quantum circuit (PQC) as a policy generator involves normalizing and encoding the aggregated state vectors of the local and neighbor states sensed by each node in real time into a quantum initial state, inputting it into a parameterized quantum circuit composed of a rotation gate and an entanglement gate, outputting a policy probability distribution through quantum measurement, and performing gradient updates in conjunction with payoff deviation feedback to achieve decentralized game optimization of the distributed optical storage system.
[0158] In an optional implementation, in step S300, the policy generator may also select a deep neural network policy network, input the local and neighbor aggregated state vectors perceived in real time by each node into the pre-trained deep neural network policy model, and the model outputs the policy probability distribution or directly generates control actions, and updates the parameters through the profit deviation feedback mechanism.
[0159] In an optional implementation, in step S300, the strategy generator can also be a fuzzy logic controller, which takes the local and neighbor aggregated state vectors perceived by each node in real time as input, performs reasoning through preset membership functions and fuzzy rule base, generates fuzzy output of control actions, and obtains specific charging and discharging power and grid interaction instructions after defuzzification, and performs strategy evaluation and rule optimization based on actual benefit deviation.
[0160] In this embodiment of the application, based on the angle parameter φ k The initial quantum state is represented as:
[0161]
[0162] In the formula, ψ input Let φ be the initial angle of the mapped quantum state; RY(φ) k ) is a revolving door; x i (t) is the state vector; f k This is the normalized mapping function;
[0163] In this embodiment, the parameterized quantum circuit (PQC) composed of multiple rotating gates and entanglement gates is represented as follows:
[0164]
[0165] In the formula, L is the circuit depth (number of layers); n is the number of qubits; θ l,j Let RY(θ) be the rotation angle of the j-th bit in the l-th layer. l,j ) represents the rotating door at the rotation angle of the j-th bit in the l-th layer, CNOT j,j+1 Let j be the entanglement gate for the j-th bit.
[0166] In this embodiment of the application, generating a strategy probability distribution through quantum measurement includes setting the initial quantum state |ψ input The output state U(θ)|ψ is obtained by inputting parameterized quantum circuit U(θ). input >;
[0167] The output state is measured to obtain the policy s. i Conditional probability in a given state:
[0168] π i (s i |x i ,θ)=| i |U(θ)|ψ input >| 2
[0169] In the formula, π i For the policy output distribution; s i For game strategy; x i For node states; U(θ) is the parameterized quantum circuit (PQC); ψ input Let be the initial angle of the mapped quantum state.
[0170] In this embodiment of the application, negative expected return is defined as the loss function:
[0171]
[0172] In the formula, This represents the expected return.
[0173] Update the parameter θ using the policy gradient algorithm:
[0174]
[0175] In the formula, θ is the rotation angle; α is the update factor; Let θ be the gradient; L(θ) is the loss function for maximizing the profit.
[0176] In this embodiment, the PQC module is integrated into the game optimization controller defined in S100; the controller execution flow is as follows:
[0177]
[0178] In this embodiment, the PQC embedded strategy controller control flow includes: the system reading the state vector of the current node; quantum encoding the state information to convert it into an input form suitable for quantum processing; using parameterized quantum circuits (PQC), generating a probability distribution of the strategy by adjusting the angle of the rotating gate; sampling specific decision actions from this probability distribution based on quantum measurement results; and using the sampled decision output to control the energy storage charging and discharging power and the power interaction with the grid, thereby achieving intelligent and decentralized control of the distributed photovoltaic-storage system. The entire process demonstrates the application of quantum computing in strategy generation, combined with classical control logic to complete closed-loop decision-making.
[0179] In this embodiment, the policy search mechanism integrating parameterized quantum circuits (PQC) includes input state vectors being encoded, and quantum operations performed through rotation gates and entanglement gates to generate a policy distribution. This process utilizes quantum superposition and interference effects to enhance policy exploration capabilities, and achieves efficient generation of the policy distribution by controlling the quantum circuitry through an adjustable parameter θ.
[0180] In this embodiment of the application, obtaining the optimal initial parameter set in step S400 includes:
[0181] Construct a historical task set containing multiple typical scheduling scenarios;
[0182] Define a meta-learning optimization function;
[0183] Based on the meta-learning optimization function, a model-independent meta-learning algorithm is adopted. The initial parameters are iteratively optimized by alternately executing the inner loop task and the outer loop meta-update.
[0184] When the initial parameters converge, the optimal set of initial parameters is obtained;
[0185] In this embodiment, the Model-Independent Meta-Learning (MAML) framework is used as the meta-learning initialization mechanism. This includes constructing a task set based on multiple historical scheduling tasks, performing inner loop gradient updates on each task to evaluate adaptability, and optimizing shared initial parameters through an outer loop to finally obtain the optimal initial parameter set that can be quickly transferred to new tasks. This set is then used for policy generation and initialization of parameterized quantum circuits at each node.
[0186] In an optional implementation, the meta-learning initialization mechanism in step S400 can also be based on a pre-trained model of feature transfer, extract the environmental features of the current scheduled task, map them to the embedding space through the task encoder, retrieve similar historical tasks, and load the optimal parameters of their pre-trained models as initial parameters.
[0187] In an optional implementation, the meta-learning initialization mechanism in step S400 can also be based on Bayesian optimization of prior parameter selection. According to the environmental characteristics of the current task, the optimal parameter distribution is predicted from the historical task performance database using a Gaussian process regression model, and the initial parameter with the highest expected performance is selected as the starting point of each node's policy generator through Bayesian optimization.
[0188] In this embodiment of the application, a task set T = {T1, T2, ..., T...} containing multiple historical tasks is defined. M}, where each task T i Includes state-action pairs Node revenue function u i And the PQC structure U(θ).
[0189] Based on multiple historical task sets, the optimization objective of meta-learning is defined as follows:
[0190]
[0191] In the formula, For task T i The policy loss function (such as negative expected return) is given, where α is the inner loop learning rate; Indicates in task T i The parameters are updated after the first gradient update.
[0192] In this embodiment of the application, inner loop task adaptation includes adapting to each task T. i ∈T, perform one gradient descent under the current shared parameters θ:
[0193]
[0194] Obtain the task-specific updated parameter θ′ i .
[0195] In this embodiment, the outer loop element update includes the updated parameter θ′. i Evaluate performance and backpropagate to update shared parameters:
[0196]
[0197] In the formula, β is the meta-update learning rate.
[0198] Repeat the inner and outer loop iterations until the parameter θ converges, obtaining the final optimal initial parameter set:
[0199]
[0200] The θ * It contains generalized knowledge across tasks and will be used as the initialization parameter for generating PQC strategy in subsequent online scheduling.
[0201] When the system faces a new scheduling task θ′ new At that time, with θ * Adapt quickly from the starting point:
[0202]
[0203] In the formula, θ′ new These are the initialization parameters for the new task; α is the learning rate within the task. Let θ be the gradient. For task T mew The PQC strategy loss function is as follows.
[0204] In this embodiment, the embedded structure of the meta-learning module in the parameterized quantum circuit (PQC) policy generator includes training to obtain the optimal initialization parameters θ starting from a multi-task dataset through a meta-learning modeling framework (using a MAML structure). * This parameter is then fed into a parameterized quantum circuit (PQC) as the starting point for training each node, ultimately generating a policy distribution for a specific task. This design ensures that the PQC can quickly adapt to new tasks after a few updates, significantly improving the efficiency and response speed of policy generation.
[0205] In this embodiment of the application, step S500, which obtains the optimal game strategy for each node within the current scheduling period, includes:
[0206] Define the joint objective function of the system;
[0207] The parameterized quantum circuits of each node are initialized based on the optimal initial parameter set obtained from meta-learning training.
[0208] Each node inputs its real-time sensed local state vector into the parameterized quantum circuit to generate an initial policy distribution;
[0209] All nodes participate in the dynamic game in parallel, and the game equilibrium mechanism is selected according to the type of interaction relationship between nodes to carry out strategy interaction;
[0210] Based on the profit deviation generated by the strategy execution result, a feedback-driven parameter optimization mechanism is executed to update the adjustable parameters of the parameterized quantum circuit in reverse.
[0211] Through a multi-round iterative strategy generation process that converges to equilibrium, the optimal game strategy for the current scheduling period is finally output.
[0212] In this embodiment, the joint objective function of the system is expressed as:
[0213]
[0214] In the formula, λ i (t) represents the nodal electricity price information; This refers to the charging power. P is the discharge power; grid,i (t) represents the power purchased from or fed into the grid; c1 is the energy storage dispatch cost factor; c2 is the frequency deviation penalty factor; Δf i (t) represents the node frequency modulation deviation index.
[0215] In this embodiment, the node closed-loop scheduling execution process includes: the node obtaining its current local running state through state awareness; generating an initial policy based on the parameterized quantum circuit (PQC) initialized by meta-learning; after the policy is output, the node participates in multi-agent game and system coordination scheduling to achieve distributed collaborative decision-making; updating the node state according to the policy execution result, and re-entering the next round of policy generation and game process, forming a continuously iterative and adaptively optimized closed-loop control mechanism.
[0216] In this embodiment of the application, the local state vector includes load demand, photovoltaic output, energy storage status, electricity price information, and neighbor node aggregation information;
[0217] In this embodiment of the application, each distributed optical storage smart node acquires its local state vector during the scheduling period t:
[0218]
[0219] In the formula, D i (t) represents the local load demand; P pv,i (t) represents photovoltaic output; SOC i (t) represents the energy storage state; λ i (t) represents the predicted or published electricity price; This includes the average price and edge node load information.
[0220] Based on the acquired local state vector, it is encoded as the initial angle of the quantum state;
[0221] The quantum state is input into the parameterized quantum circuit U(θ) i ), where θ i These are the initialization parameters optimized for meta-learning.
[0222] Generate the probability distribution of the current game strategy using quantum measurement:
[0223] π i (s i (t))=| i (t)|U(θ i )|ψ(x i (t))>| 2
[0224] In the formula, π i For the policy output distribution; s i (t) represents the game strategy of the node; ψ(x) i (t) represents the state vector x. i The initial angle of the quantum state mapped at time (t); x i (t) is the state vector; U(θ) i ) is a parameterized quantum circuit (PQC).
[0225] Based on the generated policy probability distribution π i (s i (t)), calculate the expected game strategy:
[0226]
[0227] In the formula, E[s i [t] represents the node s i The expected value under the game strategy (t); s is the node; π i (s) represents the policy output distribution at node s.
[0228] The expected game strategy E[s] i [(t)] is mapped to a specific control action vector:
[0229]
[0230] In the formula, E[s i [t] represents the node s i The expected value under the game strategy (t); s is the node; π i (s) represents the policy output distribution at node s.
[0231] All nodes are based on the strategies they generate. i (t) Parallel participation in the game;
[0232] In this embodiment, the system adaptively adopts the corresponding game equilibrium mechanism according to the type of interaction relationship between nodes; when the nodes are of equal status (such as multiple microgrids coordinating autonomously), a local Nash equilibrium is adopted; when there is a scheduling dominance relationship (such as the master station issuing instructions and energy storage responding), the Stackelberg mechanism is adopted; the entire system can dynamically identify the structure type and automatically switch the game mechanism to achieve more flexible and efficient coordination.
[0233] In an optional implementation, the game equilibrium mechanism in step S500 can also adopt a consensus mechanism, in which each node exchanges information with the policy state of its neighboring nodes based on the communication topology in each iteration, updates its own policy through a consensus protocol, and gradually converges to a common policy of group collaboration.
[0234] In an optional implementation, the game equilibrium mechanism in step S500 can also adopt an auction mechanism, where each node, as a bidder, submits its local energy storage charging and discharging capacity and price information, determines the clearing price and energy allocation scheme through distributed or local clearing rules, and adjusts its own control strategy according to the auction results.
[0235] In this embodiment of the application, the local Nash equilibrium is represented as:
[0236]
[0237] In the formula, u i s is the node revenue function; i The game strategy for nodes; This is the optimal strategy; For other node strategies. The optimal payoff obtained by a node under the local Nash equilibrium strategy; s represents the payoff obtained by a node under a local Nash equilibrium strategy; i ∈S i Game-theoretic actions as a strategy;
[0238] In this embodiment, the Stackelberg approximation game involves the leading node making a first decision, the following nodes responding, and solving for the Stackelberg equilibrium.
[0239] In this embodiment of the application, the feedback-driven parameter optimization mechanism described in step S500 includes:
[0240] In the distributed game process, each node executes an action based on the strategy output by the current parameterized quantum circuit and obtains the actual payoff value.
[0241] Calculate the deviation between the actual return and the target return to obtain the return deviation.
[0242] The aforementioned profit bias is used as a gradient signal for reinforcement learning, and the adjustable parameters in the parameterized quantum circuit are adjusted through the backpropagation algorithm.
[0243] The steps of calculating the profit deviation and adjusting the adjustable parameters constitute a parameter update. The parameter update is iterated alternately with other policy execution steps in each scheduling cycle until the policy converges.
[0244] It should be noted that the optimal game strategy generated by this invention is achieved through local observation, local game and limited communication collaboration of each node under a decentralized scheduling architecture. This architecture provides system-level support for online solution and distributed execution of the optimal strategy.
[0245] In this embodiment of the application, the decentralized scheduling architecture includes system structure design, information exchange modeling, decentralized game mechanism, communication load analysis, and mathematical modeling of information exchange constraints and scheduling.
[0246] In this embodiment, the system architecture design includes constructing a decentralized network composed of multiple autonomous intelligent nodes. The node set is N = 1, 2, ..., N, and each node forms an adjacency graph G = (N, ε) through communication links, where (i, j) ∈ ε indicates that node i and node j can exchange information. The topology satisfies graph connectivity but does not require full connectivity, supports sparse communication structures, and reduces network dependencies.
[0247] Each node's operating modules include a local perceptron (for acquiring state); a PQC policy generator; a meta-learning policy transferor; a lightweight communication module (for interacting with neighbors); and a decentralized game executor.
[0248] To reduce communication overhead, each node only obtains information digests from its neighboring nodes. The average price and edge node load information are represented as follows:
[0249]
[0250] In the formula, x j (t) is the state vector; N i Let be the neighbor set of node i; the Aggregate function can be a weighted average; O(d) is the number of communications, where d is the node state dimension.
[0251] In this embodiment, the information aggregation method includes each node obtaining the local state information of its neighboring nodes through lightweight communication, aggregating the neighboring state vectors using a weighted average function, and generating neighbor aggregation information as one of the inputs for generating its own strategy, so as to achieve local collaborative optimization without relying on global information.
[0252] In one alternative implementation, the information aggregation method can also use maximum / minimum aggregation. Each node obtains the local state information of its neighboring nodes through lightweight communication, and uses the maximum or minimum value function to aggregate key state variables (such as electricity price and load demand) to generate extreme value feature information as a strategy input, so as to enhance the response capability to emergency events or extreme scenarios.
[0253] In an alternative implementation, the information aggregation method can also use a graph attention mechanism, where each node obtains neighbor state information through lightweight communication, introduces an attention network to calculate the dynamic weights between itself and each neighbor node, and performs weighted aggregation of neighbor states to generate differentiated feature representations, which are used as policy inputs to highlight the information of high-influence nodes.
[0254] Using a weighted average as the aggregation function:
[0255]
[0256] In the formula, N i Let λ be the set of neighbors of node i; j (t) represents electricity price information.
[0257] Each node independently solves its game strategy based on its local state and aggregated information:
[0258]
[0259] In the formula, s i (t) represents the game strategy of the node; s i ∈S i The strategy is the game action; E is the expected value. Game strategy based on price mean and edge node load information; This determines the actual revenue that a node receives.
[0260] Define the system objective function The global optimum can be approximated under the following conditions: the topology is connected; the policy updates of each node are diminishing; and the information aggregation function is convergent.
[0261] Compared to traditional centralized scheduling, which requires each node to upload its full state x_i(t) and receive the central policy, resulting in a communication complexity of O(N·d), this invention only requires each node to exchange information with its k neighbors, resulting in a communication volume of O(k·d), where k << N. The communication savings are:
[0262]
[0263] In the formula, O(N·d) represents that the amount of communication data consumed increases linearly with the number of nodes N and the state dimension d of each node; o(k·d) represents that the amount of communication data consumed increases linearly with the number of nodes N and the state dimension k of each node.
[0264] Considering actual communication delays, this invention introduces a time lag term in the policy update:
[0265]
[0266] In the formula, T ij For communication latency (ms), asynchronous updates are allowed; For quantum policy function (PQC + meta-learning optimizer); For price average and edge node load information; x i (t) is the state variable.
[0267] Even if there is a delay in neighbor information, nodes can still continue to optimize their strategies based on the latest available data, ensuring robust system operation under non-ideal network conditions.
[0268] Example 3 is an embodiment of the present invention. This embodiment differs from the first embodiment in that it provides a distributed optical storage dynamic game optimization design system.
[0269] It should be noted that the technical solution of the distributed optical-storage dynamic game optimization design system and the technical solution of the above-mentioned distributed optical-storage dynamic game optimization design method belong to the same concept. For details not described in detail in the technical solution of the distributed optical-storage dynamic game optimization design system in this embodiment, please refer to the description of the technical solution of the above-mentioned distributed optical-storage dynamic game optimization design method.
[0270] This embodiment presents a distributed optical-storage dynamic game-theoretic optimization design system, comprising:
[0271] The node modeling module is used to obtain the photovoltaic module parameters, energy storage unit characteristic parameters and local load data of each distributed node, and to construct the node model of the distributed photovoltaic-storage system in combination with grid operation constraints and control objectives.
[0272] The game modeling module is used to construct a dynamic game model with incomplete information between nodes based on the node model of a distributed optical storage system by defining the local state interaction, strategy space coupling and payoff function dependency between nodes.
[0273] The quantum strategy generation module is used to introduce parameterized quantum circuits as a strategy generator based on the dynamic game model with incomplete information, and to obtain local game strategies by utilizing the properties of quantum superposition and entanglement.
[0274] The meta-learning parameter optimization module is used to obtain the optimal initial parameter set by performing inner loop adaptation and outer loop optimization on multiple historical scheduling tasks based on local game strategy and by adopting a model-independent meta-learning framework.
[0275] The real-time game execution module is used to initialize parameterized quantum circuits and generate initial strategies for each node based on the optimal initial parameter set. Through parallel game and parameter updates driven by payoff deviation, the system converges to equilibrium after multiple rounds of iteration, thus obtaining the optimal game strategy within the current scheduling period.
[0276] This embodiment also provides an electronic device applicable to a distributed optical-storage dynamic game optimization design method, including:
[0277] The system includes a memory and a processor. The memory stores computer-executable instructions, and the processor executes these instructions to implement a distributed optical storage dynamic game optimization design method as proposed in the above embodiments.
[0278] This embodiment also provides a storage medium on which a computer program is stored. When the program is executed by a processor, it implements a distributed optical storage dynamic game optimization design method as proposed in the above embodiments.
[0279] The storage medium proposed in this embodiment and the method for implementing a distributed optical storage dynamic game optimization design proposed in the above embodiments belong to the same inventive concept. Technical details not described in detail in this embodiment can be found in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.
[0280] Based on the above description of the implementation methods, those skilled in the art can clearly understand that the present invention can be implemented using software and necessary general-purpose hardware, and of course, it can also be implemented using hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as a computer floppy disk, read-only memory (ROM), random access memory (RAM), flash memory, hard disk, or optical disk, etc., including several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods of the various embodiments of the present invention.
[0281] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A distributed optical-storage dynamic game-theoretic optimization design method, characterized in that, include: Obtain the photovoltaic module parameters, energy storage unit characteristic parameters, and local load data of each distributed node, and construct a node model of the distributed photovoltaic-storage system in combination with grid operation constraints and control objectives; Based on the node model of the distributed optical storage system, a dynamic game model of incomplete information between nodes is constructed by defining the local state interaction, strategy space coupling and payoff function dependency between nodes. Based on the incomplete information dynamic game model, a parameterized quantum circuit is introduced as a strategy generator, and the local game strategy is obtained by utilizing the properties of quantum superposition and entanglement. Based on the local game strategy, a model-independent meta-learning algorithm is used to perform inner loop adaptation and outer loop optimization on multiple historical scheduling tasks to obtain the optimal initial parameter set. Based on the optimal initial parameter set, each node initializes the parameterized quantum circuit and generates an initial strategy. Through parallel game and payoff bias-driven parameter updates, the system converges to equilibrium after multiple rounds of iteration, thus obtaining the optimal game strategy within the current scheduling period.
2. The distributed optical-storage dynamic game-theoretic optimization design method as described in claim 1, characterized in that: The construction of the distributed optical storage system node model includes: Each operating unit in the distributed photovoltaic-storage system is abstracted as an autonomous intelligent agent node consisting of photovoltaic power generation equipment, energy storage unit and game optimization controller, and a collaborative scheduling network is built through the interconnection of multiple nodes; A photovoltaic power output model is established based on the photovoltaic equipment parameters and real-time irradiance conditions of the intelligent agent nodes, and the photovoltaic power output of each node at different times is calculated. Based on the electrical characteristics of the energy storage unit, an energy storage dynamic model with the state of charge as the core is established to calculate the energy storage charging and discharging power, and set the upper and lower limits of the energy storage charging and discharging power and the range of the state of charge value. Based on photovoltaic power output and energy storage charging and discharging power, combined with local load demand, nodal power balance equations are constructed to calculate grid interaction power; Based on grid interaction power and energy storage charging and discharging power, an economic objective function is defined. By integrating the photovoltaic output model, energy storage dynamic model, power balance equation and economic objective function, as well as the state input and control output definitions of the game-theoretic optimization controller, a node model of a distributed photovoltaic-energy storage system is constructed.
3. The distributed optical-storage dynamic game-theoretic optimization design method as described in claim 1 or 2, characterized in that: The imperfect information dynamic game model between the constructed nodes includes: Map the economic objective function of each autonomous intelligent agent node to a game payoff function; Based on the defined state input vector and control policy output vector, the strategy space of each node in the game is defined. Based on the policy space, a continuous decision sequence of each node at multiple time steps is established to form a dynamic decision structure. Based on the dynamic decision-making structure, each node is set to only be able to obtain local state and neighbor information, thus constructing a partially visible information structure. Based on the game payoff function, strategy space, dynamic decision structure, and information structure, a dynamic game model with incomplete information between nodes is constructed.
4. The distributed optical-storage dynamic game optimization design method as described in claim 3, characterized in that: The obtained local game strategy includes: The state vector of the agent node is converted into angle parameters through a normalized mapping function, and the initial quantum state is constructed based on the angle parameters; Based on the initial quantum state, a parameterized quantum circuit consisting of multiple rotating gates and entanglement gates is constructed; Based on parameterized quantum circuits and initial quantum states, a strategy probability distribution is generated through quantum measurement; Based on the policy probability distribution and the game payoff function, negative expected payoff is defined as the loss function, and the policy gradient algorithm is used to update the parameters of the parameterized quantum circuit. By embedding optimized parameterized quantum circuits into a game optimization controller, a local game strategy is obtained.
5. The distributed optical-storage dynamic game-theoretic optimization design method as described in claim 4, characterized in that: The process of obtaining the optimal initial parameter set includes: Construct a historical task set containing multiple typical scheduling scenarios; Define a meta-learning optimization function; Based on the meta-learning optimization function, a model-independent meta-learning algorithm is adopted. The initial parameters are iteratively optimized by alternately executing the inner loop task and the outer loop meta-update. When the initial parameters converge, the optimal set of initial parameters is obtained.
6. The distributed optical-storage dynamic game-theoretic optimization design method as described in claim 5, characterized in that: The process of obtaining the optimal game strategy within the current scheduling period includes: Define the joint objective function of the system; The parameterized quantum circuits of each node are initialized based on the optimal initial parameter set obtained from meta-learning training. Each node inputs its real-time sensed local state vector into the parameterized quantum circuit to generate an initial policy distribution; All nodes participate in the dynamic game in parallel, and the game equilibrium mechanism is selected according to the type of interaction relationship between nodes to carry out strategy interaction; Based on the profit deviation generated by the strategy execution result, a feedback-driven parameter optimization mechanism is executed to update the adjustable parameters of the parameterized quantum circuit in reverse. By repeatedly executing strategy generation, game interaction, and parameter updates, the optimal game strategy for the current scheduling period is finally output.
7. The distributed optical-storage dynamic game-theoretic optimization design method as described in claim 6, characterized in that: The feedback-driven parameter optimization mechanism includes: In the distributed game process, each node executes an action based on the strategy output by the current parameterized quantum circuit and obtains the actual payoff value. Calculate the deviation between the actual return and the target return to obtain the return deviation. The aforementioned profit bias is used as a gradient signal for reinforcement learning, and the adjustable parameters in the parameterized quantum circuit are adjusted through the backpropagation algorithm. The steps of calculating the profit deviation and adjusting the adjustable parameters constitute a parameter update. The parameter update is iterated alternately with other policy execution steps in each scheduling cycle until the policy converges.
8. A distributed optical-storage dynamic game-theoretic optimization design system, employing the method described in any one of claims 1-7, characterized in that, include: The node modeling module is used to obtain the photovoltaic module parameters, energy storage unit characteristic parameters and local load data of each distributed node, and to construct the node model of the distributed photovoltaic-storage system in combination with grid operation constraints and control objectives. The game modeling module is used to construct a dynamic game model with incomplete information between nodes based on the node model of a distributed optical storage system by defining the local state interaction, strategy space coupling and payoff function dependency between nodes. The quantum strategy generation module is used to introduce parameterized quantum circuits as a strategy generator based on the dynamic game model with incomplete information, and to obtain local game strategies by utilizing the properties of quantum superposition and entanglement. The meta-learning parameter optimization module is used to obtain the optimal initial parameter set by performing inner loop adaptation and outer loop optimization on multiple historical scheduling tasks based on local game strategy and by adopting a model-independent meta-learning framework. The real-time game execution module is used to initialize parameterized quantum circuits and generate initial strategies for each node based on the optimal initial parameter set. Through parallel game and parameter updates driven by payoff deviation, the system converges to equilibrium after multiple rounds of iteration, thus obtaining the optimal game strategy within the current scheduling period.
9. An electronic device, characterized in that, include: Memory, used to store programs; A processor for loading the program to perform the steps of the method as claimed in any one of claims 1-7.
10. A computer-readable storage medium storing a program, characterized in that, When the program is executed by a processor, it implements the steps of the method as described in any one of claims 1-7.