AI large model source network load storage optimization scheduling method and system

Through the coordinated scheduling method of large language model and neural optimal transmission theory, the problem of insufficient characterization of multi-scenario features of the power system is solved, efficient and reliable power system scheduling is achieved, and the economy and safety of the system are significantly improved.

CN120806633AInactive Publication Date: 2025-10-17JINGNENG VISION LINGJINZHIHUI (BEIJING) TECH CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510930708.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-07
Publication Date
2025-10-17
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing technologies are difficult to accurately characterize the multi-scenario characteristics of power systems, scheduling strategies lack adaptability, traditional probability sampling methods are inefficient, risk assessments are inaccurate, and there is a lack of effective feedback mechanisms, resulting in insufficient reliability and economy of scheduling schemes.

Method used

A source-grid-load-storage coordinated scheduling method based on a large language model and neural optimal transmission theory is adopted. Through multi-dimensional scenario recognition and expert model combination, probabilistic causal reasoning of variational entropy coding and differential Monte Carlo continuous sampling are used to construct an online iterative optimization mechanism to achieve efficient parameter optimization and risk control.

Benefits of technology

It improves the adaptability and explainability of scheduling strategies, ensures the safety and economy of scheduling plans, reduces operating costs, and improves the system's response speed and decision-making accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120806633A_ABST
    Figure CN120806633A_ABST
Patent Text Reader

Abstract

The invention provides an AI large model source network load storage optimization scheduling method and system, and relates to the technical field of intelligent scheduling, and the method comprises the steps: receiving power system source end power generation, power grid transmission, user load and energy storage equipment data, and carrying out the time label alignment and preprocessing to obtain a system feature data set; determining a current operation scene by using a multi-dimensional scene identifier of energy flow decomposition, and calling an expert model combination to generate an initial scheduling parameter; inputting the parameters into a pre-trained large language model, and generating an optimized scheduling parameter set through probability causal reasoning of variational entropy coding; constructing a differential Monte Carlo continuous sampling flow, adjusting the sampling probability density by adopting a neural optimal transmission theory, and screening parameter subsets meeting risk constraints; and according to the market price signal and the operation constraint, determining an optimal scheduling parameter to execute coordinated scheduling, and feeding back an execution result to the model to perform online iterative optimization. According to the invention, intelligent coordinated dispatching of the power system is realized, and the safety and economy of system operation are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of intelligent scheduling, in particular to an AI large model source network load storage optimization scheduling method and system. BACKGROUND

[0002] With the continuous expansion of the scale of new energy grid connection, the uncertainty of each link of the source network load storage of the power system significantly increases, and the traditional deterministic scheduling method is difficult to adapt to the complex and variable operation scenarios. In recent years, artificial intelligence technology has been widely applied in the field of power system scheduling, and through deep learning, reinforcement learning and other methods, the intelligent level of system scheduling has been improved. In particular, the emergence of large language models provides a new technical path for knowledge expression, scene understanding and decision optimization of the power system.

[0003] However, the existing technology still has several problems. A single neural network model cannot accurately depict the multi-scene characteristics of the power system, resulting in insufficient adaptability of the scheduling strategy. Traditional probability sampling methods are prone to low sampling efficiency and distribution degradation when dealing with high-dimensional parameter space. The existing methods are not accurate enough in quantifying the system risk, making it difficult to ensure the reliability of the scheduling scheme. There is a lack of effective feedback mechanism for the execution results of scheduling, which cannot realize the continuous optimization of the model.

[0004] In summary, the present application proposes a source network load storage coordinated scheduling method based on large language model and neural optimal transport theory, which improves the adaptability of the scheduling strategy through multi-dimensional scene recognition and expert model combination, improves the explainability of the decision by using variational entropy coding probability causal reasoning, realizes efficient parameter optimization by using differential Monte Carlo continuous sampling and neural optimal transport, and constructs an online iterative optimization mechanism to ensure the continuous improvement of the scheduling performance. Thus, the coordinated scheduling goal of safety, economy and environmental protection of the power system is realized. SUMMARY

[0005] The embodiment of the present application provides an AI large model source network load storage optimization scheduling method and system, which can solve the problems in the prior art.

[0006] The first aspect of the embodiment of the present application is,

[0007] An AI large model source network load storage optimization scheduling method is provided, comprising:

[0008] Receiving source power generation data, power grid transmission data, user load data and energy storage device data in the power system, and obtaining system feature data set through time tag alignment and data preprocessing;

[0009] Based on the system feature dataset, peak period, valley period, new energy consumption period, load response period and extreme weather period are divided, the current running scene is identified and determined through the multi-dimensional scene identifier of energy flow decomposition, the corresponding expert model combination is called, and the initial scheduling parameter is generated;

[0010] The initial scheduling parameter is input into the pre-trained large language model, the probability causal reasoning of the variation entropy coding is constructed to quantize the causal strength distribution, the bias evaluation is performed, and the optimized scheduling parameter set is generated;

[0011] Samples are extracted from the optimized scheduling parameter set, a differential Monte Carlo continuous sampling flow is constructed, the sampling probability density is adjusted based on the neural optimal transport theory, the risk probability distribution is calculated, and a scheduling parameter subset meeting the risk constraint is screened;

[0012] According to the real-time electricity market price signal and the system operation constraint condition, the optimal scheduling parameter is selected from the scheduling parameter subset, and the source, network, load and storage coordinated scheduling is performed;

[0013] The evaluation parameters of the scheduling execution result are fed back to the expert model and the large language model for online iterative optimization.

[0014] In an optional embodiment,

[0015] The multi-dimensional scene identifier of energy flow decomposition identifies and determines the current running scene, calls the corresponding expert model combination, and generates the initial scheduling parameter, which includes:

[0016] The source power, grid power, load power and energy storage power are constructed into corresponding energy flow feature vectors, the energy flow feature vectors are subjected to time dimension decomposition to obtain energy flow components, and the energy flow components are combined to form an energy flow feature matrix;

[0017] An information entropy calculation model based on the energy flow feature matrix is constructed, an optimization objective function of information entropy maximization is established, and the optimization objective function is solved to obtain optimal weight coefficients of each energy flow component. The product of the optimal weight coefficients and the corresponding energy flow component is combined into a multi-dimensional scene feature matrix;

[0018] Based on the multi-dimensional scene feature matrix, a Gaussian mixture probability model is constructed through attention enhanced feature extraction, model parameters are optimized using the expectation maximization algorithm, and the running scene probability distribution of each period is established. The period with the maximum probability value is determined as the current running scene type;

[0019] According to the current running scene type, an expert model combination is called from the expert model library corresponding to each period, historical scheduling data of the expert model combination is extracted, and a historical running error sequence of each expert model is calculated;

[0020] An exponential transformation and normalization processing is performed on the historical operation error sequence to generate an expert model combination weight coefficient, a multi-dimensional scene feature matrix is input into the expert model combination to obtain a scheduling instruction set;

[0021] The scheduling instruction set is weighted and combined according to the expert model combination weight coefficient to generate an initial scheduling parameter.

[0022] In an optional embodiment,

[0023] A Gaussian mixture probability model is constructed based on the multi-dimensional scene feature matrix through attention-enhanced feature extraction, model parameters are optimized using an expectation maximization algorithm, and a probability distribution of the operation scene of each period is established. The period with the maximum probability value is determined as the current operation scene type, including:

[0024] The multi-dimensional scene feature matrix is input into a multi-head attention network, the multi-head attention network calculates multiple groups of attention features through a learnable projection matrix, and scene correlation features are obtained by splicing the multiple groups of attention features;

[0025] The scene correlation features are input into a conditional variational autoencoder, which encodes the scene correlation features into a posterior distribution of a first scene variable under the condition of a period;

[0026] The first scene variable is input into a recurrent neural network to obtain a second scene variable, and the Gaussian mixture probability model of each period is constructed based on the second scene variable, including a period distribution weight, a period distribution mean, and a period distribution covariance;

[0027] The posterior probability of each period distribution in the Gaussian mixture probability model is calculated through an expectation maximization algorithm, the posterior probability is multiplied by the corresponding relationship of the period and summed to obtain a probability distribution of the operation scene of each period. The period with the maximum probability value in the operation scene probability distribution is determined as the current operation scene type.

[0028] In an optional embodiment,

[0029] The initial scheduling parameter is input into a pre-trained large language model, a causal strength quantization distribution is constructed through a variational entropy encoding probability causal inference, a bias evaluation is performed, and an optimized scheduling parameter set is generated, including:

[0030] The initial scheduling parameter is input into a pre-trained large language model, a probability representation of the initial scheduling parameter is constructed using a variational entropy encoding, and a first parameter probability representation in a Gaussian distribution form is generated through an encoding mapping function;

[0031] constructing a hierarchical probabilistic causal inference model based on the first parameter probability representation, calculating conditional probability log ratio values and parameter coupling strengths between the initial scheduling parameters, and obtaining a causal strength quantization distribution by optimizing a time-varying causal strength function

[0032] calculating a two-norm deviation of the initial scheduling parameters from expected target values, constructing a variational evidence lower bound function including a reconstruction error term and a KL divergence regularization term based on the causal strength quantization distribution, and iteratively optimizing variational entropy encoding parameters by maximizing the variational evidence lower bound function

[0033] performing Monte Carlo sampling on the second parameter probability representation to obtain candidate parameters, verifying whether the candidate parameters satisfy power balance constraints of power generation loads, charge and discharge capacity constraints of energy storage devices, unit ramp rate constraints, and voltage limit constraints of grid nodes, and grouping candidate parameters satisfying all constraint conditions into an optimized scheduling parameter set

[0034] In an optional embodiment,

[0035] constructing a hierarchical probabilistic causal inference model based on the first parameter probability representation, calculating conditional probability log ratio values and parameter coupling strengths between the initial scheduling parameters, and obtaining a causal strength quantization distribution by optimizing a time-varying causal strength function

[0036] constructing the first parameter probability representation of the initial scheduling parameters as a probabilistic causal inference model, the probabilistic causal inference model using a Student-t distribution to construct a probability distribution of each initial scheduling parameter, the Student-t distribution including a degree of freedom, a location parameter, and a scale parameter, to obtain a parameter probability representation

[0037] inputting the parameter probability representation into a multi-level causal graph, the multi-level causal graph including an overall constraint layer of a power system, a subsystem coupling relationship layer, and a scheduling parameter interaction layer, calculating time sequence dependency relationships between the initial scheduling parameters based on a time sequence characteristic function and a corresponding weight, and calculating parameter coupling strengths by a coupling measure function

[0038] constructing a joint distribution between any two initial scheduling parameters based on a Copula function, calculating a causal direction strength by calculating a conditional probability log ratio value of the joint distribution and an independent distribution, combining the causal direction strength and the parameter coupling strength, and constructing a multi-dimensional causal strength matrix

[0039] calculating a conditional probability log ratio value between any two initial scheduling parameters, constructing an information gain measure based on the conditional probability log ratio value, and determining a causal strength confidence as a ratio of the information gain measure and the conditional probability log ratio value of the initial scheduling parameters

[0040] A time-varying causal strength function is constructed including a historical causal strength term, an incremental term and a real-time feedback term, an optimization objective function is constructed based on the causal strength confidence and the time-varying causal strength function, and a causal strength quantitative distribution is obtained through an expectation value of the conditional probability logarithmic ratio.

[0041] In an alternative embodiment,

[0042] Samples are extracted from the optimization scheduling parameter set to construct a differential Monte Carlo continuous sampling flow, a sampling probability density is adjusted based on a neural optimal transport theory, a risk probability distribution is calculated, and a scheduling parameter subset satisfying a risk constraint is screened, including:

[0043] Samples are extracted from the optimization scheduling parameter set to construct an initial sampling set, the initial sampling set is input into a differential Monte Carlo continuous sampling flow, a stochastic differential equation containing a drift term and a Brownian motion diffusion term is constructed in the differential Monte Carlo continuous sampling flow, and continuous sampling trajectories are obtained by solving the stochastic differential equation;

[0044] A second-order Wasserstein distance between an initial probability density and a target probability density of the continuous sampling trajectories is calculated, a neural network based on a self-attention mechanism is constructed to fit an optimal transport mapping, an adjustment strategy for the sampling probability density is obtained by determining an overall loss function and combining a dynamic step strategy;

[0045] The sampling probability density of the continuous sampling trajectories is adjusted according to the adjustment strategy, a risk probability distribution of the adjusted sampling probability density is calculated, and the risk probability distribution is composed of a weighted combination of a deviation probability of the generated power and the load power, a node voltage out-of-limit probability and an energy storage capacity out-of-limit probability;

[0046] The risk probability distribution is compared with a preset system risk threshold, scheduling parameters satisfying a risk constraint are screened, and an optimization scheduling parameter subset is formed.

[0047] In an alternative embodiment,

[0048] A second-order Wasserstein distance between an initial probability density and a target probability density of the continuous sampling trajectories is calculated, a neural network based on a self-attention mechanism is constructed to fit an optimal transport mapping, an adjustment strategy for the sampling probability density is obtained by determining an overall loss function and combining a dynamic step strategy, including:

[0049] The initial probability density function is obtained by kernel density estimation of the continuous sampling trajectories, and a second-order Wasserstein distance between the initial probability density function and a target probability density function is calculated;

[0050] A deep neural network is constructed according to a second-order Wasserstein distance, as an optimization transport mapping function, input data is converted into a query matrix, a key matrix and a value matrix, a dot product of the query matrix and the key matrix is calculated to obtain an attention weight, and the attention weight is multiplied with the value matrix to obtain a feature representation;

[0051] A transport error loss between the optimization transport mapping function and a theoretical optimal transport mapping is calculated, a gradient regularization loss of the optimization transport mapping function is calculated, a cyclic consistency loss of the optimization transport mapping function is calculated, and the transport error loss, the gradient regularization loss and the cyclic consistency loss are combined to obtain a total loss function;

[0052] An optimization gradient is calculated according to the total loss function, a dynamic step size is calculated using an exponential function of an initial step size and a decay factor, a minimum step size threshold is set, and the dynamic step size is used to iteratively update the sampling points to obtain an adjustment strategy of the sampling probability density.

[0053] The second aspect of the embodiment of the application,

[0054] An AI large model source network load storage optimization scheduling system is provided, comprising:

[0055] A first unit is configured to receive source power generation data, power grid transmission data, user load data and energy storage device data in a power system, and obtain a system feature data set through time tag alignment and data preprocessing;

[0056] A second unit is configured to divide peak period, valley period, new energy consumption period, load response period and extreme weather period based on the system feature data set, identify a current operation scenario through a multi-dimensional scene identifier of energy flow decomposition, call a corresponding expert model combination, and generate initial scheduling parameters;

[0057] A third unit is configured to input the initial scheduling parameters into a pre-trained large language model, construct a causal strength quantization distribution through a probability causal inference of variational entropy coding, perform bias evaluation, and generate an optimized scheduling parameter set;

[0058] A fourth unit is configured to extract samples from the optimized scheduling parameter set, construct a differential Monte Carlo continuous sampling flow, adjust a sampling probability density based on a neural optimal transport theory, calculate a risk probability distribution, and screen a scheduling parameter subset meeting a risk constraint;

[0059] A fifth unit is configured to screen optimal scheduling parameters from the scheduling parameter subset according to real-time electricity market price signals and system operation constraint conditions, and perform source network load storage coordinated scheduling;

[0060] The sixth unit is configured to feed back the evaluation parameters of the dispatch execution result to the expert model and the large language model for online iterative optimization.

[0061] The third aspect of the embodiment of the application,

[0062] An electronic device is provided, comprising:

[0063] A processor;

[0064] A memory for storing processor-executable instructions;

[0065] The processor is configured to invoke the instructions stored in the memory to execute the method described above.

[0066] The fourth aspect of the embodiment of the application,

[0067] A computer-readable storage medium is provided, which stores computer program instructions, and the computer program instructions are executed by a processor to implement the method described above.

[0068] In the embodiment of the application, the multi-dimensional scene identifier based on energy flow decomposition can automatically distinguish peak / valley, new energy consumption, load response and extreme weather operation scenes, and call corresponding expert model combinations to generate initial scheduling parameters with strong pertinence, thereby significantly improving the scene matching degree and response speed of the scheduling strategy; the pre-trained large language model is combined with the variational entropy coding to construct a probability causal reasoning framework to evaluate and optimize the deviation of the initial scheduling parameters, which not only ensures the global optimal trend of the scheduling result, but also enhances the explainability and anti-disturbance ability of the model output; based on the differential Monte Carlo continuous sampling flow and the neural optimal transport theory, the sampling probability density is dynamically adjusted and the risk distribution is calculated, so that the scheduling parameter subset with the minimum risk can be selected under the premise of meeting the system safety and reliability constraints, and the prevention and control of rare extreme events are realized; the real-time power market price signal and the system operation constraint are combined to perform real-time parameter screening, so that the final scheduling scheme not only meets the technical feasibility, but also can obtain the maximum economic benefit in the power market, thereby significantly reducing the operating cost. BRIEF DESCRIPTION OF DRAWINGS

[0069] Figure 1 FIG. 1 is a flowchart of an AI large model source network load storage optimization scheduling method according to an embodiment of the application;

[0070] Figure 2 FIG. 6 is an actual benefit comparison radar chart of energy management according to an embodiment of the application;

[0071] Figure 3 FIG. 7 is a multi-level causal graph structure diagram according to an embodiment of the application;

[0072] Figure 4 FIG. 8 is a sampling strategy optimization effect thermodynamic diagram according to an embodiment of the application. DETAILED DESCRIPTION

[0073] To make the objects, technical solutions and advantages of the embodiments of the present application clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative work belong to the scope of protection of the present application.

[0074] The technical solutions of the present application will be described in detail below with specific embodiments. The following specific embodiments can be combined with each other, and some embodiments may not be described again for the same or similar concepts or processes.

[0075] Figure 1 The flowchart of the source network load storage optimization scheduling method of the AI large model embodiment of the present application is shown in FIG. 1, which comprises the following steps. Figure 1

[0076] Receiving source power generation data, power grid transmission data, user load data and energy storage device data in the power system, and obtaining system feature data set through time tag alignment and data preprocessing;

[0077] Based on the system feature data set, dividing peak period, valley period, new energy consumption period, load response period and extreme weather period, identifying and determining the current running scene through the multi-dimensional scene identifier of energy flow decomposition, calling the corresponding expert model combination, and generating initial scheduling parameters;

[0078] Inputting the initial scheduling parameters into the pre-trained large language model, constructing the causal strength quantitative distribution through the probability causal reasoning of the variational entropy coding, performing the bias evaluation, and generating the optimization scheduling parameter set;

[0079] Extracting samples from the optimization scheduling parameter set, constructing the differential Monte Carlo continuous sampling flow, adjusting the sampling probability density based on the neural optimal transport theory, calculating the risk probability distribution, and screening the scheduling parameter subset meeting the risk constraint;

[0080] According to the real-time electricity market price signal and the system operation constraint condition, screening the optimal scheduling parameter from the scheduling parameter subset, and performing the source network load storage coordination scheduling;

[0081] Feeding the evaluation parameters of the scheduling execution result to the expert model and the large language model, and performing the online iterative optimization.

[0082] In an optional embodiment,

[0083] ​The current operation scene is identified by a multi-dimensional scene identifier based on energy flow decomposition, a corresponding expert model combination is called, and initial scheduling parameters are generated, including:

[0084] The source power, grid power, load power, and energy storage power are constructed into corresponding energy flow feature vectors, and time dimension decomposition is performed on the energy flow feature vectors to obtain energy flow components, which are combined to form an energy flow feature matrix;

[0085] An information entropy calculation model based on the energy flow feature matrix is constructed, an optimization objective function maximizing information entropy is established, and the optimization objective function is solved to obtain optimal weight coefficients of each energy flow component, which are combined with the corresponding energy flow components to form a multi-dimensional scene feature matrix;

[0086] A Gaussian mixture probability model is constructed based on the multi-dimensional scene feature matrix through attention-enhanced feature extraction, model parameters are optimized using the expectation maximization algorithm, and the operation scene probability distribution of each period is established, and the period with the maximum probability value is determined as the current operation scene type;

[0087] According to the current operation scene type, an expert model combination is called from the expert model library corresponding to each period, historical scheduling data of the expert model combination is extracted, and a historical operation error sequence of each expert model is calculated;

[0088] The historical operation error sequence is subjected to exponential transformation and normalization processing to generate expert model combination weight coefficients, and the multi-dimensional scene feature matrix is input into the expert model combination to obtain a scheduling instruction set;

[0089] The scheduling instruction set is weighted and combined according to the expert model combination weight coefficients to generate initial scheduling parameters.

[0090] In a specific embodiment, real-time data of source power (such as photovoltaic power generation, wind power generation), grid power (such as input and output power), load power (such as residential electricity load, commercial load), and energy storage power (such as battery charging and discharging power) are collected. For example, a certain microgrid collects data every 5 minutes in a 24-hour period to form a source power vector PS = [5.2kW, 5.5kW,..., 0.5kW], a grid power vector PG = [2.1kW, 2.3kW,..., 4.5kW], a load power vector PL = [7.8kW, 8.1kW,..., 6.2kW], and an energy storage power vector PB = [0.5kW, -0.3kW,..., 1.2kW]. These vectors are combined to form an initial energy flow feature vector V0.

[0091] The wavelet transform is used to decompose V0 in time dimension to obtain energy flow components of different time scales. The db4 wavelet basis function is used for 3-layer decomposition to obtain high-frequency components VH1, VH2, VH3 and low-frequency component VL3. The high-frequency components reflect the short-term fluctuation characteristics of the energy flow, and the low-frequency component reflects the long-term trend. These components are combined to form an energy flow feature matrix M = [VH1, VH2, VH3, VL3].

[0092] For each component in the matrix M, its probability distribution is calculated. Taking VH1 as an example, its numerical range is divided into 10 intervals, and the frequency of data points in each interval is counted to obtain the probability distribution PH1 = [0.05, 0.12, 0.15, 0.20, 0.18, 0.12, 0.08, 0.05, 0.03, 0.02]. Similarly, the probability distributions of other components are calculated.

[0093] Based on the probability distribution of each component, the corresponding information entropy is calculated. The greater the information entropy, the more information the component contains. For the information entropy value of each component, a weight optimization model is established, with the maximum information entropy as the objective function and the constraint condition that the weight sum is 1. The Lagrange multiplier method is used to solve it, and the optimal weight coefficient W = [0.35, 0.25, 0.15, 0.25] is obtained. Multiply the weight coefficient with the corresponding energy flow component to obtain the multi-dimensional scene feature matrix wherein represents the component multiplication operation.

[0094] Based on the multi-dimensional scene feature matrix, a Gaussian mixture probability model is constructed through attention-enhanced feature extraction. Specifically, an attention mechanism is designed to calculate the attention score of each element in the matrix F. For example, the similarity between the energy flow state vector at each time point and the historical typical pattern is calculated to obtain the attention weight A = [0.12, 0.08,..., 0.15]. The enhanced feature matrix F' is obtained by combining the attention weight with the feature matrix F.

[0095] Based on F', a Gaussian mixture probability model (GMM) is constructed, with the number of mixed components set to 5, corresponding to typical scenes such as peak load, valley load, source high output, grid high input, and energy storage high charging. For each scene, a multivariate Gaussian distribution model is established, including the mean vector and the covariance matrix. For example, the Gaussian distribution parameters of the peak load scene are: mean vector μ1 = [8.5, 3.2, -0.5, 1.2], and covariance matrix Σ1 is obtained by statistical analysis of historical data.

[0096] The expectation-maximization algorithm is used to optimize the model parameters. The E-step and M-step are iteratively performed to update the model parameters until convergence. The E-step calculates the posterior probability of each data point belonging to each Gaussian distribution, and the M-step updates the Gaussian distribution parameters and mixing weights based on the posterior probability. After optimization, the probability distribution matrix P of the running scene for each period is obtained, with dimensions [period number, scene number]. The scene type with the highest probability in each period is selected as the running scene of that period. For example, if the probability distribution of a certain period is [0.15, 0.05, 0.65, 0.10, 0.05], the running scene type of that period is "source high power".

[0097] According to the current running scene type, the corresponding expert model combination is called from the expert model library, which contains multiple models optimized for different scenarios, such as energy storage priority scheduling model for peak load scenario, load forecasting model for valley load scenario, power balance model for source high power scenario, etc. Assuming the current running scene is "source high power", the corresponding three expert models are called: power balance model M1, energy storage scheduling model M2, and economic optimization model M3.

[0098] Extract the historical scheduling data of the expert model, calculate the historical running error sequence, analyze the deviation between the predicted power and actual power of M1 in the past 30 days, and obtain the error sequence E1 = [0.08, 0.12,..., 0.05]. Similarly, calculate the error sequences E2 and E3 of M2 and M3.

[0099] Perform exponential transformation and normalization on the error sequence to generate expert model combination weight coefficients, the exponential transformation formula is w'i = exp(-λ×Ei), where λ is the decay coefficient, set to 5. For example, after exponential transformation of E1, E2, and E3, w'1 = 0.67, w'2 = 0.45, and w'3 = 0.52 are obtained. After normalization, the maximum weight coefficients w1 = 0.41, w2 = 0.27, and w3 = 0.32 are obtained.

[0100] Input the multi-dimensional scene feature matrix F' into the expert model combination to obtain the scheduling instruction set. For example, M1 outputs the source power limit instruction D1 = [4.5kW, 4.8kW,..., 5.0kW], M2 outputs the energy storage charging instruction D2 = [1.2kW, 1.5kW,..., 1.8kW], and M3 outputs the economic scheduling parameter D3 = [0.42 yuan / kWh, 0.45 yuan / kWh,..., 0.48 yuan / kWh].

[0101] According to the expert model combination weight coefficient, the scheduling instruction set is weighted and combined to generate initial scheduling parameters. The source end power limit value is calculated as D1'=w1xD1, the energy storage charging power is calculated as D2'=w2xD2, and the unit electricity price is calculated as D3'=w3xD3. The combined initial scheduling parameters include the source end power limit value, the energy storage charging power, the unit electricity price, etc., and can be directly used for the execution of the microgrid control system.

[0102] In the embodiment, by constructing the source end, the network end, the load and the energy storage power into an energy flow feature vector and decomposing in the time dimension to form a matrix representation, the fine description of various energy flow behaviors is realized, which lays a high-quality data foundation for subsequent scene identification; by using an optimization objective function of information entropy maximization, the optimal weight coefficients of each energy flow component are automatically solved, so that under different operating conditions, the system can dynamically adjust the focus, and the discrimination and robustness of the feature representation are improved; based on the attention-enhanced feature extraction, a Gaussian mixture model is constructed, and the parameters are optimized by the expectation maximization algorithm, and the scene probability distribution of each period is quantitatively obtained, so that the current operating scene can be objectively and quantitatively identified, and the subjective deviation of artificial experience judgment is avoided; the exponential transformation and normalization processing of the historical operation error are introduced to generate the expert model weight for the current scene, so that each scheduling instruction obtains the optimal trust degree distribution from the multi-model fusion, and the accuracy of the initial scheduling instruction is significantly improved.

[0103] In an alternative embodiment, a Gaussian mixture probability model is constructed based on the multi-dimensional scene feature matrix through attention-enhanced feature extraction, the model parameters are optimized by the expectation maximization algorithm, the operating scene probability distribution of each period is established, and the period with the maximum probability value is determined as the current operating scene type, including:

[0104] The multi-dimensional scene feature matrix is input into a multi-head attention network, the multi-head attention network calculates multiple groups of attention features through a learnable projection matrix, and the scene correlation features are obtained by splicing the multiple groups of attention features;

[0105] The scene correlation features are input into a conditional variational autoencoder, the conditional variational autoencoder encodes the scene correlation features into the posterior distribution of the first scene variable under the period condition constraint;

[0106] The first scene variable is input into a recurrent neural network to obtain a second scene variable, and the Gaussian mixture probability model of each period is constructed based on the second scene variable, including the period distribution weight, the period distribution mean and the period distribution covariance;

[0107] The posterior probability of each time period distribution in the Gaussian mixture probability model is calculated by an expectation maximization algorithm, the posterior probability is multiplied by the corresponding relationship of the time period and summed to obtain the running scene probability distribution of each time period, and the time period with the maximum probability value in the running scene probability distribution is determined as the current running scene type.

[0108] In a specific embodiment, a multi-dimensional scene feature matrix is input into a multi-head attention network for feature enhancement. The dimension of the multi-dimensional scene feature matrix F is [time period number T x feature number D], for example, in actual application, T = 288 (representing one sample point every 5 minutes in a day), and D = 16 (containing various energy flow features). The multi-head attention network includes H = 8 attention heads, and each attention head projects the input feature to a query vector (Query), a key vector (Key) and a value vector (Value) through three learnable projection matrices WQ, WK and WV. Taking the hth attention head as an example, the dimensions of WQ, WK and WV are all [D x (D / H)], that is, the 16-dimensional feature is projected to a 2-dimensional subspace. After projection, the query matrix Qh = [q1,h, q2,h,..., qT,h], the key matrix Kh = [k1,h, k2,h,..., kT,h] and the value matrix Vh = [v1,h, v2,h,..., vT,h] are obtained, wherein the dimension of each vector is D / H = 2.

[0109] For each attention head, the similarity between the query matrix and the key matrix is calculated to obtain an attention score matrix. Taking the first attention head as an example, the dot product of q1,1 and all key vectors k1,1, k2,1,..., kT,1 is calculated, and then divided by the scaling factor (D / H) 1 / 2 = 2 1 / 2 , and then normalized by applying a softmax function to obtain the attention weight vector a1,1 = [0.02, 0.05, 0.01,..., 0.03]. Similarly, the attention weight vectors of other time periods are calculated to form an attention weight matrix A1. The attention weight matrix is multiplied by the value matrix to obtain the output feature O1 of the attention head. The above calculation is repeated for all 8 attention heads to obtain 8 groups of output features O1 to O8.

[0110] The 8 groups of output features are spliced in the feature dimension to obtain the enhanced scene correlation feature C = [O1, O2,..., O8], with a dimension of [T x D]. The scene correlation feature captures the dependency relationship between different time periods, for example, the correlation of the energy flow state of a certain time period with the previous and subsequent time periods.

[0111] The scene-related features are input into a conditional variational autoencoder (CVAE) for feature encoding. The CVAE consists of an encoder and a decoder. The encoder encodes the scene-related features into the posterior distribution of latent variables by time period conditioning. Specifically, the encoder adopts a two-layer fully connected neural network, each layer containing 128 neurons, with ReLU as the activation function. Taking the t-th time period as an example, the feature vector ct of the time period is connected with the time period encoding et and then input into the encoder, where et is the one-hot encoding vector of the time period t, with a length of T = 288. The encoder outputs two vectors μt and σt, representing the mean and standard deviation of the latent variable zt, both with a dimension of 32.

[0112] According to the reparameterization trick, the first scene variable zt is sampled from the normal distribution N(μt, σt). Taking a certain instance as an example, μt = [0.3, -0.5, 0.2,..., 0.1] and σt = [0.1, 0.2, 0.1,..., 0.15], the sampling zt = [0.32, -0.45, 0.22,..., 0.08] is obtained. The same operation is performed for all time periods to obtain the first scene variable sequence Z = [z1, z2,..., zT] with a dimension of [T × 32].

[0113] The first scene variable sequence Z is input into a bidirectional long short-term memory network (BiLSTM) for time series modeling. The BiLSTM contains two LSTM layers, each with 64 hidden units. The forward LSTM processes from z1 to zT, and the backward LSTM processes from zT to z1. The second scene variable sequence Z' = [z'1, z'2,..., z'T] is obtained by concatenating the forward and backward outputs, with a dimension of [T × 128]. The second scene variable captures the long-term dependencies and time series patterns between different time periods.

[0114] Based on the second scene variable, a Gaussian mixture probability model (GMM) is constructed for each time period, containing K = 5 mixture components, corresponding to 5 typical operating scenarios: peak load, valley load, high renewable energy output, high energy storage charging, and grid peak shaving. The second scene variable z't is mapped to the parameters of the GMM through a fully connected network, including the time period distribution weights πt = [πt,1, πt,2,..., πt,K], the time period distribution means Mt = [μt,1, μt,2,..., μt,K], and the time period distribution covariances Σt = [σt,1, σt,2,..., σt,K].

[0115] The GMM parameters are optimized by an expectation-maximization algorithm, including first step and second step alternating iterations, the first step calculates the posterior probability of each data point belonging to each mixed component, and the second step updates the model parameters. After initializing the model parameters, the E step is performed to calculate the posterior probability γt,k of each time period t corresponding to the scene k. For example, the posterior probability of the t=145 time period is calculated as γ145,1=0.12, γ145,2=0.18, γ145,3=0.55, γ145,4=0.08, and γ145,5=0.07. The M step updates the model parameters based on the posterior probability. After performing the first step and the second step for 10 iterations, the model parameters converge.

[0116] The running scene probability distribution of each time period is calculated. The posterior probability is multiplied by the corresponding relationship of the time period and summed to obtain the probability belonging to each scene type P=[p1, p2,..., pK]. Taking the t=145 time period as an example, the probability belonging to the k=3 type scene (high renewable energy output) is the highest, p3=0.55. Therefore, the t=145 time period is determined as the running scene type of "high renewable energy output".

[0117] Compared with the prior art, the scheme of the embodiment makes a significant improvement in scene recognition. The existing scene recognition technology mainly relies on simple statistical features or shallow machine learning models, and it is difficult to capture the complex spatio-temporal correlation in energy flow data. For example, the traditional method usually adopts K-means clustering or support vector machine classification, which cannot effectively handle the dependency relationship between different time periods, resulting in unclear scene boundary recognition and low accuracy.

[0118] The scheme of the embodiment significantly enhances the correlation modeling capability of energy flow features in different time periods by introducing a multi-head attention mechanism, solving the problem that the traditional method cannot capture long-distance dependency. At the same time, the application of conditional variational autoencoder realizes feature representation learning under different time period conditions, providing more rich probability semantic information. In addition, combined with the time series modeling of bidirectional LSTM and the Gaussian mixture probability model, the recognition accuracy of complex scenes is further improved.

[0119] Actual tests show that compared with the traditional method, the scene recognition accuracy of the scheme of the embodiment is improved, especially in the transition period of rapid energy flow change, the recognition accuracy is more significantly improved, and accurate scene recognition lays a solid foundation for subsequent expert model scheduling, and finally realizes more efficient and more economical energy management.

[0120] As Figure 2As shown, the comparison between the present embodiment and the prior art in terms of actual energy management benefits is shown. The data shows that the present embodiment achieves the best results in six key indicators: in terms of energy cost savings, the present embodiment reaches 18.7%, which is significantly higher than other methods (K-means clustering 8.2%, support vector machine 10.1%, random forest 12.5%, LSTM 14.5%); the peak load reduction rate reaches 23.4%, which is 5.8 percentage points higher than the closest LSTM method (17.6%); the carbon emission reduction rate is 15.3%, which is higher than the LSTM of 12.2%; the equipment utilization rate is increased by 12.8%, which is higher than the LSTM of 10.3%; the operation and maintenance cost is reduced by 20.5%, which is 5.2 percentage points higher than the LSTM of 15.3%; the system response time is only 1.2 seconds, which is much lower than other methods (K-means clustering 4.7 seconds, support vector machine 4.2 seconds, random forest 5.8 seconds, LSTM 6.5 seconds). The annual average energy cost savings reach 18.7%, which is 4.2% higher than the best existing technology; the peak load reduction rate reaches 23.4%, which is 5.8%; the carbon emission reduction rate reaches 15.3%, which is 3.1%; the equipment utilization rate is increased by 12.8%, which is increased by 2.5%. These data are based on the actual operation data of a certain industrial park (10 buildings, total area 85,000 square meters, annual electricity consumption about 12 million kilowatt-hours) throughout the year, fully verifying the significant economic and environmental benefits of the present embodiment in the actual energy management scenario.

[0121] In a specific embodiment, the initial scheduling parameters are input into a pre-trained large language model, a causal strength quantization distribution is constructed through a probabilistic causal reasoning of a variational entropy encoding, a bias evaluation is performed, and an optimized scheduling parameter set is generated, including:

[0122] The initial scheduling parameters are input into a pre-trained large language model, a probabilistic representation of the initial scheduling parameters is constructed using a variational entropy encoding, and a first parameter probability representation in the form of a Gaussian distribution is generated through an encoding mapping function;

[0123] A hierarchical probabilistic causal reasoning model is constructed based on the first parameter probability representation, the conditional probability logarithmic ratio and the parameter coupling strength between the initial scheduling parameters are calculated, and the causal strength quantization distribution is optimized through a time-varying causal strength function

[0124] The initial scheduling parameters and the expected target value are calculated for the two-norm bias, a variational lower bound function containing a reconstruction error term and a KL divergence regularization term is constructed based on the causal strength quantization distribution, and the variational entropy encoding parameters are iteratively optimized by maximizing the variational lower bound;

[0125] The candidate parameters are obtained by performing Monte Carlo sampling on the second parameter probability representation, and it is verified whether the candidate parameters meet the power balance constraints of the power generation load, the charging and discharging capacity constraints of the energy storage device, the climbing rate constraints of the unit, and the voltage limit constraints of the grid node. The candidate parameters that meet all the constraint conditions form an optimized scheduling parameter set.

[0126] In a specific embodiment, the initial scheduling parameters are input into a pre-trained large language model for encoding, including source end power limit values PS = [4.5kW, 4.8kW, …, 5.0kW], energy storage charging power PB = [1.2kW, 1.5kW, …, 1.8kW], load adjustment parameters PL = [7.8kW, 8.1kW, …, 6.2kW], and grid exchange power PG = [2.1kW, 2.3kW, …, 4.5kW], etc. These parameter sequences are serialized into structured text, such as "source end power: 4.5, 4.8, …, 5.0; energy storage power: 1.2, 1.5, …, 1.8; …", as input to the large language model. The pre-trained large language model obtains the hidden state vector H of the second-to-last layer of the model after inputting the text into the model, with a dimension of

[1536] .

[0127] A variational entropy encoder is used to construct a probability representation of the initial scheduling parameters. The variational entropy encoder is composed of three fully connected neural networks with hidden layer sizes of 512 and 256, respectively, and a GELU activation function. The input is the hidden state vector H generated by the large language model, and the output is two vectors μ and σ, representing the mean and standard deviation of the Gaussian distribution, both with a dimension of 128. Taking an actual case as an example, the encoder outputs μ = [0.25, -0.42, 0.18, …, 0.33] and σ = [0.05, 0.08, 0.06, …, 0.07].

[0128] By using the reparameterization technique, a first parameter probability representation z is sampled from the Gaussian distribution N(μ, σ). Specifically, a random vector ε is generated, where each element is sampled from a standard normal distribution, and then z = μ + σ * ε is calculated. For example, z = [0.26, -0.40, 0.17, …, 0.31] is sampled. The first parameter probability representation contains the distribution characteristics and uncertainty information of the initial scheduling parameters.

[0129] Based on the first parameter probability representation, a hierarchical probabilistic causal reasoning model is constructed. The hierarchical probabilistic causal reasoning model is composed of two layers of graph structures: a parameter-level causal graph and a system-level causal graph. The parameter-level causal graph describes the causal relationships between different scheduling parameters, and the system-level causal graph describes the causal relationships between the parameters and the system state.

[0130] In the parameter-level causal graph, the conditional probability log-ratio between any two scheduling parameters xi and xj is calculated. Specifically, the conditional probability information is extracted from the first parameter probability representation by a neural network f(z). The neural network f contains two fully connected layers with a hidden layer size of 256 and an output layer size of the square of the number of scheduling parameters. For the relationship between the source power and the energy storage power, the calculated conditional probability log-ratio is 0.75, indicating that the source power change has a strong influence on the energy storage power.

[0131] Based on the conditional probability log-ratio, the coupling strength between parameters is calculated, and the coupling strength matrix C has a dimension of [parameter number x parameter number], with each element Cij representing the influence strength of parameter i on parameter j. For example, the coupling strength of the source power on the energy storage power is 0.68, and the coupling strength on the load regulation parameter is 0.32.

[0132] The coupling strength matrix is optimized by a time-varying causal strength function. The time-varying causal strength function g(C, t) introduces a time dimension to capture the changes in parameter coupling relationships at different times. The function g is implemented by a recurrent neural network, which contains an LSTM layer with 128 hidden units. For the peak load period, the coupling strength of the source power on the energy storage power increases to 0.82; for the valley load period, this coupling strength decreases to 0.55. After optimization, the causal strength quantization distribution Q is obtained, which has the same dimension as the coupling strength matrix C.

[0133] The initial scheduling parameters are calculated with the expected target values to obtain the 2-norm deviation, and the expected target values come from historical optimal operation data or expert settings, such as the expected source power PS* = [4.6 kW, 4.9 kW,..., 5.1 kW]. The 2-norm deviation D = ||PS-PS*||2 = 0.42 is calculated, and the deviations of other parameters are calculated in the same way.

[0134] Combining the causal strength quantization distribution and the parameter deviation, a variational evidence lower bound function is constructed, which includes a reconstruction error term and a KL divergence regularization term. The reconstruction error term measures the difference between the initial scheduling parameters and the expected target values, and the KL divergence regularization term measures the difference between the distribution of the first parameter probability representation and the standard normal distribution. The variational evidence lower bound function is represented as the weighted sum of the reconstruction error term and the KL divergence regularization term, with a weight coefficient β = 0.1.

[0135] The variational evidence lower bound function is maximized by the gradient ascent method to iteratively optimize the parameters of the variational entropy encoder. The Adam optimizer is used with a learning rate of 0.001 and 500 iterations. During the optimization process, the variational evidence lower bound increases from the initial value of -245.8 to -98.3, indicating that the encoding quality has been significantly improved. The optimized encoder outputs the updates μ' = [0.28, -0.38, 0.22,..., 0.35] and σ' = [0.04, 0.06, 0.05,..., 0.06].

[0136] Based on the optimized μ' and σ', a second parameter probability representation z' is generated. Compared with the first parameter probability representation, the second parameter probability representation is closer to the distribution characteristics of the optimal scheduling parameters. Monte Carlo sampling is performed on the second parameter probability representation to generate 1000 candidate parameter sets. Taking the 500th sampling as an example, candidate source power PS' = [4.55 kW, 4.85 kW, …, 5.05 kW], candidate energy storage power PB' = [1.15 kW, 1.48 kW, …, 1.75 kW], and the like are obtained.

[0137] Verify whether the candidate parameters meet the multiple system constraint conditions. First, check the power balance constraint of the power generation load: in each time period t, the sum of the source power PS'(t) and the grid power PG'(t) should be equal to the sum of the load power PL'(t) and the energy storage power PB'(t), allowing an error of ±0.1 kW. For example, at t = 5, PS'(5) + PG'(5) = 4.85 + 2.25 = 7.1 kW, PL'(5) + PB'(5) = 8.05 + 1.48 = 7.03 kW, the error is 0.07 kW, which meets the balance constraint.

[0138] Secondly, check the charge and discharge capacity constraint of the energy storage device: the energy storage power PB'(t) should be within the rated power range of the device, i.e. -2.5 kW ≤ PB'(t) ≤ 2.5 kW. For example, all values in the PB' list are within the range of [-1.8 kW, 1.85 kW], which meets the capacity constraint.

[0139] Thirdly, check the unit ramp rate constraint: the source power change rate between adjacent time periods should not exceed a set threshold, i.e. |PS'(t+1)-PS'(t)| ≤ 0.5 kW. For example, |PS'(6)-PS'(5)| = |4.9-4.85| = 0.05 kW < 0.5 kW, which meets the ramp constraint.

[0140] Finally, check the in-out power constraint of the grid node. After verifying all the constraint conditions, select the parameter set with the maximum variational evidence lower bound value from the candidate parameters that meet the constraints as the final optimization result. Finally, the optimized scheduling parameters PS_opt, PB_opt, PL_opt and PG_opt are obtained, which are used to guide the actual operation of the energy system.

[0141] In the embodiment, the initial scheduling parameters are mapped into a probability representation in the form of Gaussian distribution by variational entropy coding, the uncertainty of the parameters is explicitly characterized, and the robustness of the scheduling scheme to input fluctuations is enhanced; the conditional probability log ratio value and the coupling strength between the parameters are calculated based on a hierarchical probabilistic causal reasoning model, and the causal dependence relationship between the scheduling parameters is quantitatively revealed by optimization of a time-varying causal strength function, and the explainability of the decision-making process is improved; a variational evidence lower bound is constructed by combining the two-norm reconstruction error and the KL divergence regularization term, and the encoding parameters are iteratively optimized by maximizing the lower bound, so that the systematic and globally optimal optimization of the scheduling parameters is realized; the candidate parameters are generated by Monte Carlo sampling, and multiple engineering constraints such as power balance, energy storage capacity, ramp rate, and voltage limit are strictly verified, so that the final optimized scheduling parameters not only meet the technical specifications but also have high feasibility.

[0142] In an alternative embodiment, a hierarchical probabilistic causal reasoning model is constructed based on the first parameter probability representation, the conditional probability log ratio value and the parameter coupling strength between the initial scheduling parameters are calculated, and the causal strength quantization distribution is obtained by optimization of a time-varying causal strength function, including:

[0143] The first parameter probability representation of the initial scheduling parameters is constructed into a probabilistic causal reasoning model, the probabilistic causal reasoning model uses a Student-t distribution to construct the probability distribution of each initial scheduling parameter, the Student-t distribution includes a degree of freedom, a location parameter and a scale parameter, and a parameter probability representation is obtained;

[0144] The parameter probability representation is input into a multi-level causal graph, the multi-level causal graph includes a power system overall constraint layer, a subsystem coupling relationship layer and a scheduling parameter interaction layer, a time sequence dependence relationship between the initial scheduling parameters is calculated based on a time sequence characteristic function and a corresponding weight, and a parameter coupling strength is calculated by a coupling degree measurement function;

[0145] A joint distribution between any two initial scheduling parameters is constructed based on a Copula function, a causal direction strength is obtained by calculating the conditional probability log ratio value of the joint distribution and an independent distribution, and a multi-dimensional causal strength matrix is constructed by combining the causal direction strength and the parameter coupling strength;

[0146] The conditional probability log ratio value between any two initial scheduling parameters is calculated, an information gain measurement is constructed based on the conditional probability log ratio value, and the causal strength confidence is determined as the ratio of the information gain measurement to the conditional probability log ratio value of the initial scheduling parameters;

[0147] A time-varying causal strength function is constructed including a historical causal strength term, an incremental term and a real-time feedback term, an optimization objective function is constructed based on the causal strength confidence and the time-varying causal strength function, and a causal strength quantization distribution is obtained by an expected value of the conditional probability log ratio value.

[0148] In one specific implementation, the first parameter probability representation of the initial scheduling parameters is constructed as a probabilistic causal inference model. For a set of initial scheduling parameters, including a source power sequence PS = [4.5 kW, 4.8 kW,..., 5.0 kW], an energy storage charging power sequence PB = [1.2 kW, 1.5 kW,..., 1.8 kW], a load adjustment parameter sequence PL = [7.8 kW, 8.1 kW,..., 6.2 kW], and a grid exchange power sequence PG = [2.1 kW, 2.3 kW,..., 4.5 kW], a Student-t distribution is constructed for each parameter sequence. Compared with Gaussian distribution, Student-t distribution has thicker tails, which can better model extreme events and outliers.

[0149] For the source power PS, its Student-t distribution parameters are determined based on historical data statistical analysis: a degree of freedom v = 5 (smaller degree of freedom makes the distribution tail thicker), a location parameter μ = 4.7 kW (close to the average value), and a scale parameter σ = 0.3 kW (reflects the fluctuation range). Similarly, the corresponding Student-t distribution parameters are determined for other parameter sequences: the distribution parameters of the energy storage power PB are v = 6, μ = 1.5 kW, and σ = 0.2 kW; the distribution parameters of the load parameter PL are v = 8, μ = 7.2 kW, and σ = 0.6 kW; and the distribution parameters of the grid power PG are v = 7, μ = 3.0 kW, and σ = 0.8 kW. The probability representation of each scheduling parameter is constructed through these distribution parameters.

[0150] The parameter probability representation is input into a multi-level causal graph for deep modeling, which is composed of three levels: a power system overall constraint layer, a subsystem coupling relationship layer, and a scheduling parameter interaction layer. In the power system overall constraint layer, a power balance constraint node is set to ensure that PS + PG = PL + PB is balanced in all time periods. In the subsystem coupling relationship layer, associated nodes between the source subsystem, the energy storage subsystem, the load subsystem, and the grid subsystem are established, such as a charging and discharging control node between the source and the energy storage subsystem. In the scheduling parameter interaction layer, causal edges between parameters are directly established, such as PS → PB representing the direct influence of the source power on the energy storage power.

[0151] The time sequence dependency between initial scheduling parameters is calculated based on time sequence characteristic function. The time sequence characteristic is extracted by sliding window method, and the window size is set to 12 time periods (1 hour). For the parameter sequence in the window, the autocorrelation coefficient, cross-correlation coefficient and trend slope are calculated. For example, the cross-correlation coefficient of PS and PB is 0.72, indicating a strong positive correlation; the 2nd order autocorrelation coefficient of PS is 0.65, indicating a strong time sequence dependency. Different types of time sequence characteristics are assigned weights: the cross-correlation coefficient weight is 0.5, the autocorrelation coefficient weight is 0.3, and the trend characteristic weight is 0.2. Based on the weighted sum of characteristic values and weights, the time sequence dependency strength of parameters on PS→PB is calculated as 0.68.

[0152] The parameter coupling strength is calculated by coupling measure function, which uses distance-based method to calculate the degree of cooperation of parameter changes. For any two parameters xi and xj, the Euclidean distance of the standardized change amount is calculated over all time periods, and then converted to coupling strength by exponential transformation. Taking PS and PB as an example, the parameter coupling strength is calculated as 0.75, indicating a strong coupling relationship; the coupling strength of PS and PL is 0.45, indicating a moderate coupling relationship.

[0153] The joint distribution between any two initial scheduling parameters is constructed based on Copula function. Copula function can connect multiple marginal distributions into joint distribution, which is suitable for describing complex dependency structure. For PS and PB, first convert the Student-t distribution of each to uniform distribution random variable UPS and UPB, and then select Gaussian Copula function to construct joint distribution. The Copula parameter ρ = 0.7, indicating a strong positive correlation. Based on the constructed joint distribution, the conditional probability P(PB|PS) = 0.82 and P(PS|PB) = 0.65 are calculated.

[0154] The causal direction strength is calculated by the conditional probability log ratio of joint distribution and independent distribution. For PS→PB direction, the log ratio is 0.65; for PB→PS direction, the log ratio is 0.42. Since the log ratio of PS→PB is larger, the causal direction is determined as PS affecting PB. Repeat the above process to calculate the causal direction strength between all parameter pairs, and combine the previously calculated parameter coupling strength to construct a multi-dimensional causal strength matrix C with dimension [parameter number × parameter number].

[0155] The conditional probability log ratio value between any two initial scheduling parameters is calculated to construct an information gain metric. The information gain metric reflects the degree of contribution of one parameter to the prediction of another parameter. For PS→PB, the information gain is 0.58; for PL→PG, the information gain is 0.63. The ratio of the information gain metric to the conditional probability log ratio value is determined as the causal strength confidence. The causal strength confidence of PS→PB is 0.58 / 0.65=0.89, indicating a high confidence of the causal relationship. The causal strength confidence of PL→PG is 0.63 / 0.67=0.94, also indicating a high confidence.

[0156] A time-varying causal strength function g(C, t) is constructed, which includes a historical causal strength term, an incremental term, and a real-time feedback term. The historical causal strength term adopts an exponentially weighted moving average with a weight coefficient α=0.8; the incremental term reflects the change rate of the current period relative to the previous period; and the real-time feedback term adjusts the causal strength based on the real-time system state. Taking the causal strength of PS→PB as an example, the historical strength is 0.72, the increment is 0.03, and the real-time feedback is 0.01, and the comprehensive time-varying causal strength is 0.76.

[0157] Based on the causal strength confidence and the time-varying causal strength function, an optimization objective function is constructed, which includes a causal relationship maintenance term and a system stability term, with weights of 0.7 and 0.3 respectively. The parameters of the time-varying causal strength function are optimized by the gradient descent method, with a learning rate of 0.01 and an iteration number of 100. During the optimization process, the value of the objective function increases from the initial -145.6 to -78.3, indicating that the quality of causal relationship modeling is significantly improved.

[0158] The causal strength quantitative distribution Q is obtained by the expected value of the conditional probability log ratio value. For the PS→PB relationship, the expected value is 0.71; for the PL→PG relationship, the expected value is 0.75. These values constitute the causal strength quantitative distribution Q, which can be directly used for subsequent optimization scheduling decisions.

[0159] Compared with the prior art, the scheme of the embodiment makes significant improvements in the modeling of the causal relationship between power scheduling parameters. Traditional causal relationship modeling methods mainly rely on linear correlation analysis or simple Bayesian networks, and cannot effectively capture the nonlinear dependence relationship and time-varying characteristics between parameters. For example, the common Granger causality test assumes a linear relationship, which is difficult to adapt to the complex nonlinear dynamics in the power system; and the traditional Bayesian network structure learning algorithm has high computational complexity and poor stability when facing high-dimensional time series data.

[0160] The improvement starting point of the embodiment is to construct a more accurate probability representation and a multi-level causal relationship model. By introducing the Student-t distribution, the modeling problem of abnormal values and extreme events in power parameters is solved; by using a multi-level causal graph, comprehensive modeling from system-level constraints to parameter-level interactions is realized; by applying a Copula function to construct a joint distribution, the limitations of traditional methods in describing complex dependence structures are overcome; and by introducing a time-varying causal strength function, the dynamically changing causal relationship in the power system is effectively captured.

[0161] Actual tests show that, compared with traditional methods, the causal strength quantification accuracy of the present application is improved, and the improvement is more significant under complex working conditions. The power dispatching scheme optimized based on the scheme of the embodiment improves system stability and economy, especially in scenarios with high renewable energy penetration, and improves balance stability. These improvements significantly enhance the ability of the power system to cope with complex working conditions and extreme situations, and provide strong support for achieving more efficient and reliable smart grid dispatching.

[0162] As shown in Figure 3 The multi-level causal graph structure and its performance advantages of the embodiment scheme are shown. The figure intuitively presents three levels: the system overall constraint layer (blue nodes), the subsystem coupling relationship layer (green nodes), and the dispatching parameter interaction layer (red nodes). The size of each node represents its importance weight in the causal network, and the numerical value in the node shows the accurate quantified strength. The line thickness represents the causal relationship strength, and the black thick line represents the strong causal relationship (strength > 0.7), and the gray thin line represents the weak causal relationship (strength 0.3-0.7). As can be seen from the figure, the nodes of the system constraint layer such as "total power generation capacity constraint" (SC1, strength 0.92) and "grid stability constraint" (SC2, strength 0.85) affect the "thermal power subsystem" (SS1, strength 0.73) and "wind power subsystem" (SS3, strength 0.81) of the subsystem layer through multiple paths, and finally conduct to the specific dispatching variables of the parameter layer. The performance indicators below the chart show that, compared with traditional methods, the structure identification accuracy of the embodiment scheme reaches 89.4%, an increase of 23.7%; the cross-layer causal relationship capture rate is 78.6%, an increase of 42.1%; the time series dependence accuracy is 83.7%, an increase of 31.5%; and the abnormal working condition prediction accuracy is 81.2%, an increase of 35.7%. Although the calculation complexity increases slightly (+15%) and the response time also increases to 247.5ms (+18.2%), considering the significant improvement in accuracy and capture ability, these small performance costs are acceptable. The chart comprehensively shows the significant advantages of the multi-level causal graph structure of the embodiment scheme in handling complex interaction relationships in power systems.

[0163] In an alternative embodiment, samples are drawn from the optimized scheduling parameter set, a differential Monte Carlo continuous sampling flow is constructed, the sampling probability density is adjusted based on the neural optimal transport theory, the risk probability distribution is calculated, and the scheduling parameter subset satisfying the risk constraint is screened, which comprises:

[0164] Samples are drawn from the optimized scheduling parameter set to construct an initial sampling set, the initial sampling set is input into a differential Monte Carlo continuous sampling flow, a stochastic differential equation containing a drift term and a Brownian motion diffusion term is constructed in the differential Monte Carlo continuous sampling flow, and the stochastic differential equation is solved to obtain a continuous sampling trajectory;

[0165] The second-order Wasserstein distance between the initial probability density of the continuous sampling trajectory and the target probability density is calculated, a neural network based on the self-attention mechanism is constructed to fit the optimal transport mapping, the adjustment strategy of the sampling probability density is obtained by determining the overall loss function and combining the dynamic step strategy;

[0166] The sampling probability density of the continuous sampling trajectory is adjusted according to the adjustment strategy, the risk probability distribution of the adjusted sampling probability density is calculated, and the risk probability distribution is composed of the weighted combination of the deviation probability of the power generation power and the load power, the probability of the node voltage exceeding the limit, and the probability of the energy storage capacity exceeding the limit;

[0167] The risk probability distribution is compared with the preset system risk threshold, the scheduling parameters satisfying the risk constraint are screened, and an optimized scheduling parameter subset is formed.

[0168] In a specific embodiment, representative samples are selected from an existing optimized scheduling parameter set using a stratified sampling method. The sampling process needs to consider the distribution characteristics of the parameters to ensure that the samples have good representativeness and coverage. The selected samples include the output of the generator set, the charge and discharge power of the energy storage device, and the load power, which are key parameters. These parameters together constitute an initial sampling set. This initial sampling set is input into a pre-constructed differential Monte Carlo continuous sampling flow. In the sampling flow, a stochastic differential equation containing a drift term and a Brownian motion diffusion term is constructed, where the drift term represents the evolution trend of the parameters, and the Brownian motion diffusion term introduces random disturbance to increase the diversity of the samples. A series of continuous sampling trajectories are obtained by solving the stochastic differential equation using numerical methods such as the Euler-Maruyama method, which reflect the dynamic change process of the scheduling parameters over time.

[0169] The second-order Wasserstein distance between the initial probability density of the sampling trajectory and the expected target probability density is calculated. This distance measure reflects the degree of difference between the current sampling distribution and the ideal distribution. Then a neural network based on self-attention mechanism is constructed to fit the optimal transport mapping. This neural network adopts a multi-head self-attention structure, which can capture complex dependencies between parameters. The input layer of the network receives the current sampling data, and through multiple self-attention modules and feedforward networks, it performs feature extraction and mapping learning, and the output layer gives the adjusted parameter distribution. During network training, a total loss function is constructed, including the Wasserstein distance, network parameter regularization term, etc. A dynamic step optimization algorithm such as the adaptive moment estimation algorithm is used to train the network to obtain the adjustment strategy of the probability density.

[0170] According to the adjustment strategy obtained by training, the probability density of the continuous sampling trajectory is dynamically adjusted. The risk probability distribution including three key indicators is calculated by risk assessment of the adjusted probability density: deviation probability of power generation and load power (reflecting the balance degree of supply and demand), node voltage out-of-limit probability (reflecting the voltage safety level), and energy storage capacity out-of-limit probability (reflecting the operation constraints of energy storage devices). These three probability indicators are combined into a comprehensive risk probability distribution through weighted combination, and the weight coefficients are determined according to the importance of each indicator.

[0171] The calculated risk probability distribution is compared with the pre-set system risk threshold. Only the scheduling parameters with a risk level below the threshold are retained, and these parameters form an optimized scheduling parameter subset as a candidate solution for subsequent optimization scheduling.

[0172] Exemplarily, assume that a power distribution system includes 3 distributed generator sets, 2 energy storage systems, and several key loads. The initial scheduling parameter set includes the output curves of these devices within 24 hours, charging and discharging plans, and other data. Through hierarchical sampling, 100 representative day-ahead scheduling schemes are selected as the initial sampling set. These samples are input into the differential Monte Carlo sampling flow to obtain 1000 continuous sampling trajectories by solving stochastic differential equations. During the neural network learning process, the multi-head self-attention mechanism focuses on different dimensional features such as power generation, energy storage state, and load prediction. The network contains 3 self-attention modules, each with 8 attention heads. After training, the sampling trajectories are adjusted, and when calculating the risk probability, the weights of the three indicators of generation-load deviation, voltage out-of-limit, and energy storage out-of-limit are set to 0.4, 0.3, and 0.3, respectively. The system risk threshold is set to 0.1, and finally 200 scheduling parameter schemes with a risk probability below the threshold are selected to form an optimized scheduling parameter subset. These schemes not only meet the system safety constraints, but also have good economic efficiency and operability.

[0173] In this embodiment, by introducing a stochastic differential equation containing drift term and diffusion term, continuous sampling trajectories are generated, breaking through the limitation of traditional discrete sampling, improving the exploration depth of scheduling parameter space and the search ability of global optimal solution; the optimal transport mapping is fitted by using the self-attention mechanism, combined with the Wasserstein distance and the dynamic step strategy, to realize the high-precision dynamic adjustment of the sampling probability density, ensure that the sampling is more concentrated in the potential high-quality solution area, and improve the calculation efficiency; the risk probability distribution of multiple sources such as deviation of power generation and load, node voltage out-of-limit, and energy storage capacity out-of-limit is constructed, and the comprehensive safety of the scheduling scheme is realized, which ensures the stability and safety of system operation; by comparing with the preset risk threshold, the parameter subset with controllable risk is selected from the optimized scheduling parameters, which ensures that the final scheduling scheme meets the operation constraints while having high confidence and risk tolerance; innovatively combining stochastic differential equation and neural network optimization, integrating physical modeling and machine learning technology, the modeling expression and adaptability of the scheduling method are enhanced, which provides theoretical and practical support for intelligent scheduling of complex energy systems.

[0174] In an optional embodiment, the second-order Wasserstein distance between the initial probability density of the continuous sampling trajectory and the target probability density is calculated, a neural network based on the self-attention mechanism is constructed to fit the optimal transport mapping, and the adjustment strategy of the sampling probability density is obtained by determining the overall loss function and combining the dynamic step strategy, which includes:

[0175] The initial probability density function is obtained by kernel density estimation of the continuous sampling trajectory, and the second-order Wasserstein distance between the initial probability density function and the target probability density function is calculated;

[0176] A deep neural network is constructed according to the second-order Wasserstein distance as an optimization transport mapping function, the input data is converted into a query matrix, a key matrix and a value matrix, the dot product of the query matrix and the key matrix is calculated to obtain an attention weight, and the attention weight is multiplied by the value matrix to obtain a feature representation;

[0177] The transport error loss between the optimization transport mapping function and the theoretical optimal transport mapping is calculated, the gradient regularization loss of the optimization transport mapping function is calculated, the cyclic consistency loss of the optimization transport mapping function is calculated, and the transport error loss, the gradient regularization loss and the cyclic consistency loss are combined to obtain the overall loss function;

[0178] The optimization gradient is calculated according to the overall loss function, the dynamic step is calculated by using the exponential function of the initial step and the decay factor, and the minimum step threshold is set, the sampling points are iteratively updated by using the dynamic step, and the adjustment strategy of the sampling probability density is obtained.

[0179] In one specific implementation, a kernel density estimation is performed on the continuous sampling trajectory to obtain an initial probability density function. Specifically, for a given set of sampling points, a kernel density estimation method is used to construct the initial probability density function. In one embodiment, a Gaussian kernel function is used as the kernel function, and a Gaussian distribution is assigned to each sampling point. The Gaussian distributions of all sampling points are then added and normalized to obtain a smooth probability density function. For example, for a dataset containing 1000 two-dimensional sampling points, a bandwidth parameter of 0.1 can be selected, and a Gaussian kernel density estimation method is applied to generate the initial probability density function.

[0180] The second-order Wasserstein distance between the initial probability density function and the target probability density function is calculated. The initial probability density function is denoted as P, and the target probability density function is denoted as Q. The second-order Wasserstein distance between P and Q is calculated by solving the optimal transport problem. In practical implementation, the Sinkhorn algorithm can be used for approximate calculation. For example, the regularization parameter of the Sinkhorn algorithm is set to 0.01, and the maximum number of iterations is set to 100. The calculated Wasserstein distance value is 0.85.

[0181] A deep neural network is constructed according to the second-order Wasserstein distance as an optimization transport mapping function. The neural network uses a Transformer architecture and includes a multi-head self-attention mechanism. Specifically, the network structure includes 3 self-attention layers, each with 8 attention heads and a hidden layer dimension of 256. After linear transformation of the input data, the data is converted into a query matrix, a key matrix, and a value matrix, each with a dimension of batch size x sequence length x hidden layer dimension. For example, for input data with a batch size of 32, after linear transformation, the query matrix, key matrix, and value matrix have dimensions of 32x100x256.

[0182] The dot product of the query matrix and the key matrix is calculated to obtain the attention score, which is scaled by the square root of the key vector dimension, and the Softmax function is applied to obtain the attention weight. For example, for a query matrix and a key matrix with dimensions of 32x100x256, the dot product is calculated to obtain an attention score with a dimension of 32x100x100, which is scaled by 16 (i.e., the square root of 256) and normalized by the Softmax function. The attention weight is multiplied by the value matrix to obtain the weighted value, and the outputs of multiple attention heads are connected and passed through a feedforward network to obtain the final feature representation.

[0183] During the training process, the transmission error loss between the optimized transmission mapping function and the theoretical optimal transmission mapping needs to be calculated. Specifically, the transmission error loss is defined as the mean absolute error between the neural network output and the theoretical optimal transmission mapping result. In one embodiment, the theoretical optimal transmission mapping can be solved by a linear programming method, and then the error between the neural network output and the theoretical optimal solution is calculated. For example, for a test sample set, the average absolute error between the transmission mapping result predicted by the neural network and the theoretical optimal solution is 0.05.

[0184] At the same time, the gradient regularization loss of the optimized transmission mapping function is calculated, which constrains the gradient of the transmission mapping function to ensure that the mapping function satisfies the Lipschitz continuity condition. Specifically, the difference between the gradient norm of the mapping function on a pair of randomly sampled input points and a preset threshold (such as 1.0) is taken as the gradient regularization loss. For example, during the training process, 100 pairs of input points can be randomly sampled, and the average value of the gradient norm on these pairs of points is 1.2, then the gradient regularization loss is 0.2.

[0185] In addition, the cycle consistency loss of the optimized transmission mapping function is calculated to ensure that the mapping function has good bidirectional consistency between the source distribution and the target distribution. Specifically, the points in the source distribution are mapped to the target distribution through the mapping function, and then mapped back to the source distribution through the inverse mapping function, and the distance between the original points and the cyclically mapped points is calculated as the cycle consistency loss. For example, for 100 sampled points in the source distribution, the average Euclidean distance between the original points and the cyclically mapped points after forward mapping and inverse mapping is 0.08.

[0186] The transmission error loss, gradient regularization loss and cycle consistency loss are combined to obtain the overall loss function. In one embodiment, the three losses can be combined in the form of weighted sum, and the weights are 0.5, 0.3 and 0.2 respectively. For example, when the transmission error loss is 0.05, the gradient regularization loss is 0.2, and the cycle consistency loss is 0.08, the overall loss function value is 0.05x0.5+0.2x0.3+0.08x0.2=0.089.

[0187] According to the overall loss function, the optimization gradient is calculated, and the dynamic step size is calculated using the exponential function of the initial step size and the decay factor. Specifically, the initial step size is set to 0.01, the decay factor is set to 0.95, and the step size is multiplied by the decay factor every 100 iterations. At the same time, the minimum step size threshold is set to 0.0001 to ensure that the step size does not become too small to cause training stagnation. The dynamic step size is used to iteratively update the sampling points to obtain the adjustment strategy of the sampling probability density. For example, after 1000 iterations, the step size gradually decreases from the initial 0.01 to 0.0058, and the sampling points gradually adjust from the initial distribution to a state close to the target distribution.

[0188] Traditional optimal transport methods usually rely on linear programming or semi-discrete optimal transport algorithms, which have high computational complexity and are difficult to apply to large-scale datasets. For example, the computational complexity of the optimal transport algorithm based on linear programming is O(n 3 log(n)), where n is the number of sampling points. When the number of sampling points reaches ten thousand, the computational overhead becomes unacceptable. While existing neural network-based methods have reduced computational complexity, they have deficiencies in ensuring that the mapping function meets the conditions of optimal transport theory, resulting in a large deviation between the generated results and the theoretical optimal solution.

[0189] The method proposed in this embodiment fuses kernel density estimation, Wasserstein distance and deep learning technology, constructs a neural network with self-attention mechanism as the optimization transport mapping function, and designs a joint optimization objective including transport error loss, gradient regularization loss and cycle consistency loss, effectively ensuring the theoretical properties of the mapping function. At the same time, a dynamic step strategy is used for iterative optimization, avoiding the problem of gradient explosion or disappearance in the training process. Experimental results show that compared with traditional methods, the method of this embodiment reduces the computational complexity to O(nlog(n)), and when processing 100,000 sampling points, the computation time is reduced from several hours to minutes, while maintaining a high consistency between the mapping results and the theoretical optimal solution, and reducing the average error rate.

[0190] As Figure 4As shown, the whole process and effect comparison of sampling distribution optimization are shown. The upper left graph shows the initial sampling distribution, and it can be seen that there is obvious local aggregation phenomenon, mainly concentrated in the three regions of (-1.5, 1.5), (1.5, 1.5) and (0, -1.5), the density variance is as high as 0.432, the coverage rate is only 0.674, and the sampling efficiency is as low as 0.531, indicating that the initial sampling is seriously uneven; the upper right graph shows the ideal target distribution, which presents a uniform circular distribution, the density variance is only 0.051, and the coverage rate and sampling efficiency are both more than 0.97; the lower left graph shows the result after optimization by the traditional method, although the initial distribution is improved, there is still obvious density unevenness, the density variance is reduced to 0.187, the coverage rate is improved to 0.836, and the sampling efficiency reaches 0.763, but there is still a large gap with the target distribution; the lower right graph shows the result after optimization by the scheme of the embodiment, the sampling point distribution is highly uniform and almost consistent with the target distribution, the density variance is only 0.067 (decreased by 64.2% compared with the traditional method), the coverage rate is 0.954 (increased by 14.1%), and the sampling efficiency is 0.947 (increased by 24.1%). The scheme of the embodiment effectively solves the local aggregation problem that the traditional method cannot handle by fusing kernel density estimation, Wasserstein distance calculation and optimized transport mapping, and realizes the efficient and uniform distribution of sampling points, which is of great significance to improve the accuracy and computational efficiency of subsequent simulation analysis.

[0191] The AI large model source network load storage optimization scheduling system of the embodiment of the application comprises:

[0192] The first unit is configured to receive source power generation data, power grid transmission data, user load data and energy storage device data in a power system, and obtain a system feature data set through time tag alignment and data preprocessing;

[0193] The second unit is configured to divide peak period, valley period, new energy consumption period, load response period and extreme weather period based on the system feature data set, identify and determine a current operation scenario through a multi-dimensional scene identifier of energy flow decomposition, call a corresponding expert model combination, and generate initial scheduling parameters;

[0194] The third unit is configured to input the initial scheduling parameters into a pre-trained large language model, construct a causal strength quantization distribution through a probability causal inference of a variational entropy coding, perform bias evaluation, and generate an optimized scheduling parameter set;

[0195] The fourth unit is configured to extract samples from the optimized scheduling parameter set, construct a differential Monte Carlo continuous sampling flow, adjust a sampling probability density based on a neural optimal transport theory, calculate a risk probability distribution, and screen a scheduling parameter subset meeting a risk constraint;

[0196] A fifth unit is configured to select optimal dispatch parameters from the dispatch parameter subset according to a real-time electricity market price signal and system operation constraints, and perform source-grid-load-storage coordinated dispatch;

[0197] A sixth unit is configured to feed back evaluation parameters of the dispatch execution result to the expert model and the large language model, and perform online iterative optimization.

[0198] A third aspect of the embodiment of the application,

[0199] An electronic device is provided, comprising:

[0200] A processor;

[0201] A memory for storing processor-executable instructions;

[0202] The processor is configured to invoke the instructions stored in the memory to execute the method described above.

[0203] A fourth aspect of the embodiment of the application,

[0204] A computer-readable storage medium is provided, which stores computer program instructions, and the computer program instructions are executed by a processor to implement the method described above.

[0205] The present application can be a method, device, system and / or computer program product. The computer program product can include a computer-readable storage medium having stored thereon computer-readable program instructions that, when executed by a computer, cause the computer to carry out various aspects of the present application.

[0206] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.

Claims

1. AI large-scale model source-grid-load-storage optimization scheduling method, characterized by: include: Receive source power generation data, grid transmission data, user load data, and energy storage device data from the power system, and obtain a system feature data set through time tag alignment and data preprocessing; Based on the system characteristic data set, the system divides the peak period, off-peak period, new energy consumption period, load response period and extreme weather period. The multi-dimensional scenario identifier of energy flow decomposition is used to identify the current operating scenario, call the corresponding expert model combination, and generate the initial scheduling parameters. The initial scheduling parameters are fed into a pre-trained large language model, and a quantitative distribution of causal strength is constructed through probabilistic causal reasoning using variational entropy coding. Deviation evaluation is performed to generate an optimized scheduling parameter set. Extract samples from the optimized scheduling parameter set, construct a differential Monte Carlo continuous sampling stream, adjust the sampling probability density based on the neural optimal transmission theory, calculate the risk probability distribution, and select the scheduling parameter subset that meets the risk constraints; Based on real-time electricity market price signals and system operation constraints, the optimal dispatching parameters are selected from the dispatching parameter subset to perform source-grid-load-storage coordinated dispatching. The evaluation parameters of the scheduling execution results are fed back to the expert model and large language model for online iterative optimization.

2. The method according to claim 1, characterized in that The multi-dimensional scenario identifier of energy flow decomposition identifies the current operating scenario, calls the corresponding expert model combination, and generates the initial scheduling parameters including: Constructing source power, grid power, load power, and energy storage power into corresponding energy flow feature vectors, performing time dimension decomposition on the energy flow feature vectors to obtain energy flow components, and combining the energy flow components to form an energy flow feature matrix; Constructing an information entropy calculation model based on the energy flow characteristic matrix, establishing an optimization objective function for maximizing information entropy, solving the optimization objective function to obtain the optimal weight coefficient of each energy flow component, and combining the product of the optimal weight coefficient and the corresponding energy flow component into a multidimensional scene characteristic matrix; Based on the multi-dimensional scene feature matrix, a Gaussian mixture probability model is constructed through attention-enhanced feature extraction. The expectation-maximization algorithm is used to optimize the model parameters, and the probability distribution of the operation scenario in each time period is established. The time period with the maximum probability value is determined as the current operation scenario type. Calling an expert model combination from the expert model library corresponding to each time period according to the current operation scenario type, extracting the historical scheduling data of the expert model combination, and calculating the historical operation error sequence of each expert model; Performing exponential transformation and normalization processing on the historical operation error sequence to generate expert model combination weight coefficients, and inputting the multi-dimensional scenario feature matrix into the expert model combination to obtain a scheduling instruction set; The scheduling instruction set is weighted and combined according to the expert model combination weight coefficient to generate the initial scheduling parameters.

3. The method according to claim 2, characterized in that Based on the multi-dimensional scene feature matrix, a Gaussian mixture probability model is constructed through attention-enhanced feature extraction. The expectation maximization algorithm is used to optimize the model parameters, and the probability distribution of the operating scenarios in each time period is established. The time period with the maximum probability value is determined as the current operating scenario type, including: The multi-dimensional scene feature matrix is ​​input into a multi-head attention network, which calculates multiple sets of attention features through a learnable projection matrix and concatenates the multiple sets of attention features to obtain scene-related features; Inputting the scene-related features into a conditional variational autoencoder, wherein the conditional variational autoencoder encodes the scene-related features into a posterior distribution of a first scene variable under a time period constraint; Inputting the first scenario variable into a recurrent neural network to obtain a second scenario variable, and constructing a Gaussian mixture probability model for each time period based on the second scenario variable, including a time period distribution weight, a time period distribution mean, and a time period distribution covariance; The posterior probability of the distribution of each time period in the Gaussian mixture probability model is calculated by the expectation maximization algorithm, the corresponding relationship between the posterior probability and the time period is multiplied and summed to obtain the operation scenario probability distribution of each time period, and the time period with the maximum probability value in the operation scenario probability distribution is determined as the current operation scenario type.

4. The method according to claim 1, wherein The initial scheduling parameters are input into the pre-trained large language model, and the causal strength quantization distribution is constructed through probabilistic causal reasoning with variational entropy coding. Deviation evaluation is performed to generate the optimized scheduling parameter set including: The initial scheduling parameters are input into the pre-trained large language model, and the probability representation of the initial scheduling parameters is constructed using variational entropy coding. The probability representation of the first parameter in the form of a Gaussian distribution is generated through the encoding mapping function; Based on the first parameter probability representation, a hierarchical probabilistic causal inference model is constructed to calculate the conditional probability logarithm ratio and parameter coupling strength between the initial scheduling parameters, and the quantitative distribution of causal strength is obtained by optimizing the time-varying causal strength function. The bi-norm deviation between the initial scheduling parameters and the expected target value is calculated. Combined with the quantized distribution of causal strength, a variational evidence lower bound function containing a reconstruction error term and a KL divergence regularization term is constructed. The variational entropy coding parameters are iteratively optimized by maximizing the variational evidence lower bound. Monte Carlo sampling is performed on the probability representation of the second parameter to obtain candidate parameters. It is verified whether the candidate parameters meet the power balance constraints of the power generation load, the charge and discharge capacity constraints of the energy storage equipment, the unit ramp rate constraints, and the grid node voltage limit constraints. The candidate parameters that meet all the constraints are combined into an optimized scheduling parameter set.

5. The method according to claim 4, characterized in that Based on the first parameter probability representation, a hierarchical probabilistic causal inference model is constructed to calculate the conditional probability logarithm ratio and parameter coupling strength between the initial scheduling parameters. The quantitative distribution of causal strength is obtained by optimizing the time-varying causal strength function, including: Constructing a first parameter probability representation of the initial scheduling parameters into a probabilistic causal inference model, wherein the probabilistic causal inference model uses a Student-t distribution to construct a probability distribution of each initial scheduling parameter, wherein the Student-t distribution includes degrees of freedom, a location parameter, and a scale parameter to obtain a parameter probability representation; Inputting the parameter probability representation into a multi-level causal graph, the multi-level causal graph includes a power system overall constraint layer, a subsystem coupling relationship layer, and a scheduling parameter interaction layer, calculating the temporal dependency between the initial scheduling parameters based on a temporal characteristic function and corresponding weights, and calculating the parameter coupling strength using a coupling metric function; Based on the Copula function, a joint distribution between any two initial scheduling parameters is constructed, and the conditional probability logarithm ratio of the joint distribution to the independent distribution is calculated to obtain the causal direction strength. The causal direction strength is combined with the parameter coupling strength to construct a multidimensional causal strength matrix. Calculating a conditional probability logarithm ratio between any two initial scheduling parameters, constructing an information gain metric based on the conditional probability logarithm ratio, and determining a ratio of the information gain metric to the conditional probability logarithm ratio of the initial scheduling parameters as a causal strength confidence level; A time-varying causal strength function including a historical causal strength term, an incremental term and a real-time feedback term is constructed. An optimization objective function is constructed based on the causal strength confidence and the time-varying causal strength function. The quantitative distribution of causal strength is obtained through the expected value of the conditional probability logarithm ratio.

6. The method according to claim 1, characterized in that Samples are extracted from the optimized scheduling parameter set, a differential Monte Carlo continuous sampling stream is constructed, the sampling probability density is adjusted based on the neural optimal transmission theory, the risk probability distribution is calculated, and the scheduling parameter subset that meets the risk constraints is screened, including: Extracting samples from the optimized scheduling parameter set to construct an initial sampling set, inputting the initial sampling set into a differential Monte Carlo continuous sampling stream, constructing a stochastic differential equation containing a drift term and a Brownian motion diffusion term in the differential Monte Carlo continuous sampling stream, and solving the stochastic differential equation to obtain a continuous sampling trajectory; The second-order Wasserstein distance between the initial probability density and the target probability density of the continuous sampling trajectory is calculated, and a neural network based on the self-attention mechanism is constructed to fit the optimal transmission mapping. By determining the overall loss function and combining it with a dynamic step size strategy, an adjustment strategy for the sampling probability density is obtained. Adjusting the sampling probability density of the continuous sampling trajectory according to the adjustment strategy, and calculating the risk probability distribution of the adjusted sampling probability density, wherein the risk probability distribution is composed of a weighted combination of the probability of deviation between generated power and load power, the probability of node voltage exceeding a limit, and the probability of energy storage capacity exceeding a limit; The risk probability distribution is compared with a preset system risk threshold, and scheduling parameters that meet the risk constraints are screened to form an optimized scheduling parameter subset.

7. The method according to claim 6, characterized in that The second-order Wasserstein distance between the initial probability density and the target probability density of the continuous sampling trajectory is calculated, and a neural network based on the self-attention mechanism is constructed to fit the optimal transmission mapping. By determining the overall loss function and combining it with the dynamic step size strategy, the adjustment strategy for the sampling probability density is obtained, including: Perform kernel density estimation on the continuous sampling trajectory to obtain an initial probability density function, and calculate the second-order Wasserstein distance between the initial probability density function and the target probability density function; A deep neural network is constructed based on the second-order Wasserstein distance as an optimized transfer mapping function, which converts the input data into a query matrix, a key matrix, and a value matrix. The dot product of the query matrix and the key matrix is ​​calculated to obtain the attention weight, and the attention weight is multiplied by the value matrix to obtain the feature representation. Calculating a transmission error loss between an optimized transmission mapping function and a theoretical optimal transmission mapping, calculating a gradient regularization loss of the optimized transmission mapping function, calculating a cycle consistency loss of the optimized transmission mapping function, and combining the transmission error loss, the gradient regularization loss, and the cycle consistency loss to obtain an overall loss function; The optimization gradient is calculated according to the overall loss function, the dynamic step size is calculated using the exponential function of the initial step size and the attenuation factor, and the minimum step size threshold is set. The sampling points are iteratively updated using the dynamic step size to obtain an adjustment strategy for the sampling probability density.

8. An AI large-scale model source-grid-load-storage optimization scheduling system, used to implement the method according to any one of claims 1 to 7, characterized in that: include: The first unit is used to receive source power generation data, grid transmission data, user load data, and energy storage device data in the power system, and obtain a system feature data set through time tag alignment and data preprocessing; The second unit is used to divide the system characteristic data set into peak hours, off-peak hours, new energy consumption hours, load response hours, and extreme weather hours. It uses the multi-dimensional scenario identifier of energy flow decomposition to identify the current operating scenario, calls the corresponding expert model combination, and generates initial scheduling parameters. The third unit is used to input the initial scheduling parameters into the pre-trained large language model, construct the quantized distribution of causal strength through probabilistic causal reasoning with variational entropy coding, perform deviation evaluation, and generate the optimized scheduling parameter set; The fourth unit is used to extract samples from the optimized scheduling parameter set, construct a differential Monte Carlo continuous sampling flow, adjust the sampling probability density based on the neural optimal transmission theory, calculate the risk probability distribution, and select the scheduling parameter subset that meets the risk constraints; The fifth unit is used to select the optimal scheduling parameters from the scheduling parameter subset based on real-time power market price signals and system operation constraints, and perform source-grid-load-storage coordinated scheduling; The sixth unit is used to feed back the evaluation parameters of the scheduling execution results to the expert model and the large language model for online iterative optimization.

9. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Micro-grid intelligent scheduling method and system based on AI large model

    CN121012127A

  • Digital path dynamic generation method and system

    CN121659604A

  • A method and system for dynamic generation of digital paths

    CN121659604B