Generation method and device for source load uncertainty datamation representation

Through the fuzzy clustering method of Chebishev distance and Manhattan distance, cluster clusters representing extreme fluctuation output scenarios are generated, which solves the problem of low accuracy of source load uncertainty description information in the prior art, realizes a full reflection of the diversity and volatility of new energy data, and improves the accuracy of power system optimization.

CN120354159APending Publication Date: 2025-07-22EAST CHINA BRANCH OF STATE GRID CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510274123.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The existing cluster analysis methods cannot effectively process sparse data and extreme data, resulting in low accuracy in source load uncertainty description information and cannot fully reflect the diversity and volatility of renewable energy data.

Method used

The fuzzy clustering method of Chebishev distance and Manhattan distance is used to fuzzy cluster the historical source load output time sequence data to generate multiple cluster clusters that characterize extreme fluctuation output scenarios, and the source load uncertainty data-based representation information is obtained through the success rate deviation probability distribution function.

Benefits of technology

Effectively process sparse data and extreme data, fully reflect the diversity and volatility of new energy data, improve the accuracy of uncertainty description, and provide a more flexible and accurate data basis for power system optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120354159A_ABST
    Figure CN120354159A_ABST
Patent Text Reader

Abstract

The invention discloses a source load uncertainty datamation representation generation method and device, relates to the technical field of power system optimization, and mainly aims to solve the problem of low accuracy of source load uncertainty datamation representation. The method mainly comprises the steps that multiple pieces of historical source load output time sequence data of target new energy are acquired, and each piece of historical source load output time sequence data comprises multi-period output data; fuzzy clustering is carried out on the historical source load output time sequence data according to a preset clustering distance calculation strategy, a clustering result is obtained, the preset clustering distance calculation strategy comprises a Chebyshev distance calculation strategy, and the clustering result comprises a plurality of clustering clusters representing an extreme fluctuation output scene; and for each cluster in the clustering result, generating a power deviation probability distribution function, and obtaining source load uncertainty datamation representation information of the target new energy meeting different scenes. The method is mainly used for generating source load uncertainty datamation representation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of power system optimization, and particularly to a method and device for generating a digital representation of the uncertainty of power sources and loads. Background Art

[0002] In the fields of energy management, power system optimization, and load forecasting, the uncertainty of power sources and loads (i.e., power sources and loads) is an important consideration factor. In order to accurately describe and quantify this uncertainty, a digital representation method is particularly important. Among them, clustering analysis, as an unsupervised learning method, is widely used for classifying and extracting features from the uncertainty data of power sources and loads.

[0003] Existing clustering analysis methods cannot effectively process sparse data and extreme data, resulting in a decrease in the accuracy of clustering results in extreme cases and an inability to fully reflect the diversity and volatility of renewable energy data, resulting in low accuracy of the obtained uncertainty description information. Summary of the Invention

[0004] In view of this, the present invention provides a method and device for generating a digital representation of the uncertainty of power sources and loads, mainly aiming to solve the problem of low accuracy of uncertainty description information when existing clustering analysis methods process sparse data and extreme data.

[0005] According to one aspect of the present invention, a method for generating a digital representation of the uncertainty of power sources and loads is provided, including:

[0006] Obtaining multiple pieces of historical power output time series data of a target new energy source, where the time scales of each piece of the historical power output time series data are the same, and each piece of the historical power output time series data includes multi-period power output data;

[0007] Performing fuzzy clustering on the historical power output time series data according to a preset clustering distance calculation strategy to obtain a clustering result, where the preset clustering distance calculation strategy includes a Chebyshev distance calculation strategy, and the clustering result includes multiple clustering clusters representing extreme fluctuation power output scenarios;

[0008] For each clustering cluster in the clustering result, generating a power deviation probability distribution function respectively to obtain the digital representation information of the uncertainty of power sources and loads of the target new energy source under different scenarios.

[0009] Further, the performing fuzzy clustering on the historical power output time series data according to a preset clustering distance calculation strategy to obtain a clustering result includes:

[0010] Initializing a type of clustering center and a type of membership function value of each piece of historical power output time series data;

[0011] For each piece of the historical source-load output time-series data, calculate the Chebyshev distance between the historical source-load output time-series data and each first-class clustering center, and update the first-class membership function values based on the Chebyshev distance to obtain a first-class membership matrix;

[0012] Update the first-class clustering centers according to the first-class membership matrix of the historical source-load output time-series data to obtain updated first-class clustering centers;

[0013] Return to the step of calculating the Chebyshev distance between the historical source-load output time-series data and each first-class clustering center, and iteratively update the first-class membership matrix and the first-class clustering centers until the iteration stop condition is met, to obtain multiple clustering clusters representing extreme fluctuation output scenarios.

[0014] Further, in the first-class membership matrix, the membership formula for each piece of historical source-load output time-series data with respect to the first-class clustering center is:

[0015] where, u ij is the membership of the i-th piece of historical source-load output time-series data x i with respect to the j-th first-class clustering center v 1j , u ij ∈[0,1]; v 1k is any first-class clustering center, c1 is the first confidence level, m1 is the fuzzy index, V1 is the total number of first-class clustering centers, and N is the total number of pieces of historical source-load output time-series data.

[0016] Further, the preset clustering distance calculation strategy further includes a Manhattan distance calculation strategy, and the clustering result further includes multiple clustering clusters representing different typical output scenarios. Fuzzy clustering the historical source-load output time-series data according to the preset clustering distance calculation strategy to obtain a clustering result, including:

[0017] Initialize the second-class clustering centers and the second-class membership function values of each piece of historical source-load output time-series data;

[0018] For each piece of the historical source-load output time-series data, calculate the Manhattan distance between the historical source-load output time-series data and each second-class clustering center, and update the second-class membership function values based on the Manhattan distance to obtain a second-class membership matrix;

[0019] Update the second-class clustering centers according to the second-class membership matrix of the historical source-load output time-series data to obtain updated second-class clustering centers;

[0020] Return to the step of calculating the Manhattan distances between the historical source-load output time series data and each of the secondary clustering centers, and iteratively update the secondary membership matrix and the secondary clustering centers until the iteration stop condition is met, to obtain multiple clustering clusters representing typical output scenarios.

[0021] Further, in the secondary membership matrix, the membership degree calculation formula of each historical source-load output time series data with respect to the secondary clustering center is:

[0022] where, u ij ′ is the membership degree of the i-th historical source-load output time series data x i with respect to the j-th secondary clustering center v 2j , v 2k is any secondary clustering center, c2 is the second confidence level, m2 is the second fuzzy index, and V2 is the total number of secondary clustering centers.

[0023] Further, for each of the clustering clusters in the clustering result, generating a power deviation probability distribution function respectively, includes:

[0024] For each of the clustering clusters, calculate the power difference between each historical source-load output time series data and the corresponding clustering center in each time period, to obtain a high-dimensional power deviation set;

[0025] Expand the high-dimensional power deviation set into a one-dimensional array, and respectively calculate the probability density function of each element in the one-dimensional array according to the kernel function smoothing algorithm;

[0026] Construct the power deviation probability distribution function of the clustering cluster according to the probability density functions of the respective elements, to obtain the power deviation probability distribution functions of each clustering cluster.

[0027] Further, after obtaining the source-load uncertainty data representation information of the target new energy satisfying different scenarios, the method further includes:

[0028] Store the source-load uncertainty data representation information into the target storage space, where the target storage space stores the source-load uncertainty data representation information of different types of new energy within different power jurisdiction areas;

[0029] In response to the retrieval instruction of the source-load uncertainty information, determine at least one new energy to be optimized according to the power system optimization task carried by the retrieval instruction, where the power system optimization task includes power load forecasting task, new energy access management task, power system risk assessment task, distributed energy system optimization task, intelligent microgrid management task, energy trading and pricing task;

[0030] Send the source-load uncertainty data representation information that matches the new energy to be optimized to the target terminal, so that the target terminal executes the power system optimization task according to the source-load uncertainty data representation information.

[0031] According to another aspect of the present invention, there is provided a device for generating a source-load uncertainty data representation, including:

[0032] An acquisition module, configured to acquire multiple pieces of historical source-load output time series data of the target new energy, wherein the time scales of each piece of the historical source-load output time series data are the same, and each piece of the historical source-load output time series data includes multi-period output data;

[0033] A clustering module, configured to perform fuzzy clustering on the historical source-load output time series data according to a preset clustering distance calculation strategy to obtain a clustering result, wherein the preset clustering distance calculation strategy includes a Chebyshev distance calculation strategy, and the clustering result includes multiple clustering clusters representing extreme fluctuation output scenarios;

[0034] A generation module, configured to respectively generate a power deviation probability distribution function for each clustering cluster in the clustering result to obtain the source-load uncertainty data representation information of the target new energy that meets different scenarios.

[0035] According to still another aspect of the present invention, there is provided a storage medium in which at least one executable instruction is stored, and the executable instruction causes the processor to perform operations corresponding to the method for generating the source-load uncertainty data representation as described above.

[0036] According to still another aspect of the present invention, there is provided a terminal, including: a processor, a memory, a communication interface, and a communication bus, and the processor, the memory, and the communication interface complete communication with each other through the communication bus;

[0037] The memory is used to store at least one executable instruction, and the executable instruction causes the processor to perform operations corresponding to the method for generating the source-load uncertainty data representation as described above.

[0038] By means of the above technical solutions, the technical solutions provided by the embodiments of the present invention at least have the following advantages:

[0039] The present invention provides a method and device for generating a digital representation of source-load uncertainty. In the embodiments of the present invention, multiple pieces of historical source-load output time series data of a target new energy source are acquired. Among them, the time scales of each piece of historical source-load output time series data are the same, and each piece of historical source-load output time series data includes multi-period output data. Fuzzy clustering is performed on the historical source-load output time series data according to a preset clustering distance calculation strategy to obtain a clustering result. Among them, the preset clustering distance calculation strategy includes a Chebyshev distance calculation strategy, and the clustering result includes multiple clustering clusters representing extreme fluctuation output scenarios. For each clustering cluster in the clustering result, a power deviation probability distribution function is generated respectively to obtain the digital representation information of the source-load uncertainty of the target new energy source under different scenarios. Through fuzzy clustering based on the Chebyshev distance, the situation where the source-load output data is sparse data and extreme data can be effectively processed, and the diversity and volatility of new energy data can be fully reflected.

[0040] The above description is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present invention more obvious and understandable, the following specifically illustrates the specific embodiments of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention. Moreover, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:

[0042] Figure 1 shows a flowchart of a method for generating a digital representation of source-load uncertainty provided by an embodiment of the present invention;

[0043] Figure 2 shows a block diagram of a device for generating a digital representation of source-load uncertainty provided by an embodiment of the present invention;

[0044] Figure 3 shows a schematic structural diagram of a terminal provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0045] Hereinafter, the exemplary embodiments of the present disclosure will be described in more detail with reference to the drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully conveyed to those skilled in the art.

[0046] In view of the problem that the accuracy of uncertainty description information is relatively low when existing clustering analysis methods are used to process sparse data and extreme data, an embodiment of the present invention provides a method for generating a digital representation of source-load uncertainty data, as Figure 1 shown. The method includes:

[0047] 101. Obtain multiple pieces of historical source-load output time series data of the target new energy.

[0048] In an embodiment of the present invention, the target new energy may be any one of photovoltaic, wind power, and hydropower energy, and the embodiment of the present invention does not make specific limitations. The multiple pieces of historical source-load output time series data are the result of equally dividing the source-load data of the target new energy in the historical time period, that is, the time scales of each piece of historical source-load output time series data are the same. Among them, the time scale may be 1 natural day, 4 natural days, etc., and the embodiment of the present invention does not make specific limitations. Taking the time scale of 1 natural day as an example, the source-load output data of the target new energy in the past 365 days is divided to obtain 365-day historical source-load output time series data. Among them, each piece of historical source-load output time series data includes multi-period output data. The period is usually divided by hour. For a time scale of 1 natural day, it is 24 periods, and for a time scale of 4 natural days, it is 96 periods. Among them, the output data of each period includes the power generation output data on the power generation side of the target new energy or the load output data on the power consumption side.

[0049] 102. Perform fuzzy clustering on the historical source-load output time series data according to a preset clustering distance calculation strategy to obtain a clustering result.

[0050] In an embodiment of the present invention, all the historical source-load output time series data of the target new energy is used as a data set, and each piece of historical source-load output time series data is an independent clustering object in this data set. The fuzzy C-means clustering algorithm is used to perform soft clustering on the data in this data set, that is, it focuses on using membership degrees to distinguish clustering objects. Each clustering object has a membership degree for each cluster class, and the sum of all membership degrees is 1. For each clustering object, the higher the membership degree with respect to a certain cluster class, the higher the similarity, and the closer it is to this cluster class. Among them, the essence of the membership degree is to calculate the ratio of the distance from each clustering object to each clustering center to the sum of the distances from each clustering object to all clustering centers. The distance between the clustering object and the clustering center is calculated according to a preset clustering distance calculation strategy. Each clustering cluster corresponds to different clustering representative scenarios that can reflect the output characteristics of the source-load (such as renewable energy sources such as wind power and photovoltaics) under different load periods. For example, high-output scenarios, low-output scenarios, positive peak shaving scenarios, and negative peak shaving scenarios. Different scenarios correspond to different weather or seasons. For example, the high-output scenario can correspond to spring or windy weather for wind power, and can also correspond to summer or strong sunlight weather for photovoltaic power generation.

[0051] It should be noted that the preset clustering distance calculation strategy includes the Chebyshev distance calculation strategy. That is, the Chebyshev distance between the clustering object and the clustering center is used as the distance for calculating the membership degree. The Chebyshev distance can measure the maximum absolute value of the differences in each dimension between two pieces of data. In the case where the output data varies greatly in the time dimension, the clustering result can reflect the extreme situations in the source-load output dataset, that is, the clustering result includes multiple clustering clusters representing extreme fluctuation output scenarios. For example, the output data of photovoltaic power generation over 24 hours a day is in a fluctuating state. The difference between the hour with the maximum output and the hour with the minimum output in a day is the output fluctuation of the day. Through fuzzy clustering based on the Chebyshev distance calculation strategy, clustering clusters corresponding to the maximum fluctuation output scenario and clustering clusters corresponding to the minimum fluctuation output scenario can be obtained. Existing fuzzy clustering only relies on the Euclidean distance for clustering, resulting in a decrease in accuracy in extreme situations, and the membership degree calculation is only based on the Euclidean distance, which cannot fully reflect the diversity and volatility of new energy data. Through fuzzy clustering based on the Chebyshev distance, the situation where the source-load output data is sparse data and extreme data can be effectively handled, and the diversity and volatility of new energy data can be fully reflected.

[0052] In an embodiment of the present invention, for further illustration and limitation, the step of performing fuzzy clustering on the historical source-load output time series data according to the preset clustering distance calculation strategy to obtain a clustering result includes:

[0053] Initializing a type of clustering center and a type of membership degree function value for each piece of historical source-load output time series data;

[0054] For each piece of the historical source-load output time series data, calculate the Chebyshev distance between the historical source-load output time series data and each type of clustering center respectively, and update the type of membership degree function value according to the Chebyshev distance to obtain a type of membership degree matrix;

[0055] Update the type of clustering center according to the type of membership degree matrix of the historical source-load output time series data to obtain an updated type of clustering center;

[0056] Return to the step of calculating the Chebyshev distance between the historical source-load output time series data and each type of clustering center, and iteratively update the type of membership degree matrix and the type of clustering center until the iteration stop condition is satisfied, and obtain multiple clustering clusters representing extreme fluctuation output scenarios.

[0057] In the embodiments of the present invention, a type of clustering center, a type of membership function value, and a type of membership matrix are the clustering center, membership function value, and membership matrix in the fuzzy clustering process based on the Chebyshev distance. The process of iteratively updating the type of membership matrix and the type of clustering center is essentially a process of continuously optimizing the membership matrix and the clustering center under the constraint conditions based on the objective function. Set the historical source load output time series data as the data set X, expressed as: X = {x1, x2, …, x i , … x N}. In the fuzzy clustering process for the time scale, each independent x1 in the data set is a historical source load output time series data. Assume that it is divided into V1 categories by using the fuzzy C-means clustering algorithm, where V1 is a positive integer greater than 1, and the clustering centers of the V1 categories are {v1, v2, …, v j , … v V} respectively. The main goal of fuzzy C-means clustering is to find the distance from each historical source load output time series data to the clustering center, and minimize the sum of the weighted products of the membership degrees, while satisfying the necessary condition that the sum of the membership degrees is 1. Thus, the objective function of fuzzy clustering is:

[0058]

[0059] where u ij is the membership degree of the i-th historical source load output time series data x i with respect to the j-th type of clustering center v 1j , and u ij ∈ [0, 1]. m1 is the first fuzzy exponent, 0 < m < ∞, and generally takes the value of 2. i = 1, 2, 3, …, N; j = 1, 2, 3, …, V1. ||x i - v 1j ‖ ∞ = max(|x1 - v 1j |, |x2 - v1j, …, xi - v1j represents the Chebyshev distance between xi and v1j. The optimization constraint conditions of the above objective function are:

[0060] Fuzzy clustering is mainly used to process unclear and difficult-to-separate data sets in historical source load output data, such as wind and solar power generation historical data, fluctuating load data, etc. Fuzzy clustering is based on the distance between data to identify the fuzzy relationship between data points, so as to better understand the internal structure of the data, and classify the huge source load output data set according to the characteristic function between clustering clusters, the degree of probability closeness, and the similarity of data within the set.

[0061] After determining the objective function and constraints, the membership degree calculation formula of each historical source-load output time series data in a type of membership degree matrix with respect to the type of clustering center is derived based on Formula (1). Specifically, to minimize the objective function, the Lagrange multiplier method is used for the objective function (1) under the condition of satisfying the constraints, and the formula is obtained:

[0062]

[0063] After transformation, the formula is obtained:

[0064]

[0065] where λ i is the Lagrange multiplier. To adjust the update speed of the membership degree, the influence of the confidence c is considered in the optimization process, and the partial derivative is taken with respect to u ij in the above formula to obtain the type of membership degree calculation formula:

[0066]

[0067] where u ij is the membership degree of the i-th historical source-load output time series data x i with respect to the j-th type of clustering center v 1j , u ij ∈[0,1]. v 1k is any type of clustering center, c1 is the first confidence, m1 is the first fuzzy index, V1 is the total number of type of clustering centers, and N is the total number of historical source-load output time series data. Among them, the size of the confidence can adjust the optimization speed of the membership degree. The larger the membership degree, the faster the convergence speed of the algorithm, but the clustering accuracy will decrease; the smaller the confidence, the slower the convergence speed of the algorithm, but the clustering result will be more accurate. Therefore, the value of the first confidence can be configured according to the requirements of the specific application scenario, and the embodiments of the present invention do not make specific limitations.

[0068] Further, the partial derivative is taken with respect to v 1j to obtain the calculation formula of the type of clustering center as:

[0069]

[0070] In an embodiment of the present invention, for further illustration and limitation, the step of performing fuzzy clustering on the historical source-load output time series data according to the preset clustering distance calculation strategy to obtain a clustering result includes:

[0071] Initializing the type-two clustering center and the type-two membership degree function values of each historical source-load output time series data;

[0072] For each piece of the historical source-load output time series data, calculate the Manhattan distance between the historical source-load output time series data and each second-class clustering center, and update the second-class membership function values based on the Manhattan distance to obtain a second-class membership matrix;

[0073] Update the second-class clustering center according to the second-class membership matrix of the historical source-load output time series data to obtain an updated second-class clustering center;

[0074] Return to the step of calculating the Manhattan distance between the historical source-load output time series data and each second-class clustering center, and iteratively update the second-class membership matrix and the second-class clustering center until the iteration stop condition is satisfied, so as to obtain multiple clustering clusters representing typical output scenarios.

[0075] In the embodiments of the present invention, the second-class clustering center, the second-class membership function value, and the second-class membership matrix are the clustering center, the membership function value, and the membership matrix in the fuzzy clustering process based on the Manhattan distance. The process of iteratively updating the second-class membership matrix and the second-class clustering center is the same as that of iteratively updating the first-class membership matrix and the first-class clustering center, which will not be elaborated herein. The preset clustering distance calculation strategy also includes the Manhattan distance calculation strategy. Fuzzy clustering based on the Manhattan distance can measure the sum of the absolute values of the differences between each corresponding element of two pieces of data. In the processing of sparse data and data sets with outliers, compared with the fuzzy clustering algorithm based on the Euclidean distance, it can identify clearer clustering boundaries, and for the source-load output data, it can more reasonably distinguish different characteristics at different time points, so as to obtain multiple clustering clusters for representing different typical output scenarios.

[0076] In the process of iteratively updating the second-class membership matrix and the second-class clustering center, the formula for calculating the second-class membership is:

[0077]

[0078] where, u ij ′ is the membership of the i-th piece of historical source-load output time series data x i with respect to the j-th second-class clustering center v 2j , u ij ′ ∈[0,1]. v 2k is any second-class clustering center, c2 is the second confidence level, m2 is the second fuzzy index, and V2 is the total number of second-class clustering centers. Correspondingly, the calculation formula for the second-class clustering center is:

[0079]

[0080] It should be noted that the derivation processes of formula (7) and formula (5), and formula (8) and formula (6) are the same, only the Chebyshev distance ||x i - v 1j || ∞ is replaced by the Manhattan distance ||x i - v 2j || 1 , The derivation processes of formula (7) and (8) will not be elaborated here. The first confidence level and the second confidence level can be configured to be the same confidence level or different; the first fuzzy exponent and the second fuzzy exponent can be configured to be the same fuzzy exponent or different. For the configuration of the confidence level and the fuzzy exponent, the embodiments of the present invention do not make specific limitations.

[0081] 103. For each clustering cluster in the clustering result, a power deviation probability distribution function is generated respectively to obtain the source-load uncertainty data representation information of the target new energy satisfying different scenarios.

[0082] In the embodiments of the present invention, considering that in the actual working scenario, the prediction data of wind-solar load is increasingly affected by uncertainty, and the deviation of the received power prediction data is also increasing. Therefore, it is necessary to generate the probability distribution of the power generation deviation by using historical power data, so as to adjust the scheduling strategy for the real-time power deviation in the day-ahead to real-time stage. In order to more precisely represent the source-load uncertainty, for each clustering cluster corresponding to each scenario in the clustering result, according to the occurrence probability of the clustering cluster and characteristic parameters such as upper and lower bounds, a power deviation probability distribution function is generated respectively, and this function is used as the source-load uncertainty data representation information of the corresponding scenario, so as to obtain the source-load uncertainty data representation information of the target new energy satisfying different scenarios. By digitally describing the source-load uncertainty of different scenarios, the fluctuation characteristics of renewable energy can be more accurately reflected, and a more flexible and accurate data basis can be provided for subsequent regulation, analysis and optimization based on the source-load uncertainty.

[0083] In an embodiment of the present invention, for further illustration and limitation, the generating a power deviation probability distribution function for each clustering cluster in the clustering result includes:

[0084] For each clustering cluster, calculate the power difference between each historical source-load output time series data and the corresponding clustering center in each time period to obtain a high-dimensional power deviation set;

[0085] Expand the high-dimensional power deviation set into a one-dimensional array, and calculate the probability density function of each element in the one-dimensional array respectively according to the kernel function smoothing algorithm;

[0086] Construct the power deviation probability distribution function of the clustering cluster according to the probability density functions of the respective elements, and obtain the power deviation probability distribution functions of the respective clustering clusters.

[0087] In the embodiment of the present invention, the power deviation probability distribution function of each clustering cluster is determined based on the operation method of the probability distribution function of kernel density estimation. In the kernel density estimation algorithm, the kernel function smoothing algorithm is used to estimate the probability density function, and a continuous probability density function is constructed by the linear summation of discrete sample points, so as to obtain a smooth sample distribution. Since the elements in the high-dimensional power deviation set are the power differences between each historical source-load output time series data in each clustering cluster and the corresponding clustering center in each time period, and each historical source-load output time series data includes the power of multiple time periods, therefore, this high-dimensional power deviation set is a high-dimensional element set. In order to facilitate kernel density estimation, it is necessary to reduce its dimension, expand it into a one-dimensional array, and then calculate the power deviation probability distribution function based on the formula of kernel density estimation. Among them, the formula of kernel density estimation is:

[0088]

[0089] Among them, is the value of the estimated probability density function at point y, M is the total number of samples, h is the bandwidth, K is the kernel function, and y i is the sample data point.

[0090] Furthermore, considering that wind-solar power generation and load are easily affected by natural conditions or human conditions, and the power deviation is divided into positive deviation and negative deviation. When the data volume is sufficient, the overall is evenly distributed above and below 0. Therefore, the Gaussian function that most conforms to the data characteristics is used as the kernel function of kernel density estimation, and the hyperparameter of bandwidth is considered in the Gaussian kernel function. Among them, the Gaussian function formula is:

[0091] Considering the bandwidth hyperparameter in the Gaussian kernel function, the formula is obtained:

[0092]

[0093] Substitute formula (10) into formula (11), and then substitute the substitution result into formula (9), the formula is obtained:

[0094]

[0095] Take the elements in the one-dimensional array as the sample data points y i in the formula. Based on formula (12), the power deviation probability distribution functions of the respective clustering clusters can be calculated.

[0096] In an embodiment of the present invention, for further illustration and limitation, after obtaining the source-load uncertainty digital representation information of the target new energy to meet the requirements of different scenarios, the method further includes:

[0097] Storing the source-load uncertainty digital representation information in a target storage space;

[0098] In response to a retrieval instruction for source-load uncertainty information, determining at least one new energy to be optimized according to the power system optimization task carried by the retrieval instruction;

[0099] Sending the source-load uncertainty digital representation information matching the new energy to be optimized to a target terminal, so that the target terminal executes the power system optimization task according to the source-load uncertainty digital representation information.

[0100] In an embodiment of the present invention, after obtaining the source-load uncertainty digital representation information of the target new energy, the information is stored in a pre-constructed target storage space for subsequent task invocation. Among them, the target storage space can be local storage, blockchain or cloud storage, and the embodiments of the present invention do not make specific limitations. The target storage space stores the source-load uncertainty digital representation information of different types of new energy in different power jurisdiction areas. That is, after representing the source-load uncertainty of different types of new energy in different power jurisdiction areas, they are all stored in the target storage space for unified information management. When a power system optimization task needs to use the source-load uncertainty digital representation information of a certain new energy and retrieves the information in the target storage space from the current execution entity, at least one new energy to be optimized is determined according to the power system optimization task carried by the retrieval instruction, so as to send the source-load uncertainty digital representation information matching each new energy to be optimized to the target terminal specified by the retrieval instruction. For example, if the power system optimization task is the energy system optimization task of distributed photovoltaic power generation and wind power generation in Area A, the new energy to be optimized is the distributed photovoltaic power generation and wind power generation corresponding to Area A.

[0101] It should be noted that the power system optimization tasks include power load forecasting tasks, new energy access management tasks, power system risk assessment tasks, distributed energy system optimization tasks, intelligent microgrid management tasks, and energy trading and pricing tasks. In the power load forecasting task, based on the digital characterization information of source-load uncertainty data, the fluctuation of power load can be accurately predicted, providing strong support for the stable operation of the power system. In the new energy access management task, through the digital characterization information of source-load uncertainty data, the access of new energy can be managed more effectively, and the operation mode of the power system can be optimized. In the power system risk assessment task, through the digital characterization information of source-load uncertainty data, the risk level of the power system can be evaluated more comprehensively, providing a decision-making basis for the safe operation of the power system. In the distributed energy system optimization task, through the digital characterization information of source-load uncertainty data, the coordinated operation of multiple energies can be optimized more accurately, improving energy utilization efficiency. Moreover, a more scientific and reasonable energy management strategy can be formulated to achieve the efficient utilization and sustainable development of energy. In the intelligent microgrid management task, since the uncertainty of source-load has an important impact on the operation stability of the microgrid, through the digital characterization information of source-load uncertainty data, the operation mode of the microgrid can be optimized, and the stability and reliability of the microgrid can be improved. In the energy trading and pricing task, since the uncertainty of source-load has an important impact on energy trading and pricing, through the digital characterization information of source-load uncertainty data, the value and risk of energy can be evaluated more accurately, providing strong support for energy trading and pricing.

[0102] The present invention provides a method for generating digital characterization of source-load uncertainty. In an embodiment of the present invention, multiple historical source-load output time series data of a target new energy are obtained, where the time scales of each piece of historical source-load output time series data are the same, and each piece of historical source-load output time series data includes multi-period output data; fuzzy clustering is performed on the historical source-load output time series data according to a preset clustering distance calculation strategy to obtain a clustering result, where the preset clustering distance calculation strategy includes a Chebyshev distance calculation strategy, and the clustering result includes multiple clustering clusters representing extreme fluctuation output scenarios; for each clustering cluster in the clustering result, a power deviation probability distribution function is generated respectively to obtain the digital characterization information of source-load uncertainty of the target new energy under different scenarios. Through fuzzy clustering based on the Chebyshev distance, the situation where source-load output data is sparse data and extreme data can be effectively processed, and the diversity and volatility of new energy data can be fully reflected.

[0103] Furthermore, as an implementation of the above Figure 1 shown method, an embodiment of the present invention provides a device for generating digital characterization of source-load uncertainty, as Figure 2 shown, the device includes:

[0104] An acquisition module 21, configured to acquire multiple pieces of historical source-load output time-series data of a target new energy source, where the time scales of each piece of the historical source-load output time-series data are the same, and each piece of the historical source-load output time-series data includes multi-period output data;

[0105] A clustering module 22, configured to perform fuzzy clustering on the historical source-load output time-series data according to a preset clustering distance calculation strategy to obtain a clustering result, where the preset clustering distance calculation strategy includes a Chebyshev distance calculation strategy, and the clustering result includes multiple clustering clusters representing extreme fluctuation output scenarios;

[0106] A generation module 23, configured to generate a power deviation probability distribution function for each clustering cluster in the clustering result respectively, to obtain digital representation information of the source-load uncertainty of the target new energy source under different scenarios.

[0107] Further, the clustering module 22 includes:

[0108] A first initialization unit, configured to initialize a type of clustering center and a type of membership function value of each piece of historical source-load output time-series data;

[0109] A first calculation unit, configured to calculate the Chebyshev distance between each piece of the historical source-load output time-series data and each type of clustering center respectively, and update the type of membership function value according to the Chebyshev distance to obtain a type of membership matrix;

[0110] An update unit, configured to update the type of clustering center according to the type of membership matrix of the historical source-load output time-series data to obtain an updated type of clustering center;

[0111] An iteration unit, configured to return to the step of calculating the Chebyshev distance between the historical source-load output time-series data and each type of clustering center, and perform iterative updates on the type of membership matrix and the type of clustering center until an iteration stop condition is met, to obtain multiple clustering clusters representing extreme fluctuation output scenarios.

[0112] In a specific application scenario, in the type of membership matrix, the membership calculation formula of each piece of historical source-load output time-series data with respect to the type of clustering center is:

[0113] where, u ij is the membership of the i-th piece of historical source-load output time-series data x i with respect to the j-th type of clustering center v 1j , u ij ∈[0,1]; v 1kis any type of clustering center, c1 is the first confidence level, m1 is the fuzzy exponent, V1 is the total number of clustering centers of one type, and N is the total number of historical source load output time series data.

[0114] Further, the clustering module 22 further includes:

[0115] A second initialization unit for initializing the clustering centers of the second type and the membership function values of the second type for each historical source load output time series data;

[0116] A second calculation unit for calculating the Manhattan distance between each historical source load output time series data and each clustering center of the second type respectively, and updating the membership function values of the second type based on the Manhattan distance to obtain a membership matrix of the second type;

[0117] A second update unit for updating the clustering centers of the second type according to the membership matrix of the second type of the historical source load output time series data to obtain the updated clustering centers of the second type;

[0118] A second iteration unit for returning to the step of calculating the Manhattan distance between each historical source load output time series data and each clustering center of the second type, and iteratively updating the membership matrix of the second type and the clustering centers of the second type until the iteration stop condition is met, to obtain multiple clustering clusters representing typical output scenarios.

[0119] In a specific application scenario, in the membership matrix of the second type, the membership calculation formula of each historical source load output time series data with respect to the clustering centers of the second type is:

[0120] where u ij ′ is the membership of the i-th historical source load output time series data x i with respect to the j-th clustering center v 2j of the second type, v 2k is any clustering center of the second type, c2 is the second confidence level, m2 is the second fuzzy exponent, and V2 is the total number of clustering centers of the second type.

[0121] Further, the generation module 23 includes:

[0122] A third calculation unit for calculating the power difference between each historical source load output time series data and the corresponding clustering center in each time period for each clustering cluster to obtain a high-dimensional power deviation set;

[0123] A fourth calculation unit for expanding the high-dimensional power deviation set into a one-dimensional array and calculating the probability density function of each element in the one-dimensional array respectively according to the kernel function smoothing algorithm;

[0124] A construction unit is used to construct a power deviation probability distribution function of the clustering cluster according to the probability density functions of the respective elements, so as to obtain the power deviation probability distribution functions of the respective clustering clusters.

[0125] Furthermore, the device further includes:

[0126] A storage module is used to store the source-load uncertainty data characterization information into a target storage space, and the target storage space stores the source-load uncertainty data characterization information of different types of new energy in different power jurisdiction areas;

[0127] A determination module is used to, in response to a retrieval instruction for source-load uncertainty information, determine at least one new energy to be optimized according to the power system optimization task carried by the retrieval instruction, where the power system optimization task includes a power load prediction task, a new energy access management task, a power system risk assessment task, a distributed energy system optimization task, a smart microgrid management task, and an energy trading and pricing task;

[0128] A sending module is used to send the source-load uncertainty data characterization information matching the new energy to be optimized to a target terminal, so that the target terminal executes the power system optimization task according to the source-load uncertainty data characterization information.

[0129] The present invention provides a generating device for source-load uncertainty data characterization. In an embodiment of the present invention, multiple historical source-load output time series data of a target new energy are obtained, where each piece of historical source-load output time series data has the same time scale, and each piece of historical source-load output time series data includes multi-period output data; fuzzy clustering is performed on the historical source-load output time series data according to a preset clustering distance calculation strategy to obtain a clustering result, where the preset clustering distance calculation strategy includes a Chebyshev distance calculation strategy, and the clustering result includes multiple clustering clusters representing extreme fluctuation output scenarios; for each clustering cluster in the clustering result, a power deviation probability distribution function is respectively generated to obtain the source-load uncertainty data characterization information of the target new energy satisfying different scenarios. Through fuzzy clustering based on the Chebyshev distance, the situation where the source-load output data is sparse data and extreme data can be effectively processed, and the diversity and volatility of new energy data can be fully reflected. According to an embodiment of the present invention, a storage medium is provided, and the storage medium stores at least one executable instruction, and the computer executable instruction can execute the method for generating source-load uncertainty data characterization in any of the above method embodiments.

[0130] Figure 3 The structural schematic diagram of a terminal provided according to an embodiment of the present invention is shown. The specific implementation of the terminal is not limited in the specific embodiment of the present invention.

[0131] AsFigure 3 As shown in Figure 3 , the terminal may include: a processor 302, a communications interface 304, a memory 306, and a communication bus 308.

[0132] Among them: The processor 302, the communications interface 304, and the memory 306 communicate with each other through the communication bus 308.

[0133] The communications interface 304 is used for network communication with other devices such as clients or other servers.

[0134] The processor 302 is used to execute the program 310, and specifically can execute the relevant steps in the embodiment of the method for generating the source-load uncertainty data representation described above.

[0135] Specifically, the program 310 may include program code, and the program code includes computer operation instructions.

[0136] The processor 302 may be a central processing unit CPU, or a specific integrated circuit ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present invention. One or more processors included in the terminal may be of the same type of processor, such as one or more CPUs; or may be of different types of processors, such as one or more CPUs and one or more ASICs.

[0137] The memory 306 is used to store the program 310. The memory 306 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk memory.

[0138] The program 310 is specifically used to cause the processor 302 to perform the following operations:

[0139] Obtain multiple pieces of historical source-load output time series data of the target new energy, where the time scales of each piece of the historical source-load output time series data are the same, and each piece of the historical source-load output time series data includes multi-period output data;

[0140] Perform fuzzy clustering on the historical source-load output time series data according to a preset clustering distance calculation strategy to obtain a clustering result, where the preset clustering distance calculation strategy includes a Chebyshev distance calculation strategy, and the clustering result includes multiple clustering clusters representing extreme fluctuation output scenarios;

[0141] For each clustering cluster in the clustering result, a power deviation probability distribution function is generated respectively, so as to obtain the source-load uncertainty digital characterization information of the target new energy satisfying different scenarios.

[0142] Obviously, those skilled in the art should understand that the above-mentioned modules or steps of the present invention can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. Optionally, they can be implemented by program codes executable by the computing device. Thus, they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a sequence different from that here, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module to implement. In this way, the present invention is not limited to any specific combination of hardware and software.

[0143] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for generating a digital representation of source-load uncertainty data, characterized in that, Including: Obtain multiple historical source-load output time series data of the target new energy. Among them, the time scales of each piece of the historical source-load output time series data are the same, and each piece of the historical source-load output time series data includes multi-period output data; Perform fuzzy clustering on the historical source-load output time series data according to a preset clustering distance calculation strategy to obtain a clustering result. Among them, the preset clustering distance calculation strategy includes a Chebyshev distance calculation strategy, and the clustering result includes multiple clustering clusters representing extreme fluctuation output scenarios; For each clustering cluster in the clustering result, generate a power deviation probability distribution function respectively to obtain the digital representation information of the source-load uncertainty of the target new energy under different scenarios.

2. The method according to claim 1, wherein The step of performing fuzzy clustering on the historical source-load output time series data according to a preset clustering distance calculation strategy to obtain a clustering result includes: Initialize a type of clustering center and the membership function values of each piece of historical source-load output time series data; For each piece of the historical source-load output time series data, calculate the Chebyshev distance between the historical source-load output time series data and each type of clustering center respectively, and update the membership function values of the type according to the Chebyshev distance to obtain a type membership matrix; Update the type of clustering center according to the type membership matrix of the historical source-load output time series data to obtain an updated type of clustering center; Return to the step of calculating the Chebyshev distance between the historical source-load output time series data and each type of clustering center, and perform iterative updates on the type membership matrix and the type of clustering center until the iterative stop condition is met to obtain multiple clustering clusters representing extreme fluctuation output scenarios.

3. The method according to claim 1, characterized in that, In the type membership matrix, the membership degree calculation formula of each piece of historical source-load output time series data with respect to the type of clustering center is: Among them, u ij is the time series data x of the i-th historical source-load output i with respect to the j-th type-1 clustering center v 1j of membership degree, u ij ∈ [0, 1]; v 1k is any type-1 clustering center, c1 is the first confidence level, m1 is the fuzzy index, V1 is the total number of type-1 clustering centers, and N is the total number of historical source-load output time series data.

4. The method according to claim 1, wherein The preset clustering distance calculation strategy further includes a Manhattan distance calculation strategy, and the clustering result further includes multiple clustering clusters representing different typical output scenarios. The step of performing fuzzy clustering on the historical source-load output time series data according to a preset clustering distance calculation strategy to obtain a clustering result includes: Initialize a second type of clustering center and the membership function values of each piece of historical source-load output time series data; For each piece of the historical source-load output time series data, calculate the Manhattan distance between the historical source-load output time series data and each second type of clustering center respectively, and update the membership function values of the second type according to the Manhattan distance to obtain a second type membership matrix; Update the second type of clustering center according to the second type membership matrix of the historical source-load output time series data to obtain an updated second type of clustering center; Return to the step of calculating the Manhattan distance between the historical source-load output time series data and each second type of clustering center, and perform iterative updates on the second type membership matrix and the second type of clustering center until the iterative stop condition is met to obtain multiple clustering clusters representing typical output scenarios.

5. The method according to claim 4, characterized in that In the second type membership matrix, the membership degree calculation formula of each piece of historical source-load output time series data with respect to the second type of clustering center is: Among them, u ij ′ is the time series data x of the historical source-load output of the i-th item i with respect to the membership degree of the j-th second-class clustering center v 2j , v 2k is any second-class clustering center, c2 is the second confidence level, m2 is the second fuzzy index, and V2 is the total number of second-class clustering centers.

6. The method according to claim 5, characterized in that, For each clustering cluster in the clustering result, generating a power deviation probability distribution function respectively, including: For each of the clustering clusters, calculating the power difference between each piece of historical source-load output time series data and the corresponding clustering center in each time period, to obtain a high-dimensional power deviation set; Expanding the high-dimensional power deviation set into a one-dimensional array, and respectively calculating the probability density function of each element in the one-dimensional array according to the kernel function smoothing algorithm; Constructing the power deviation probability distribution function of the clustering cluster according to the probability density function of each element, to obtain the power deviation probability distribution functions of each clustering cluster.

7. The method according to claim 1, wherein After obtaining the digital representation information of the source-load uncertainty for the target new energy to meet different scenarios, the method further includes: Storing the digital representation information of the source-load uncertainty in a target storage space, where the target storage space stores the digital representation information of the source-load uncertainty of different types of new energy in different power jurisdiction areas; In response to a retrieval instruction of source-load uncertainty information, determining at least one new energy to be optimized according to the power system optimization task carried by the retrieval instruction, where the power system optimization task includes power load forecasting task, new energy access management task, power system risk assessment task, distributed energy system optimization task, intelligent microgrid management task, energy trading and pricing task; Sending the digital representation information of the source-load uncertainty matching the new energy to be optimized to a target terminal, so that the target terminal executes the power system optimization task according to the digital representation information of the source-load uncertainty.

8. A generating device for digitizing and representing source-load uncertainty data, characterized in that, Including: An acquisition module, configured to acquire multiple pieces of historical source-load output time series data of a target new energy, where the time scales of each piece of the historical source-load output time series data are the same, and each piece of the historical source-load output time series data includes multi-period output data; A clustering module, configured to perform fuzzy clustering on the historical source-load output time series data according to a preset clustering distance calculation strategy, to obtain a clustering result, where the preset clustering distance calculation strategy includes a Chebyshev distance calculation strategy, and the clustering result includes multiple clustering clusters representing extreme fluctuation output scenarios; A generation module, configured to respectively generate a power deviation probability distribution function for each clustering cluster in the clustering result, to obtain the digital representation information of the source-load uncertainty for the target new energy to meet different scenarios.

9. A storage medium, in which at least one executable instruction is stored, and the executable instruction causes a processor to perform operations corresponding to the method for generating the digital representation of source-load uncertainty according to any one of claims 1-7.

10. A terminal, comprising: A processor, a memory, a communication interface, and a communication bus, where the processor, the memory, and the communication interface complete communication with each other through the communication bus; The memory is used to store at least one executable instruction, and the executable instruction causes the processor to perform operations corresponding to the method for generating the digital representation of source-load uncertainty according to any one of claims 1-7.