Spot day-ahead market decision sample matrix clustering preprocessing method and system

By employing model-driven projection clustering and nuclear norm regularization, the problems of dimensionality curse and outlier interference in the day-ahead electricity market decision-making sample matrix in high-dimensional space are solved. This approach enables stable identification of typical operating modes and effective identification of outliers, thereby improving the efficiency and accuracy of electricity market decision-making.

CN121365263APending Publication Date: 2026-01-20ELECTRIC POWER RESEARCH INSTITUTE OF STATE GRID SHANDONG ELECTRIC POWER COMPANY
View PDF 13 Cites 0 Cited by

Patent Information

Application Number
CN202511946721.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-23
Publication Date
2026-01-20

AI Technical Summary

Technical Problem

The decision sample matrix of the day-ahead electricity market suffers from the curse of dimensionality and outlier interference in high-dimensional space, which degrades the performance of traditional clustering algorithms and makes it difficult to effectively uncover typical operating patterns and intrinsic subspace structures.

Method used

A model-driven projection clustering and nuclear norm regularization method is adopted. Through normalization processing, fuzzy membership matrix and intra-cluster weight matrix, combined with an alternating iterative optimization strategy, abnormal samples are identified and clustering results are cleaned to construct a stable typical operation mode.

Benefits of technology

It significantly alleviates the curse of dimensionality in high-dimensional spaces, improves the compactness and interpretability of clustering results, and enhances the robustness and efficiency of prediction and optimization models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121365263A_ABST
    Figure CN121365263A_ABST
Patent Text Reader

Abstract

The invention provides a spot day-ahead market decision sample matrix clustering preprocessing method and system, and belongs to the technical field of electricity market and data preprocessing, and the method comprises the steps: determining the time granularity and statistical period of a sample, and constructing a sample matrix; introducing a dimension weight matrix for each cluster, defining a cluster center matrix, updating a cluster center under a typical fuzzy clustering framework, and adaptively updating a dimension weight according to a weighted variance in the cluster; explicitly introducing a nuclear norm regular term into the clustering objective function; solving the minimum value of the clustering objective function by adopting an alternating iterative optimization strategy; and outputting the cluster to which each sample belongs, the membership of the cluster, the typical operation mode corresponding to each cluster and the abnormal sample set. According to the method, abnormal working condition samples such as extreme climate, sudden maintenance and data errors can be effectively identified, the robustness of a subsequent prediction and optimization model is improved, and the efficiency and precision of links such as day-ahead market quotation strategy optimization, flexibility evaluation and safety check are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of power market and data preprocessing, and particularly relates to a spot day-ahead market decision sample matrix clustering preprocessing method and system. BACKGROUND

[0002] The statements in this section merely provide background information related to the present application and do not necessarily constitute prior art.

[0003] The goal of the spot day-ahead power market in the clearing is to determine the generation plan, market price and node price in each 15-minute or 1-hour trading period of the next day under the premise of meeting the safety constraints of the power system, with the goal of minimizing the total social electricity purchase cost or maximizing the social welfare. While the bidding optimization is a decentralized decision made by market participants to maximize individual interests. The goal of the clearing of the spot day-ahead power market is to determine the generation plan, market price and node price in each 15-minute or 1-hour trading period of the next day on the basis of meeting the safety constraints of the power system, with the goal of minimizing the total social electricity purchase cost or maximizing the social welfare. The bidding optimization is a decentralized decision made by market participants based on individual interests.

[0004] Therefore, when the spot day-ahead power market is clearing and bidding optimization, it needs to consider multiple input factors and output decision results comprehensively, including but not limited to: multi-element load and its time series characteristics; uncertain power output of wind power, photovoltaic and other power sources and their random characteristics; probability correlation between various uncertain factors; independent energy storage operation strategy and available capacity; aggregated load and dispatchable resources of virtual power plant (VPP); environmental conditions such as weather and climate; medium and long-term constraint factors such as unit maintenance plan, power grid construction plan, and power source retirement plan. Correspondingly, the output of the spot day-ahead market decision includes: generation plan of thermal power units in each period; wind power and photovoltaic output plan or reduction ratio; independent energy storage charging and discharging power and state trajectory; aggregated output and decomposition instruction of virtual power plant, etc.

[0005] In order to train high-quality prediction and decision models using historical data, it is necessary to construct a set of spot day-ahead market decision sample matrices and establish a mapping relationship between the historical data of the above input and output variables. However, such sample matrices usually have extremely high dimensions, contain multi-source features with different dimensions, different time scales, strong and weak correlations, and the dimensionality disaster and space phenomenon exist in high-dimensional space. The traditional clustering algorithm based on Euclidean distance such as K-means has a sharp performance degradation in high-dimensional space. There are abnormal points such as abnormal working conditions, extreme weather, and sudden maintenance in the samples, which will seriously interfere with model training. The spot day-ahead market decision sample matrix has a natural matrix structure and time series correlation. If the kernel norm and other low-rank constraints can be used in the clustering process, it will be beneficial to mine the typical operating modes and internal subspace structure. SUMMARY

[0006] In order to overcome the above-mentioned deficiencies of the prior art, the present application provides a spot day-ahead market decision sample matrix clustering preprocessing method and system, which realizes spot day-ahead market decision sample matrix clustering preprocessing based on model-driven projection clustering and kernel norm regularization, reduces the workload of manual screening and experience segmentation of modeling personnel, and improves the efficiency and accuracy of business links such as day-ahead market quotation strategy optimization, flexibility evaluation and safety checking.

[0007] In order to achieve the above-mentioned purpose, one or more embodiments of the present application provide the following technical solutions: In a first aspect, a spot day-ahead market decision sample matrix clustering preprocessing method is disclosed, comprising: determining the time granularity and statistical period of the sample and constructing a sample matrix; normalizing each dimension feature in the sample matrix to obtain normalized features, and defining statistical indicators of each dimension based on the normalized features for feature selection and dimension compression; setting the number of clusters for clustering, defining a fuzzy membership matrix, and obtaining membership constraints based on the fuzzy membership matrix; introducing a dimension weight matrix for each cluster, defining a cluster center matrix, updating the cluster center in a typical fuzzy clustering framework, and updating the dimension weight according to the weighted variance within the cluster; On the basis of model-driven clustering, a probability model is introduced to describe the distribution of samples within the cluster; explicitly introducing a kernel norm regularization term in the clustering objective function; adopting an alternating iterative optimization strategy to solve the minimum value of the clustering objective function; Based on the minimum value of the solved clustering objective function, further abnormal sample identification and clustering result cleaning are performed, and the cluster to which each sample belongs and its membership, the typical operating mode corresponding to each cluster, and the abnormal sample set are output.

[0008] As a further technical solution, when constructing the sample matrix, it specifically comprises: constructing samples according to the time step of day-ahead market rolling optimization; determining the historical time range of the sample to ensure both the number of samples and the coverage of various climates and operating conditions; extracting input class features and output class features for each sample; unifying and splicing all input class features and output class features in a fixed order to obtain a sample feature vector; stacking all samples by rows to obtain a spot day-ahead market decision sample matrix.

[0009] As a further technical solution, the extracted input class features include: time-period load prediction value; available output prediction of wind power and photovoltaic power and related uncertainty quantification index; probability correlation estimation result between load and wind-solar output; available capacity, SOC state of charge and operation constraint parameter of independent energy storage; available regulation capacity and controllable load scale of aggregated resource in virtual power plant; meteorological and climate information; unit maintenance plan; grid power supply construction / retirement state.

[0010] As a further technical solution, the extracted output class features include: output curve or time-of-use output decision of each thermal power unit in day-ahead clearing result; output plan or curtailment / abandoned wind and light amount of each wind farm and photovoltaic power station; charge-discharge power trajectory and initial and final SOC of independent energy storage; aggregated output of virtual power plant and internal resource decomposition instruction.

[0011] As a further technical solution, the statistical indicators of each dimension include mean, variance and correlation coefficient, and the statistical indicators are used for: eliminating features with minimum variance; merging or screening highly correlated features.

[0012] As a further technical solution, based on the minimum value of the solved clustering objective function, further abnormal sample identification and clustering result cleaning are performed, specifically including: defining the minimum cluster distance of each sample, and calculating the distance of the sample to the nearest cluster according to the final weighted distance; setting distance threshold and membership threshold, if a sample meets the set condition, the sample is marked as an abnormal sample, and for the sample marked as abnormal, it can be clustered separately or removed from the training set according to its distribution characteristics.

[0013] In a second aspect, a spot day-ahead market decision sample matrix clustering preprocessing system is disclosed, including: a sample matrix construction module configured to determine the time granularity and statistical period of the sample and construct the sample matrix; a normalization processing module configured to perform normalization processing on each dimension feature in the sample matrix to obtain normalized features, and define statistical indicators of each dimension based on the normalized features, which are used for feature selection and dimension compression; The updating module is configured to set the number of clusters of the clustering, define a fuzzy membership matrix, obtain membership constraints based on the fuzzy membership matrix, introduce a dimension weight matrix for each cluster, define a cluster center matrix, update the cluster center in a typical fuzzy clustering framework, and update the dimension weight adaptively according to the weighted variance within the cluster. The clustering objective function construction module is configured to introduce a probability model to describe the distribution of samples within a cluster on the basis of model-driven clustering, and explicitly introduce a kernel norm regularization term in the clustering objective function. The preprocessing output module is configured to solve the minimum value of the clustering objective function by using an alternating iterative optimization strategy. Based on the minimum value of the solved clustering objective function, further abnormal sample identification and clustering result cleaning are performed, and the cluster to which each sample belongs and the membership of each sample, the typical operation mode corresponding to each cluster and the abnormal sample set are output.

[0014] The above one or more technical solutions have the following beneficial effects: The technical scheme of the present application can automatically highlight the key feature dimensions of various operation scenarios in a high-dimensional space, and can extract stable typical operation modes by using the low-rank structure of the sample matrix. Thus, the dimension disaster and space problem faced by traditional clustering algorithms in a high-dimensional scenario are significantly alleviated, the compactness and interpretability of the clustering result are improved, and a unified and reliable data basis is provided for constructing typical days / typical scenarios, decision sample compression and sample weighting.

[0015] The technical scheme of the present application can effectively identify and mark abnormal working condition samples such as extreme climate, sudden maintenance and data errors by explicitly describing the weighted distance, membership and kernel norm of the sub-matrix within the cluster of the sample and the cluster during the clustering process, and combining the distance-membership joint criterion of abnormal samples, thereby significantly improving the robustness of subsequent prediction and optimization models to abnormal data.

[0016] Meanwhile, the method and system of the present application can be directly embedded in an existing scheduling and trading platform to realize automatic clustering preprocessing of the spot day-ahead market decision sample matrix, reduce the workload of manual screening and experience segmentation of modeling personnel, and improve the efficiency and accuracy of business links such as day-ahead market quotation strategy optimization, flexibility evaluation and safety checking.

[0017] The advantages of the additional aspects of the present application will be partially given in the following description, partially become obvious from the following description, or be understood through the practice of the present application. BRIEF DESCRIPTION OF DRAWINGS

[0018] The accompanying drawings, which form a part of this specification, are included to provide a further understanding of the application and are incorporated in and constitute a part of this specification. The embodiments of the application, and their

[0019] Figure 1 The method flowchart of the embodiment of the application is shown in the figure. DETAILED DESCRIPTION

[0020] It should be noted that the following detailed description is exemplary in nature and is intended to provide further description of the application. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs.

[0021] It should be noted that the terms used herein are merely for the purpose of describing specific embodiments and are not intended to limit the exemplary embodiments according to the application.

[0022] In the case of no conflict, the embodiments in the application and the features in the embodiments can be combined with each other.

[0023] Embodiment one Referring to the accompanying Figure 1 The embodiment discloses a spot day-ahead market decision sample matrix clustering preprocessing method, which comprises the following steps: Step one: sample matrix construction: determine the sample time granularity and statistical period, select the historical time range, extract the input class features, extract the output class features, splice them into the sample feature vector in a fixed order, and stack them into the sample matrix by rows; In this step, by constructing a set of spot day-ahead market decision sample matrices, the input quantities such as multivariate load, uncertain power output and random characteristics, uncertain factors, probability correlation between independent energy storage, virtual power plant, climate, unit maintenance plan, network power supply construction, power supply retirement, and the output quantities such as thermal power, wind power, photovoltaic power, independent energy storage, virtual power plant are uniformly represented as sample matrices; Step two: data normalization and feature selection: normalize each feature dimension of the sample matrix, calculate the mean and variance of each feature dimension, calculate the correlation coefficient between features, eliminate features with extremely small variance, merge or select highly correlated features, and obtain the normalized and compressed sample matrix; Step three: clustering model initialization: set the number of clustering clusters , initialize the fuzzy membership matrix, cluster center matrix and dimension weight matrix, set the convergence threshold and maximum iteration number, and apply the membership constraint; Step four: building a probability model and objective function: assuming that each feature dimension within a cluster is approximately independent, a multi-dimensional Gaussian probability model is established for each cluster, a mixed probability model of samples is constructed, a clustering loss term containing fuzzy membership and dimension weight is constructed, a kernel norm regularization term of intra-cluster sub-matrix is added to the objective function, and an overall clustering objective function is obtained; Step five: alternating iterative optimization: calculate the weighted distance of samples to each cluster, update the fuzzy membership matrix according to the weighted distance, update the cluster center matrix, update the dimension weight matrix according to the weighted variance within the cluster, update the kernel norm related quantity of each cluster sub-matrix based on the current cluster division, and calculate the change of the objective function or the change of the membership; If yes, go to step six, otherwise continue alternating iterative optimization; Step six: abnormal sample identification: according to the final membership matrix, cluster center matrix and dimension weight matrix, the weighted distance of each sample to the nearest cluster is calculated, the maximum membership of each sample is calculated, the distance threshold and membership threshold are set, the samples that meet the conditions of too large distance or too small membership are marked as abnormal samples, the abnormal sample set is obtained, and the clustering preprocessing of the spot day-ahead market decision sample matrix set is completed; Step seven: clustering result output and application: output the cluster and membership of each sample, output the typical operation mode corresponding to each cluster, output the abnormal sample set, and use the results for typical day or typical scenario construction, sample screening and model training.

[0024] In the present embodiment, the logical relationship between the above steps is as follows: Each dimension of the sample matrix is normalized to obtain normalized features, and statistical indicators of each dimension are defined based on the normalized features for feature selection and dimension compression. First, data normalization is performed to eliminate differences in different feature dimensions and scales. The normalized features help ensure that each feature contributes equally to the result in the clustering process, regardless of its original dimension such as load, temperature, etc. Then, based on these features, statistical indicators such as mean, variance, correlation coefficient, etc. are used to select and reduce the dimension of the features, ensuring that the clustering algorithm only focuses on those features with important information, and removes redundant or noisy features.

[0025] Before clustering, the number of clusters K needs to be set. Fuzzy clustering allows a sample to belong to multiple clusters at the same time, and the membership matrix represents the membership degree of each sample to each cluster, usually with a value between 0 and 1. This membership constraint will serve as the basis for subsequent clustering calculations, controlling the allocation of samples. Fuzzy clustering can effectively handle the mixed characteristics of market states.

[0026] In the clustering process, a dimension weight matrix is introduced for each cluster to reflect the attention of different clusters in different feature dimensions. The center of each cluster is the weighted average of all samples in the cluster, and the cluster center matrix defines the representative value of each cluster in each dimension. By weighting the variance, the weight of each dimension can be dynamically adjusted, so that the dimensions with smaller variance have higher weight in the clustering process, thereby enhancing the consistency within the cluster and suppressing the influence of unimportant features.

[0027] On the basis of model-driven clustering, a probability model is introduced to describe the distribution of samples within the cluster. In traditional clustering methods, the relationship between samples is usually defined based on distance metrics (such as Euclidean distance). By introducing a probability model such as Gaussian mixture model to describe the distribution of samples within each cluster, the statistical characteristics and distribution rules of samples within the cluster can be better captured, and the accuracy and robustness of clustering can be improved.

[0028] In the clustering objective function, a kernel norm regularization term is explicitly introduced. Finally, in order to ensure that the clustering result not only depends on the distance relationship of the samples, but also maintains the low-rank structure of the sub-matrix within the cluster, a kernel norm regularization term is introduced in the clustering objective function. Kernel norm regularization encourages the sample matrix within the cluster to have a low-rank structure, which helps to mine stable patterns in the data, prevent overfitting, and enhance the interpretability and stability of the clustering result.

[0029] The above technical solutions of the embodiment first reduce noise and redundant information through data normalization and feature selection, then assign different feature importance to each cluster through fuzzy membership matrix and dimension weight. Then, the probability model helps to describe the distribution of samples within the cluster, and finally the kernel norm regularization optimizes the clustering result to ensure the stability and low-rank of the structure of samples within the cluster.

[0030] The core goal of the above series of steps is to make the clustering not only focus on the distance relationship between samples, but also capture the complex structure of data, such as the pattern of time series, the nonlinear correlation between variables, etc., so as to improve the accuracy and interpretability of clustering.

[0031] In one embodiment, specifically in step one, the various input and output quantities of spot day-ahead market decision are sorted and expressed to form a set of historical sample matrices that can be used for clustering analysis, and a set of spot day-ahead market decision sample matrices is constructed, including unified mathematical representation of input variables and output variables.

[0032] In this step, first determine the time granularity and statistical period of the sample.

[0033] Construct samples according to the time step (such as 1h, 30min) of the rolling optimization of the day-ahead market. The historical time range of the sample can cover the past 1-3 years, which not only ensures the number of samples, but also covers various climates and operating conditions.

[0034] Let is the number of historical samples, such as the number of historical days if it is unit by "day", or the number of historical time periods if it is unit by "time period"; is the total dimension of the characteristics of each sample, including input and output characteristics.

[0035] For each sample , the following information is extracted from the scheduling and trading system: (1) input characteristics . Hourly load forecasting values, including different levels of network / region / user side; wind power and photovoltaic available power forecasting and related uncertainty quantification indicators, such as prediction interval, variance; probability correlation estimation results between load and wind and light output, such as indicators calculated in advance by correlation coefficient, Copula, etc.; available capacity of independent energy storage, SOC state of charge, operating constraint parameters; available regulation capacity of aggregated resources in virtual power plant, controllable load size, etc.; meteorological and climate information, temperature, humidity, wind speed, irradiance, extreme weather markers, etc.; unit maintenance plan, whether to be maintained, maintenance capacity, start and end time during the corresponding period; power grid construction / retirement status, such as whether a certain line is put into operation, whether a certain unit has been retired, and transmission channel constraints, etc.

[0036] (2) output characteristics . The output curve of each thermal power unit in the day-ahead clearing result or the time-period output decision; the output plan or the amount of curtailment / wind and light curtailment of each wind farm and photovoltaic power station; the charging and discharging power trajectory of independent energy storage, the initial and final SOC; the aggregated output of virtual power plant and internal resource decomposition instructions, etc.

[0037] All input characteristics and output characteristics are unified and concatenated in a fixed order to obtain the sample feature vector:

[0038] wherein: is the feature vector of the th sample, is the feature dimension of each sample, contains input and output characteristics, representing a D-dimensional space, represents the total dimension of the characteristics of each sample, belongs to dimensional space, i.e. the sample has characteristics; : input variable vector, including multivariate load, meteorology, wind and light available power, maintenance plan, grid and power construction status, etc. is the vector space of input variables, represents the input variable vector feature dimension of each sample; : Output variable vector, including the output of thermal power, wind power, photovoltaic, independent energy storage, virtual power plant, etc. or the bidding / clearing result, is the vector space of output variables, represents the feature dimension of the output variable vector of each sample; .

[0039] Stack all samples by rows to get the spot day-ahead market decision sample matrix:

[0040] wherein: is the spot day-ahead market decision sample matrix; represents the value of the th sample at the th feature dimension; ; .

[0041] It should be noted that is a N × D matrix containing D feature values of all N samples.

[0042] In an embodiment, in step two, regarding normalization and statistical feature analysis: due to the large difference in dimension and magnitude of each feature, for example, load is measured in MW, temperature is measured in °C, and output ratio may be 0-1, normalization is needed. That is: to eliminate the influence of different dimensions, each dimension is normalized. Let the minimum and maximum values of the th feature in all samples be:

[0043] Then the normalized feature is:

[0044] wherein: is the normalized feature value; is a small constant to prevent the denominator from being zero.

[0045] After normalization, it satisfies:

[0046] Based on the normalized data, statistical indicators of each dimension can be defined for feature selection and dimension compression: (1) Mean:

[0047] (2) Variance:

[0048] (3) Correlation coefficient, any two dimensions :

[0049] In the formula: : The sample mean of the first dimension feature; : The sample variance of the first dimension feature; : The linear correlation coefficient between the first dimension and the second dimension feature.

[0050] Statistical indicators can be used to: eliminate features with minimal variance (low information); merge / filter highly correlated features, above the threshold. These operations can reduce the dimension of noise and improve the efficiency and stability of subsequent high-dimensional clustering.

[0051] The output of step two above is: (1) Normalized sample matrix: After normalizing each feature dimension in the sample matrix, the normalized matrix is obtained. The normalized matrix makes the dimensions and scales of different features consistent, ensuring that the subsequent clustering algorithm is not affected by the difference in feature dimensions.

[0052] (2) Statistical indicators of each feature dimension: By calculating the mean, variance, correlation coefficient and other statistical indicators, each feature dimension is analyzed, important features are selected, and dimension compression is performed. The specific output includes: mean: the mean of each feature dimension in all samples. Variance: the variance of each feature dimension, used to identify redundant features with small variance. Correlation coefficient: used to measure the correlation between feature dimensions, helping to identify and remove highly correlated features.

[0053] (3) Selected feature dimensions: based on statistical indicators (such as mean, variance and correlation coefficient), select and compress feature dimensions, remove redundant or irrelevant features, and retain the most useful information.

[0054] In short, the output of step two is: the normalized feature matrix, which is used to eliminate the dimensional differences of different features. The results of feature selection and dimension compression include the selection of meaningful features and the feature set after dimension reduction.

[0055] In one embodiment, step three is fuzzy membership and basic constraints: this step introduces the idea of fuzzy clustering, allowing a sample to belong to multiple clusters to different degrees, to better reflect the mixed conditions in the spot day-ahead market, such as having high load characteristics and being affected by extreme weather on a certain day.

[0056] The input of Step Three is the output of Step Two: the normalized sample matrix and the statistical feature analysis results, including mean, variance, correlation coefficient, etc.

[0057] The processing of Step Three: Based on the normalized data and statistical analysis results, Step Three will construct a fuzzy membership matrix, update the cluster centers, and optimize the clustering process through basic constraints such as dimension weight. Specific steps include calculating the similarity between samples, calculating membership, updating cluster centers, adjusting dimension weight, etc.

[0058] The output of Step Three: Fuzzy membership matrix, representing the membership of each sample to each cluster. Cluster center matrix, representing the center position of each cluster. Dimension weight matrix, reflecting the importance of different feature dimensions in clustering.

[0059] The goal of Step Two is to provide clean and effective input data for subsequent clustering through normalization and statistical feature analysis. It removes redundancy, standardizes data, making subsequent clustering more accurate. Step Three uses these cleaned data, through fuzzy clustering method and basic constraints, to further optimize clustering results. This process allows each sample to be flexibly assigned to multiple clusters, while considering the importance of features and updating the representative center of clusters.

[0060] Let the number of clusters be Define the fuzzy membership matrix:

[0061] Where, represents the membership of the th sample to the th cluster, K is the number of clusters.

[0062] The membership constraint can be written as:

[0063] In the formula: represents the membership of the th sample to the th cluster; is the number of clusters; : total number of samples; : sample index; : cluster index.

[0064] In specific implementation examples: U storage: can be stored in memory with Double-precision matrix storage is also available for large-scale samples. Initial value setting: initial membership can be randomly generated to meet the membership constraint; or it can be initialized uniformly in each coarse category according to load level, season, etc., to improve convergence speed.

[0065] The physical meaning is: if a sample has multiple typical mode characteristics in terms of load, climate, etc., its membership may not be zero in multiple clusters; this can avoid information loss caused by hard partitioning and better meet the "fuzzy boundary" characteristics of actual system operation.

[0066] Regarding dimension weight and cluster center: to depict the focus of different clusters in different feature dimensions, dimension weight and cluster center specific to the cluster are introduced.

[0067] A dimension weight matrix is introduced for each cluster:

[0068] And satisfy:

[0069] Wherein: : cluster Importance weight of cluster in the th dimension;

[0070] Define the cluster center matrix:

[0071] Wherein is the center (prototype) value of cluster in the th dimension.

[0072] In the typical fuzzy clustering framework, the update of the cluster center can be written as:

[0073] Wherein: is the fuzzification index; : weighted membership; other symbols are the same as before.

[0074] Cluster center updating is a core step in fuzzy clustering, determining the quality and accuracy of the clustering results. By calculating and updating the membership matrix, cluster centers can be dynamically adjusted to better reflect the sample distribution. Cluster center updates use a weighted average to ensure that each cluster center better represents the samples within the cluster; samples with higher membership degrees have a greater impact on the cluster center. This process is repeated until the cluster centers converge, ultimately yielding a more stable and accurate clustering result. Within the fuzzy clustering framework, cluster center updates adjust the cluster centers based on the membership degree of each sample using a weighted average, progressively optimizing the clustering results. The updated cluster centers provide a representative pattern for each cluster, helping to identify typical operating states of the samples. The output of this step is the final cluster centers, membership matrix, and clustering results for the samples. These results are used in subsequent steps for outlier identification by calculating the distance or membership degree between samples and cluster centers, and can also provide input for market decision-making, strategy optimization, and machine learning models.

[0075] Dimension weights can be adaptively updated based on the intra-cluster weighted variance, for example:

[0076] In the formula: the numerator represents the first... Dimension in cluster The smaller the variance of the "compactness" in the cluster, the more important the dimension is in the cluster.

[0077] The denominator is a normalization factor, guaranteeing the formula Established.

[0078] In fuzzy clustering, adaptive updates of dimension weights adjust the importance of each feature dimension by considering the weighted variance within each cluster. Specifically, when a feature dimension has a highly concentrated sample distribution and low variance within a cluster, its weight is increased, indicating that this dimension provides more stable information in the cluster. Conversely, dimensions with high variance have lower weights because their changes may reflect more complex or unstable features. In this way, adaptive updates of dimension weights automatically adjust the influence of each dimension during the clustering process, ensuring that the most important features dominate the clustering.

[0079] The connection between this adaptive update and subsequent steps is mainly reflected in optimizing clustering results and improving model stability. In steps such as cluster center updating, outlier identification, and application of clustering results, adjusting dimensional weights allows the model to focus more on key features and reduce interference from noisy dimensions. Especially when the dataset contains multiple related or redundant features, updating dimensional weights ensures the effectiveness of the clustering model in high-dimensional data, thereby improving the accuracy of subsequent decision support, market optimization, and prediction models.

[0080] It should be noted that when the variance of a dimension in the cluster is small, it means that the samples in the cluster are highly similar in the dimension and the pattern is stable, and the weight of the dimension will be amplified; if the dimension changes dramatically in the cluster, it means that the dimension does not have good clustering consistency in the cluster, and the weight will be suppressed; through the adaptive weight mechanism, the subspace clustering is realized: different clusters can automatically select a number of dimensions that are most important to themselves, and all dimensions do not need to play a role.

[0081] In an embodiment, the specific process of step four is as follows: 1. Assuming that the feature dimensions in each cluster are independent.

[0082] 2. Model the samples in each cluster as a product of multiple single-dimensional Gaussian distributions, that is, the samples in each cluster are a multi-dimensional distribution composed of multiple independent Gaussian distributions.

[0083] 3. Combine the distributions of all clusters into a mixture Gaussian model, and fuse multiple Gaussian distributions into an overall model by weighting to describe the distribution of the samples.

[0084] This step uses a probability model to depict the distribution of the samples in the cluster.

[0085] In an embodiment, the specific step four probability model and clustering criterion: based on model-driven clustering, the present application introduces a probability model to depict the distribution of the samples in the cluster. Assuming that the features in the cluster are approximately independent, the conditional probability density of the cluster can be modeled as a product of multiple-dimensional Gaussian distributions:

[0086] Wherein: : the i-th cluster; : the probability density of the sample under the condition of belonging to the cluster : the mean parameter of the cluster in the i-th dimension (equivalent or associated with : the variance parameter of the cluster in the i-th dimension; other symbols are the same as before. Define the cluster prior weight , which satisfies , then the overall mixture model is:

[0087] ​​​​​

[0088] in: :sample The overall probability density; :cluster The role of the overall mixture model is to effectively model data by calculating the probability distribution of samples, thereby helping to discover potential patterns in the data, classify them, and make subsequent optimizations and decisions.

[0089] By comparing the distances between model distributions under different parameters, a clustering objective function was constructed. This invention further introduces nuclear norm regularization to form a matrix-structured objective function.

[0090] In practical implementation, a clustering objective function is constructed based on maximum likelihood to make the model... The goal is to approximate the true data distribution as closely as possible. This idea provides the theoretical foundation for the subsequent introduction of nuclear norm regularization.

[0091] In addition, the clustering objective function of the nuclear norm is introduced: the spot market decision samples have a matrix structure and time correlation. For example, the load and output of 24 time periods in a day form a "row vector" in the sample; the output of multiple units / power plants has a high correlation or low-rank structure in the column direction.

[0092] To fully utilize this structure, this invention explicitly introduces a nuclear norm regularization term into the clustering objective function, encouraging the sample submatrices within clusters to exhibit a low-rank structure, corresponding to the "typical operating mode".

[0093] Let the first Each cluster contains a set of sample indexes. The corresponding intra-cluster submatrix can then be represented as:

[0094] in: : No. Sample submatrices of each cluster; :cluster Medium sample size.

[0095] Define the nuclear norm as the sum of the singular values ​​of a matrix:

[0096] in: For nuclear norm; :matrix The One singular value; : The rank of the submatrix within the cluster.

[0097] Taking into account fuzzy membership, dimensional weights, and nuclear norm constraints, this example constructs the following clustering objective function:

[0098] where, : objective function, is the function to be minimized. It consists of two main parts: the first part is the loss function of the clustering or matrix factorization model, which calculates the prediction error. The second part is the regularization term, which is used to prevent the model from overfitting. : number of clusters or number of latent factors, represents the number of clusters or latent factors in the model. : number of samples, represents the number of samples in the dataset. : represents the D-dimensional space. : weight coefficient, represents the degree of sample belongs to cluster or the membership of the sample to the cluster. It is usually used in fuzzy clustering, representing the "membership" of the sample to the cluster , usually in the range [0, 1]. : weight term in matrix factorization, represents the relationship between cluster and feature . It is usually learned during the matrix factorization process, used to represent the contribution of different clusters to a specific feature. : value of sample on feature , is the observed value in the data matrix, usually representing the actual observed value of a sample on a certain feature. : cluster center or latent variable, represents the typical value of cluster on feature , is the latent variable in the model. It is usually learned through optimization algorithm, representing the central feature of the cluster. : regularization parameter, controls the strength of the regularization term. Larger values mean stronger penalties for the regularization term, the model will pay more attention to avoid overfitting; smaller values mean more attention to fitting data. : nuclear norm regularization, is a low-rank constraint for a matrix, used to promote the sparsity or low-rank structure of the matrix. The nuclear norm is the sum of all singular values of a matrix, used to ensure that the representation of the cluster is as simple as possible (low-rank), thereby avoiding overfitting and promoting the generalization ability of the model.

[0099] The first term is the projection clustering cost based on weighted Euclidean distance, reflecting the fuzzy membership and dimension weight; mainly measures the difference between each sample and its cluster center. Through the weighted way (controlled by and ), it calculates the actual observed value of the sample and the feature value of the cluster.The difference between them. This is aimed at minimizing the prediction error, that is, as close as possible to each sample of its corresponding cluster representative features.

[0100] The second is the nuclear norm regularization term, is a balance parameter, used to control the degree of low rank of the intra-cluster sub-matrix, when is larger, the algorithm tends to get more "low rank" intra-cluster structure, and has a stronger inhibitory effect on outliers; when is smaller, it is more biased towards distance minimization, and the clustering is closer to traditional soft clustering.

[0101] The complexity of the model is controlled by the nuclear norm constraint, which encourages the cluster centers or latent feature representations to have a low rank structure and prevents overfitting. This regularization term can help the model better capture the low rank structure in the data and has a stronger inhibitory effect on outliers.

[0102] The regularization parameter : The role of is to balance the weight between the loss function and the regularization term. A larger will make the model pay more attention to simplifying the structure and avoid complexity; a smaller will make the model more focused on fitting the data itself.

[0103] The objective function continues the framework of model-driven projection clustering + constraint conditions, and introduces the nuclear norm as an explicit regularization term to take advantage of the matrix structure and low rank characteristics, according to the characteristics of the spot day-ahead market decision sample matrix.

[0104] In an embodiment, the iterative optimization and update strategy in step six: to solve the minimum value of the above objective function, an alternating iterative optimization strategy similar to EM and fuzzy C-means is used.

[0105] Given , the is updated iteratively by alternating optimization, so that the objective function converges to a local optimal solution. The iterative process is designed as follows: 1. Initialization: (1) According to the load level or season (such as spring, summer, autumn and winter), the samples are roughly divided, and the center of each class is used as the initial cluster center , or random samples can be selected as cluster centers; (2) Initialize the dimension weight ; (3) Initialize the membership matrix satisfies the formula: .

[0106] 2. Membership update: For each sample and cluster , the weighted distance is calculated:

[0107] Then the membership is updated:

[0108] where, : the sample to the weighted Euclidean distance of the cluster , which represents the distance between the sample and the cluster center, and the weight is controlled by . : the weight between the cluster and the feature . This weight reflects the importance of the cluster to the feature . The larger the weight, the greater the contribution of the feature to the clustering result. : the actual observation value of the sample on the feature . : the center value of the cluster , which represents the typical or average value of the cluster on the feature . : the dimension of the feature, which represents the dimension of the feature space of the sample data. : the membership of the sample to the cluster , which represents the degree to which the sample belongs to the cluster . The value of the membership is between 0 and 1. : represents the exponential function, which is used to calculate the exponential decay of the distance between the sample and the cluster center. The closer the distance, the greater the membership. : the temperature parameter, which is used to control the "fuzziness" of the membership. A smaller value makes the boundary of the cluster more fuzzy, and the change of the membership more smooth; a larger value makes the boundary of the cluster more clear, and the change of the membership more drastic. : the square of the weighted distance of the sample to the cluster . The same as the calculation of the weighted distance, but it is calculated for the cluster . : the number of clusters, which represents the total number of clusters that the sample can belong to.

[0109] 3. Cluster center and dimension weight update: When is fixed, the cluster center is updated according to the formula , and the dimension weight is updated according to the formula Updating dimension weights .

[0110] 4. Kernel norm related update: According to the current membership, samples are divided into clusters , obtaining sub-matrix ; singular value decomposition is performed on each and is calculated, and if necessary, the related auxiliary variables are updated by the sub-gradient or proximal operator of the kernel norm to approximate the minimization of the second term in the objective function.

[0111] 5. Convergence criterion: If the membership matrix changes between the adjacent two iterations satisfy

[0112] The algorithm is considered to converge, and the iteration is stopped.

[0113] Where: : preset convergence threshold; : number of iterations. : represents the membership of sample to cluster in the th iteration. The membership is the probability or degree to which a sample belongs to a cluster. : represents the membership of sample to cluster in the th iteration. As the clustering process proceeds, the membership is gradually adjusted. : This is the maximum value of the membership change, which is calculated as the maximum change in the membership of all samples and clusters. By comparing the change in membership of samples to clusters in each iteration, it is determined whether the clustering result is stable. The maximum change value indicates which sample has the maximum membership change to a cluster among all clusters and all samples. If this change is less than a predetermined threshold ( ), the algorithm is considered to have converged. Convergence threshold, used to determine whether clustering has converged. is a very small number set in advance, usually close to zero. When the maximum membership change is less than this threshold, the algorithm is considered to have converged, and the iteration is stopped. The value of depends on the specific application scenario. A smaller will require more accurate convergence, while a larger may lead to earlier stopping.

[0114] In an embodiment, regarding step seven: abnormal sample identification and sample cluster update: to characterize abnormal conditions, the membership matrix obtained in the final iteration is used to identify abnormal samples.Below, this step further carries out abnormal sample identification and cluster result cleaning, to provide high-quality data for subsequent spot day-ahead market modeling.

[0115] 1. Abnormal metric calculation: for each sample Define its minimum cluster distance, and calculate its distance to the nearest cluster according to the final weighted distance:

[0116] Wherein: : Abnormal metric of sample ; : Defined weighted distance; .

[0117] 2. Abnormal determination rule: set distance threshold and membership degree threshold , if: or

[0118] Sample is marked as an abnormal sample, wherein: : Distance threshold; : Membership degree threshold. For samples marked as abnormal, they can be clustered separately or removed from the training set according to their distribution characteristics, thereby completing the clustering preprocessing and abnormal point removal of the spot day-ahead market decision sample matrix set.

[0119] Such samples may correspond to actual situations including rare load curves and output characteristics caused by extreme weather, special clearing results caused by temporary maintenance or large-scale failure, and abnormal points caused by measurement errors or data entry errors.

[0120] 3. Sample cluster update strategy: for samples marked as abnormal, the following processing methods are available: (1) Establish an "abnormal scenario set" separately in subsequent model training, for safety review or risk analysis; (2) Or remove it from the "typical sample set" to avoid bias to the regression / prediction model; (3) If the number of abnormal samples is large and shows a certain pattern, a new cluster can also be defined and the clustering algorithm can be run again.

[0121] At this point, the present application completes the clustering preprocessing of the spot day-ahead market decision sample matrix set, obtaining: (1) The cluster to which each sample belongs and its membership degree; (2) The typical operating mode corresponding to each cluster (cluster center , dimension weight and intra-cluster sub-matrix Commonly characterized) (3) Abnormal sample set.

[0122] These outputs can be directly used for: (1) Building typical days / typical scenarios of day-ahead market; (2) Providing inputs for bidding strategy optimization, flexibility assessment, and security constraint unit partitioning; (3) Providing structured and cleaned training data for subsequent machine learning prediction models.

[0123] In this embodiment, a set of spot day-ahead market decision sample matrices is constructed, including unified mathematical representations of input variables and output variables; normalization and statistical feature analysis are performed on the sample matrices to establish a feature selection and dimension compression mechanism; a high-dimensional projection clustering framework based on a probability model is constructed, fuzzy membership and dimension weight are introduced; a kernel norm regularization term is explicitly introduced into the clustering objective function, and the low-rank structure of the sample cluster sub-matrix is used to improve the clustering quality; an iterative optimization algorithm combining the EM idea is designed to alternately update the membership matrix, cluster center vector, dimension weight and kernel norm related variables; combined with the projection distance and membership of the sample points to the cluster, the abnormal point recognition and sample cluster updating strategy is completed; The historical samples are traversed to generate the clustering preprocessing results of the set of spot day-ahead market decision sample matrices, which are used for subsequent prediction and decision optimization models.

[0124] Among them, the unified sample matrix modeling for spot day-ahead market decision: the multivariate load, uncertain power output, probability correlation, energy storage and virtual power plant state, climate, maintenance plan, power construction and retirement, etc. Time-varying factors are embedded into the high-dimensional sample matrix structure together with the corresponding thermal power, wind power, photovoltaic, energy storage, and virtual power plant output decision, providing a unified data carrier for clustering.

[0125] High-dimensional projection clustering mechanism combined with dimension weight: Unlike traditional K-means clustering in full-dimensional space, the invention introduces a dimension weight matrix for each cluster, automatically learns the contribution of each dimension feature in the corresponding cluster, realizes the description of the subspace clustering structure, and improves the clustering performance and interpretability under high-dimensional data.

[0126] Clustering determination and target construction based on probability model: using the approximate independence assumption of high-dimensional samples in each dimension, modeling each cluster as the product form of multi-dimensional Gaussian distribution, constructing a mixed probability model, and realizing clustering determination through likelihood function or equivalent objective function, improving the fitting ability of complex sample distribution of spot day-ahead market.

[0127] Introducing kernel norm regularization into matrix structure clustering: In view of the fact that the spot day-ahead market decision sample naturally has a matrix and low rank structure, a kernel norm item of a sub-matrix of a cluster is introduced into a clustering objective function, the compactness and pattern consistency of the cluster are enhanced by controlling the rank of each cluster sub-matrix, and stable extraction of the operation mode is realized.

[0128] Abnormal sample identification and clustering update mechanism: A joint criterion combining projection distance and membership degree is used to propose an abnormal sample identification method for abnormal conditions (extreme load, abnormal weather, sudden maintenance, etc.) of the spot day-ahead market, and the cluster structure is dynamically updated according to the distribution of abnormal points to improve the robustness of the clustering result to abnormal data.

[0129] Seamless connection with the spot day-ahead clearing / optimization model: The clustering result is directly used in sample screening, typical day / typical scenario construction, sample weighting and other links, and provides a data basis for the day-ahead market pricing strategy optimization, safety constraint unit bearing, flexibility evaluation and other models.

[0130] The embodiment of the present application realizes unified modeling: In a unified data structure, various input and output quantities of the spot day-ahead market decision are sorted out and expressed to form a historical sample matrix set that can be used for clustering analysis; high-dimensional clustering is realized: in the case of extremely high dimension and different importance of different dimensions, dimension disaster is avoided, and efficient clustering of the sample matrix is realized; feature selection and dimension reduction are realized: the characteristics of the spot day-ahead market decision sample matrix are analyzed from the aspects of data structure, magnitude, dimension and dimension, and feature selection and dimension compression are realized in combination with statistical indicators; clustering robustness is realized: cluster division and cluster update are effectively performed under the premise of abnormal operation conditions and noise data; matrix structure utilization is realized: the low rank property of the sample matrix is utilized, the stability and interpretability of the clustering cluster structure are realized through kernel norm constraint, and a clustering preprocessing result that can be used for spot day-ahead market decision is formed.

[0131] Embodiment: Specific application in a regional power grid spot day-ahead market: The embodiment takes a regional power grid as an example to illustrate a specific application process of the method of the present application.

[0132] 1, data range and scale: historical data: spot day-ahead market operation data in the last 3 years; time granularity: daily sample, including 24 time periods of decision results; sample number: days; feature dimension: about , including system load curve characteristics, prediction power and uncertainty parameters of each wind farm / photovoltaic station, maintenance state, network constraint indicators, weather characteristics, etc.; output dimension: about , including the output of more than ten thermal power units, the output of multiple wind farms / photovoltaic power stations, the output trajectory or related indicators of a number of independent energy storage and virtual power plants; total dimension: .

[0133] 2. Parameter setting: cluster number: , selected according to actual load and weather pattern, seasonal division and model evaluation; fuzzy index: ; temperature parameter: Take the empirical value (such as 0.5-1.5), determine by cross-validation; kernel norm weight parameter: Take 0.01-0.1 range of tuning; convergence threshold: .

[0134] 3. Running process: use the system module described in the specific embodiment to extract nearly 3 years of data from EMS / SCADA, trading system, etc.; construct a sample matrix , normalize and feature selection, get ; set the cluster number , initialize ; execute kernel norm constraint clustering iteration until the convergence criterion is met; calculate and , mark abnormal samples according to formula , or .

[0135] 4. Result interpretation and application: the clustering shows that among the 8 clusters, 2 clusters correspond to high load and high temperature scenarios in summer, characterized by high load peak, high air conditioning load ratio, and relatively stable wind and light output characteristics; 2 clusters correspond to spring and autumn wind and light power output scenarios, and the dimension weight shows that wind power and photovoltaic prediction error and wind and light abandonment indexes have larger weights in these two clusters; The remaining clusters correspond to "low load at night", "high heat load in winter" and other characteristic scenarios. By analyzing the cluster center and dimension weight , it can be clearly seen in each cluster: which characteristics (such as wind power prediction error, energy storage output) play a leading role; which output decisions (such as thermal power unit start-stop combination, energy storage charging and discharging strategy) are highly consistent in the cluster mode.

[0136] The abnormal sample set is mainly concentrated in: days of extreme cold waves, heat waves; dates of sudden large-area maintenance or failure leading to abnormal fluctuations in clearing results. These abnormal samples will be marked as risk scenarios in subsequent bidding strategy training and included in simulation and stress testing in a separate way, rather than directly used for "typical scenario" statistical modeling.

[0137] Embodiment: comparison application with traditional clustering method To further illustrate the beneficial effects of the present application, this embodiment gives a comparison with traditional K-means, fuzzy C-means (FCM) and other methods on the same spot day-ahead market data set.

[0138] 1. Comparison scheme setting Scheme A: traditional K-means clustering (Euclidean distance, full dimension); Scheme B: traditional FCM clustering (fuzzy membership, but no dimension weight and kernel norm constraint); Scheme C: kernel norm high-dimensional clustering method of the present application (with dimension weight and kernel norm constraint, using formula (20) objective function).

[0139] 2. Evaluation index Sum of Squared Errors (SSE); inter-cluster distance index (such as inter-class variance); application effect of clustering results in subsequent prediction model (such as the mean square error RMSE of the daily load prediction model trained by each cluster); interpretability of clustering results (by analyzing whether the importance of each cluster feature conforms to the operation experience).

[0140] 3. Comparison conclusion Scheme A is difficult to significantly reduce SSE and has unclear cluster structure due to "dimension disaster" in high-dimensional scenarios.

[0141] Scheme B improves clustering flexibility to some extent, but still treats all dimensions equally and is difficult to highlight key features such as wind power prediction error and energy storage output.

[0142] Scheme C, i.e. the present application, significantly improves the consistency of each cluster in key feature dimensions by introducing dimension weight and kernel norm constraint , and suppresses the noise influence in secondary feature dimensions; the subsequent prediction model RMSE brought by this is significantly lower than that of schemes A and B, and the operation mode reflected by the cluster center has good physical interpretation.

[0143] In the technical scheme of the embodiment, firstly, multiple loads, uncertain power output, random characteristics, probability correlation between uncertain factors, independent energy storage, virtual power plant, climate, unit maintenance plan, power grid construction, power supply retirement, and other input quantities, and thermal power, wind power, photovoltaic power, independent energy storage, virtual power plant, and other output quantities are sorted out to construct a unified spot day-ahead market decision sample matrix set; then, normalization and statistical feature analysis are performed on the sample matrix to realize feature selection and dimension compression from the aspects of data structure, magnitude, dimension, and dimension; on this basis, fuzzy membership, intra-cluster dimension weight, and probability model are introduced to construct a high-dimensional projection clustering framework suitable for spot day-ahead market decision samples, and a kernel norm regularization term of an intra-cluster sample sub-matrix is embedded in a clustering objective function to extract typical operation modes by using the low-rank structure of the sample matrix; by alternately updating the membership, cluster center, and dimension weight, and combining the weighted distance and membership threshold, the samples are clustered and abnormal samples are identified to complete the clustering preprocessing of the spot day-ahead market decision sample matrix set. The application can relieve the problems of high-dimensional data "curse of dimensionality" and noise interference, automatically highlight key feature dimensions, extract typical operation scenarios with physical meaning, and improve the robustness and modeling efficiency of subsequent prediction and optimization models for abnormal operating conditions.

[0144] Embodiment two The purpose of the embodiment is to provide a computer device including a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor implements the steps of the above method when executing the program.

[0145] The processor can be a general-purpose CPU, GPU, or special-purpose acceleration chip for executing clustering algorithms and matrix operations; the memory is used to store sample matrices , normalized matrices , membership matrices , cluster center matrices , dimension weight matrices , and intermediate calculation results; and the network interface is used for data interaction with EMS / SCADA, trading systems, meteorological systems, and the like.

[0146] The program running on the computer device includes: a data acquisition and preprocessing program; a clustering model construction and iterative optimization program for realizing related calculations in the specific embodiments; and a clustering result management and visualization program for displaying typical load curves and power generation output modes of each cluster.

[0147] Embodiment three The purpose of the embodiment is to provide a computer-readable storage medium.

[0148] A computer readable storage medium having stored thereon a computer program which, when executed by a processor, performs the steps of the above method.

[0149] The computer readable storage medium can be a disk, an optical disk, a USB disk, a ROM, a RAM or any other portable or fixed storage medium; and stores one or more programs which, when executed on the processor, cause the processor to perform all the steps of any method embodiment of the application, including: constructing a spot day-ahead market decision sample matrix ; normalizing the samples; initializing and iteratively updating the membership matrix , the cluster center matrix , the dimension weight matrix ; minimizing the objective function according to the kernel norm constraint ; calculating the anomaly degree index and marking the abnormal samples.

[0150] Through the above hardware and storage medium embodiments, the application can not only be integrated into an existing scheduling / trading platform in the form of a software module, but also can run on a dedicated server in the form of an independent system to realize the automatic clustering preprocessing of the spot day-ahead market decision sample matrix.

[0151] Embodiment Four The purpose of this embodiment is to provide a spot day-ahead market decision sample matrix clustering preprocessing system, which comprises: a sample matrix construction module configured to determine the time granularity and statistical period of the samples and construct a sample matrix; a normalization processing module configured to perform normalization processing on each dimension feature in the sample matrix to obtain normalized features, and define statistical indicators of each dimension based on the normalized features for feature selection and dimension compression; an updating module configured to set the number of clusters for clustering, define a fuzzy membership matrix, obtain membership constraints based on the fuzzy membership matrix, introduce a dimension weight matrix for each cluster, define a cluster center matrix, update the cluster center in a typical fuzzy clustering framework, and adaptively update the dimension weight according to the weighted variance within the cluster; a clustering objective function construction module configured to introduce a probability model to describe the distribution of samples within the cluster on the basis of model-driven clustering; and explicitly introduce a kernel norm regularization term in the clustering objective function; a preprocessing output module configured to solve the minimum value of the clustering objective function by using an alternating iterative optimization strategy; Further abnormal sample identification and clustering result cleaning are performed based on the minimum value of the solved clustering objective function, and the cluster to which each sample belongs and the membership thereof, the typical operating mode corresponding to each cluster and the abnormal sample set are output.

[0152] Embodiment Five The purpose of the embodiment is to provide a computer program product containing instructions which, when run on a computer, cause the computer to perform the method and functions involved in any of the above embodiments. The steps and method embodiments involved in the above embodiment of the device correspond to embodiment one, and the specific implementation can refer to the relevant description of embodiment one. The term "computer readable storage medium" should be understood to include a single medium or multiple media of one or more instruction sets; it should also be understood to include any medium capable of storing, encoding, or carrying a set of instructions for execution by a processor and causing the processor to perform any of the methods in the present application.

[0153] Those skilled in the art should understand that the above-mentioned modules or steps of the present application can be realized by a general computer device, alternatively, they can be realized by program codes executable by a computing device, so that they can be stored in a storage device for execution by a computing device, or they can be respectively made into individual integrated circuit modules, or a plurality of modules or steps among them can be made into a single integrated circuit module to realize. The present application is not limited to any specific combination of hardware and software.

[0154] Although the specific embodiments of the present application are described above in combination with the accompanying drawings, it is not a limitation on the scope of protection of the present application, and those skilled in the art should understand that various modifications or changes made by those skilled in the art on the basis of the technical solutions of the present application without creative labor are still within the scope of protection of the present application.

Claims

1. A spot-to-day-ahead market decision sample matrix clustering preprocessing method, characterized in that, The method comprises the following steps: determine the time granularity and statistical period of the sample and construct a sample matrix; normalize each dimension feature in the sample matrix to obtain normalized features, and define statistical indicators of each dimension based on the normalized features for feature selection and dimension compression; define a fuzzy membership matrix based on the fuzzy membership matrix, and obtain membership constraints based on the fuzzy membership matrix; introduce a dimension weight matrix for each cluster, define a cluster center matrix, update the cluster center in a typical fuzzy clustering framework, and update the dimension weight adaptively according to the weighted variance within the cluster; introduce a probability model to describe the distribution of samples within the cluster based on model-driven clustering; explicitly introduce a kernel norm regularization term in the clustering objective function; use an alternating iterative optimization strategy to solve the minimum value of the clustering objective function; based on the minimum value of the solved clustering objective function, further perform abnormal sample identification and clustering result cleaning, and output the cluster to which each sample belongs, the membership of each sample, the typical operation mode corresponding to each cluster, and the abnormal sample set.

2. The spot day-ahead market decision sample matrix clustering pre-processing method of claim 1, wherein, When constructing the sample matrix, the method specifically comprises the following steps: constructing samples according to the time step of rolling optimization of the day-ahead market, determining the historical time range of the samples, and ensuring the number of samples and covering various climates and operation conditions; extracting input class features and output class features for each sample; unifying and splicing all input class features and output class features in a fixed order to obtain a sample feature vector; stacking all samples in rows to obtain a spot day-ahead market decision sample matrix.

3. The spot day-ahead market decision sample matrix clustering pre-processing method of claim 2, wherein, The extracted input class features include: load prediction values in different time periods; available output prediction and related uncertainty quantification indicators of wind power and photovoltaic power; probability correlation estimation results between load and wind and light output; available capacity, SOC state of charge, and operation constraint parameters of independent energy storage; available adjustment capacity and controllable load size of aggregated resources in a virtual power plant; weather and climate information; unit maintenance plan; power grid construction / retirement status.

4. The spot day-ahead market decision sample matrix clustering pre-processing method of claim 2, wherein, The extracted output class features include: output curve or time-period output decision of each thermal power unit in the day-ahead market clearing result; output plan or curtailment / abandoned wind and light amount of each wind farm and photovoltaic power station; charge and discharge power trajectory and initial and final SOC of independent energy storage; aggregated output of the virtual power plant and internal resource decomposition instruction.

5. The spot day-ahead market decision sample matrix clustering pre-processing method of claim 1, wherein, The statistical indicators of each dimension include mean, variance, and correlation coefficient, and the statistical indicators are used for: removing features with extremely small variance; merging or screening highly correlated features.

6. The spot day-ahead market decision sample matrix clustering pre-processing method of claim 1, wherein, Further abnormal sample identification and clustering result cleaning based on the minimum value of the solved clustering objective function, specifically comprising: define the minimum cluster distance of each sample, and calculate the distance from the nearest cluster according to the final weighted distance; set a distance threshold and a membership threshold, if a sample meets the set conditions, mark the sample as an abnormal sample, and for the samples marked as abnormal, cluster them separately or remove them from the training set according to their distribution characteristics.

7. A spot-to-intra-day market decision sample matrix clustering pre-processing system characterized by, The method comprises the following steps: a sample matrix construction module configured to determine the time granularity and statistical period of the sample and construct a sample matrix; The normalization processing module is configured to perform normalization processing on each dimension feature in the sample matrix to obtain normalized features, and define statistical indexes of each dimension based on the normalized features, which are used for feature selection and dimension compression. The updating module is configured to define a fuzzy membership matrix based on the cluster number, obtain membership constraints based on the fuzzy membership matrix, introduce a dimension weight matrix for each cluster, define a cluster center matrix, update the cluster center in a typical fuzzy clustering framework, and update the dimension weight adaptively according to the weighted variance within the cluster. The clustering objective function construction module is configured to introduce a probability model to describe the distribution of samples within a cluster on the basis of model-driven clustering, and explicitly introduce a kernel norm regularization term in the clustering objective function. The preprocessing output module is configured to solve the minimum value of the clustering objective function by using an alternating iterative optimization strategy. Further abnormal sample identification and clustering result cleaning are performed based on the minimum value of the solved clustering objective function, and a cluster to which each sample belongs and a membership of the cluster, a typical operation mode corresponding to each cluster, and an abnormal sample set are output.

8. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to implement the method of any one of claims 1 to 6.

9. A computer device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor executes the program to implement the steps of the method of any one of claims 1 to 6.

10. A computer-readable storage medium having stored thereon a computer program, characterized in that, The program is executed by the processor to perform the steps of the method of any one of claims 1 to 6.

Citation Information

Patent Citations

  • Possibility fuzzy c mean clustering algorithm based on multiple kernels

    CN105894024A

  • Normalization possibilistic fuzzy entropy clustering method based on Gaussian kernel hybrid artificial bee colony algorithm

    CN106056167A

  • Discriminative dictionary learning based multi-source image fusion denoising method

    CN108198147A

  • Multi-view fuzzy clustering algorithm based on adaptive strategy

    CN118245833A

  • Failure security and protection image fuzzy clustering method based on pairwise strategy

    CN118334392A