Parameterized configuration-oriented power transaction data cleaning and settlement method and parameterized configuration-oriented power transaction data cleaning and settlement system

The power trading data cleaning method, which employs dynamic reverse learning and global calibration, addresses the issues of anomalies and missing data in power trading, thereby improving data consistency and settlement efficiency and ensuring the fairness and accuracy of the power market.

CN120973785APending Publication Date: 2025-11-18GUANGDONG POWER GRID CO LTD INFORMATION CENT
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511126117.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-12
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

The electricity trading data contains anomalies or missing data, making it difficult to ensure data consistency. Parameter adjustments are not flexible enough, and traditional manual processes are inefficient and cannot meet market demands.

Method used

A dynamic reverse learning strategy is used to identify and clean abnormal data, generate a weighted data matrix, calibrate the data through global calibration coefficients, and repair missing data by combining linear interpolation and singular value thresholding algorithms, which automatically triggers the settlement process.

Benefits of technology

Accurately identify and remove abnormal data to improve data consistency and reliability, ensure the fairness and efficiency of settlement, reduce human intervention, and lower the risk of errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120973785A_ABST
    Figure CN120973785A_ABST
Patent Text Reader

Abstract

The invention provides a parameterized-configuration-oriented power transaction data cleaning and settlement method and system, and relates to the technical field of power transaction data, and the method comprises the steps: 1, calculating a data trust value through employing a dynamic reverse learning strategy based on a competitive negative constraint relation between power transaction data packets, cleaning the abnormal data lower than a preset isolation threshold value, and outputting a cleaned data set; step 2, intercepting a time window segment with a preset optimization length for the cleaned data set, dynamically allocating high and low weight intervals according to data integrity, and generating a weighted data matrix; the power transaction data quality and the trans-provincial long-period settlement efficiency are improved through exception cleaning, weighted calibration, missing repair and automatic settlement under parameterized configuration.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of power trading data technology, and in particular to a method and system for cleaning and settling power trading data for parameterized configuration. Background Technology

[0002] With the advancement of electricity market reforms, especially the operation of inter-provincial electricity spot markets, the scale, complexity, and real-time requirements of electricity trading data have significantly increased. The quality of this data is crucial to market fairness, efficiency, and security. The main challenges currently faced include: Data is prone to anomalies or missing data, and existing processing methods are difficult to adapt to its characteristics, which may affect the accuracy of settlement. Cross-provincial transactions involve multiple data sources and lack an effective global calibration mechanism, making data consistency susceptible to problems. Different scenarios have different requirements for data processing parameters, and existing methods are not flexible enough in adjusting parameters, limiting their universality. At the same time, the transaction settlement cycle is short and the data volume is large, and traditional manual processes are inefficient and may not be able to meet the needs of market operation. Summary of the Invention

[0003] The technical problem to be solved by the present invention is to provide a method and system for cleaning and settling power trading data for parameterized configuration, so as to improve data quality and ensure fair and accurate settlement.

[0004] To solve the above-mentioned technical problems, the technical solution of the present invention is as follows: Firstly, a method for cleaning and settling power trading data oriented towards parameterized configuration, the method comprising: Step 1: Based on the competitive negative constraint relationship between power transaction data packets, a dynamic reverse learning strategy is used to calculate the data trust value, and abnormal data below the preset isolation threshold is cleaned, and the cleaned data set is output. Step 2: For the cleaned dataset, extract time window segments of a preset optimized length, dynamically allocate high and low weight intervals according to data integrity, and generate a weighted data matrix; Step 3: For the weighted data matrix, select multiple consecutive benchmark trading cycle scales on the trading time axis to form a benchmark sequence, and construct trading feature vectors according to the chronological order of the timestamps of the benchmark sequence. Generate global calibration coefficients by calculating the overall offset of the trading feature vectors. Apply the global calibration coefficients to calibrate the weighted data matrix as a whole, and output the calibrated data matrix. Step 4: Perform linear interpolation to pre-fill missing data on the calibrated data matrix, construct a feature matrix based on similar user groups, and use the singular value thresholding algorithm for secondary repair to output the complete dataset; Step 5: Input the complete dataset into the settlement process according to the settlement cycle, which will automatically trigger the cross-provincial spot daily clearing calculation and the generation of the inter-provincial transmission and distribution fee settlement list at the end of the month.

[0005] Secondly, a power trading data cleaning and settlement system oriented towards parameterized configuration includes: The anomaly cleaning module is used to calculate the data trust value based on the competitive negative constraint relationship between power transaction data packets, and to clean up the abnormal data that is below the preset isolation threshold, and output the cleaned data set. The matrix generation module is used to extract time window segments of preset optimized length from the cleaned dataset, dynamically allocate high and low weight intervals according to the data integrity, and generate a weighted data matrix. The calibration module is used to select multiple consecutive benchmark trading period scales on the trading time axis to form a benchmark sequence for the weighted data matrix, construct trading feature vectors according to the chronological order of the timestamps of the benchmark sequence, generate global calibration coefficients by calculating the overall offset of the trading feature vectors, apply the global calibration coefficients to calibrate the weighted data matrix as a whole, and output the calibrated data matrix. The repair module is used to perform linear interpolation to pre-fill missing data in the calibrated data matrix, construct a feature matrix based on similar user groups and use the singular value thresholding algorithm for secondary repair, and output a complete dataset. The processing module is used to input the complete dataset into the settlement process according to the settlement cycle, and automatically trigger the cross-provincial spot daily clearing calculation and the generation of the inter-provincial transmission and distribution fee settlement list at the end of the month.

[0006] Thirdly, a computing device, comprising: One or more processors; A storage device for storing one or more programs that, when executed by one or more processors, cause the one or more processors to implement the method.

[0007] Fourthly, a computer-readable storage medium storing a program that, when executed by a processor, implements the method.

[0008] The above-described solution of the present invention has at least the following beneficial effects: Anomaly cleaning based on competitive negation constraints accurately identifies and removes outlier data, reducing interference with settlement results. Dynamic weight allocation and global calibration enhance the consistency and reliability of multi-source data. A combination of linear interpolation pre-filling and singular value thresholding for secondary repair, along with similar user group characteristics, more accurately fills in missing data, improving dataset integrity. Benchmark sequence construction and global calibration coefficient generation effectively address issues such as timestamp synchronization and unit uniformity in cross-provincial transactions, ensuring comparability of data from different periods and sources, supporting settlement fairness. Automatic triggering of daily cross-provincial spot market clearing and month-end transmission and distribution fee settlement lists reduces manual intervention, adapts to the short-cycle settlement needs of the spot market, improves settlement efficiency, and reduces the risk of human error. Attached Figure Description

[0009] Figure 1 This is a flowchart illustrating a parameterized configuration-oriented power transaction data cleaning and settlement method provided by an embodiment of the present invention.

[0010] Figure 2 This is a schematic diagram of a power trading data cleaning and settlement system oriented towards parameterized configuration, provided by an embodiment of the present invention. Detailed Implementation

[0011] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.

[0012] like Figure 1 As shown, an embodiment of the present invention proposes a method for cleaning and settling power trading data oriented towards parameterized configuration. The method includes the following steps: Step 1: Based on the competitive negative constraint relationship between power transaction data packets, a dynamic reverse learning strategy is used to calculate the data trust value, and abnormal data below the preset isolation threshold is cleaned, and the cleaned data set is output. Step 2: For the cleaned dataset, extract time window segments of a preset optimized length, dynamically allocate high and low weight intervals according to data integrity, and generate a weighted data matrix; Step 3: For the weighted data matrix, select multiple consecutive benchmark trading cycle scales on the trading time axis to form a benchmark sequence, and construct trading feature vectors according to the chronological order of the timestamps of the benchmark sequence. Generate global calibration coefficients by calculating the overall offset of the trading feature vectors. Apply the global calibration coefficients to calibrate the weighted data matrix as a whole, and output the calibrated data matrix. Step 4: Perform linear interpolation to pre-fill missing data on the calibrated data matrix, construct a feature matrix based on similar user groups, and use the singular value thresholding algorithm for secondary repair to output the complete dataset; Step 5: Input the complete dataset into the settlement process according to the settlement cycle, which will automatically trigger the cross-provincial spot daily clearing calculation and the generation of the inter-provincial transmission and distribution fee settlement list at the end of the month.

[0013] In this embodiment of the invention, abnormal data cleaning based on competitive negation constraints can accurately identify and remove abnormal data, reducing interference with settlement results. Combined with dynamic weight allocation and global calibration, the consistency and reliability of multi-source data are enhanced. A combination of linear interpolation pre-filling and singular value thresholding algorithm for secondary repair, along with similar user group characteristics, can more accurately fill in missing data, improving dataset integrity. Through benchmark sequence construction and global calibration coefficient generation, problems such as timestamp synchronization and unit unification of multi-source data in cross-provincial transactions are effectively solved, ensuring the comparability of data from different periods and sources, supporting settlement fairness. Automatic triggering of cross-provincial daily spot clearing and the generation of month-end transmission and distribution fee settlement lists reduces manual intervention, adapts to the short-cycle settlement needs of the spot market, improves settlement efficiency, and reduces the risk of human error.

[0014] In a preferred embodiment of the present invention, step 1 above, which involves calculating a data trust value based on the competitive negation constraint relationship between power transaction data packets using a dynamic reverse learning strategy, and cleaning abnormal data below a preset isolation threshold, and outputting a cleaned data set, may include: Step 100: Based on the competitive negation constraint relationship between power transaction data packets, a dynamic reverse learning strategy is used to calculate the data trust value. The dynamic reverse learning strategy is implemented through the following process: Step 102: Identify the mutually exclusive rules between cross-provincial medium- and long-term transaction results and day-ahead security verification data as competitive negation constraints; Step 103: Calculate the probability of each data packet violating the competing negation constraint to generate an initial trust value; Step 104: Dynamically adjust the initial trust value based on the spatial distribution characteristics of the trust values ​​of adjacent data packets to generate the final trust value set; Step 105: Data packets with a final trust value lower than the preset isolation threshold are identified as abnormal data and cleaned, and the cleaned data set is output.

[0015] In this embodiment of the invention, the inter-provincial medium- and long-term transaction results include key information such as the transaction target (e.g., electricity volume, electricity price), transaction period (e.g., daily or weekly transaction periods), transaction entities, and transmission channels; the day-ahead safety verification data includes the maximum transmission capacity of the corresponding transmission channel (e.g., the maximum electricity volume that can be transmitted per hour), grid safety constraints (e.g., the upper limit of line load rate, voltage stability threshold), and maintenance plans (e.g., the channel outage time for a specific period).

[0016] Mutual Exclusion Rule Identification: By analyzing logical conflict scenarios between two types of data, mutual exclusion rules are extracted, for example: Rule 1: If the transaction volume in a certain period of the medium- and long-term transaction results exceeds the maximum transmission capacity of the corresponding transmission channel for that period (from safety verification data), then the two are mutually exclusive (the transaction volume cannot exceed the safety constraints). Rule 2: If the time frame of a medium- to long-term transaction includes the channel maintenance period marked in the security verification data, and the transaction plan does not avoid this period, then the two are mutually exclusive (transactions cannot be executed during the maintenance period). Rule 3: If the electricity price in a medium- or long-term transaction exceeds the price fluctuation range corresponding to that channel in the safety verification data (such as the maximum price limit due to grid loss), then the two are mutually exclusive (the price must comply with grid cost constraints).

[0017] Ultimately, a multi-dimensional set of mutually exclusive rules is formed, which serves as the basis for judging competitive negation constraints.

[0018] Step 103 involves matching the specific data of a single data packet (such as the transaction volume, time, and price) against all the mutually exclusive rules identified in Step 102 to determine whether a rule is violated. For example, for a data packet, check whether its transaction volume exceeds the maximum transmission capacity in Rule 1, and whether the transaction time includes the maintenance period in Rule 2, etc.

[0019] Violation probability calculation: The number of mutual exclusion rules violated by the data packet is counted and denoted as the "number of violated rules". The proportion of the "number of violated rules" to the total number of mutual exclusion rules is calculated as the "basic violation probability". If there are differences in the importance of different rules (e.g., rule 1 involves power grid security and has a higher weight than rule 3), the "basic violation probability" is adjusted by weighting: assign a weight to each rule (e.g., security rules have a weight of 0.6 and price rules have a weight of 0.4), and calculate the "weighted violation probability" (i.e., the sum of the weights of all violated rules divided by the sum of the weights of all rules); the initial trust value = 1 - weighted violation probability (the value ranges from 0 to 1, and the higher the value, the more trustworthy the data packet is).

[0020] Step 104: Define the criteria for determining "adjacency" and filter neighbor data associated with the current data packet: Time-related neighbors: Based on the timestamp of the current data packet, the three closest data packets (a total of six) within the same transmission channel and the same trading day are selected as time-related neighbors. For example, if the current data packet corresponds to the transaction data of "July 1, 2024, 10:00-11:00", then the data packets of 9:00-10:00, 8:00-9:00, 7:00-8:00 (the first three) and 11:00-12:00, 12:00-13:00, 13:00-14:00 (the last three) are selected as time neighbors.

[0021] Logical neighbor association: Select three consecutive data packets belonging to the same trading entity (such as the same power generation company or power sales company) and the same transaction batch (such as medium- to long-term transactions within the same week) as the current data packet as logical neighbors. For example, if the current data packet is the "5th weekly transaction" data of a certain company, then the 4th, 3rd (the first two), and 6th (the last one) transactions are selected as logical neighbors.

[0022] Merge Neighbor Sets: Merge time-related neighbors and logically related neighbors after deduplication to form the "neighbor set" of the current data packet (containing 6-8 data packets).

[0023] Statistical analysis was performed on the initial trust values ​​of the neighbor set to extract three types of features: mean, variance, and trend. Calculate the neighbor mean: sum the initial trust values ​​of all data packets in the neighbor set, and then divide by the number of neighbors. This value reflects the overall trustworthiness of the neighbors.

[0024] Calculate the neighbor variance: First, calculate the difference between the initial trust value of each neighbor and the neighbor mean. Then, square all the differences and sum them up. Finally, divide by the number of neighbors to get the variance. The smaller the variance, the more concentrated the neighbor trust values ​​are (high consistency); the larger the variance, the greater the difference in neighbor trust values ​​(high dispersion).

[0025] Determine neighbor trends: Sort the neighbor set by timestamp from earliest to latest, calculate the initial trust value difference between two adjacent neighbors. If more than 70% of the differences are positive, it is determined to be an "upward trend"; if more than 70% of the differences are negative, it is determined to be a "downward trend"; if the proportion of positive and negative differences is close, it is determined to be a "stable trend". At the same time, record the average change of the trend (e.g., if the average difference per step in a downward trend is -0.03, it means that the trust value of each adjacent neighbor decreases by an average of 0.03).

[0026] Adjustments are made based on the relationship between neighbor characteristics and the initial trust value of the current data packet, depending on the scenario: Scenario 1: The current initial trust value is less than the neighbor mean, and the neighbor variance is small (e.g., variance < 0.02), indicating that the neighbor trust values ​​are concentrated and generally reliable. The current data may have a low trust value due to accidental errors (e.g., typos). Adjustment method: Take the weighted average of the current initial trust value and the neighbor mean (weights are allocated according to the degree of neighbor consistency; for example, the smaller the neighbor variance, the higher the weight of the mean). For example, if the current initial trust value is 0.6, the neighbor mean is 0.7, and the variance is 0.01 (high consistency), then the final trust value = 0.6 × 0.3 + 0.7 × 0.7 = 0.67 (upward adjustment of 0.07).

[0027] Scenario 2: The current initial trust value is greater than the neighbor average, and the neighbor trend is downward. This indicates that the neighbor trust value is gradually decreasing over time, and the current high trust value may not conform to the overall trend (e.g., false reporting). Adjustment method: Adjust downwards by the average magnitude of the neighbor's downward trend. For example, if the neighbor's trust value decreases by an average of 0.03 per step, the current initial trust value is 0.8, and the neighbor average is 0.6, then the final trust value = 0.8 - (0.03 × 2) = 0.74 (adjusted by a 2-step trend magnitude, down by 0.06). If the trend is obvious (e.g., an average decrease of 0.05 per step), the downward adjustment magnitude will be increased (e.g., adjusted by a 3-step magnitude).

[0028] Scenario 3: High neighbor variance (e.g., variance > 0.05) indicates significant differences in the trustworthiness of neighbors, making them less valuable as a reference. Adjustment method: Based on the current initial trust value, make minor adjustments only according to the direction of the majority of neighbors' trust values ​​(e.g., if 60% of neighbors' trust values ​​are > the current value, adjust upwards by 3%; otherwise, adjust downwards by 3%). For example, if the current initial trust value is 0.7, and 60% of neighbors' trust values ​​are > 0.7, then the final trust value = 0.7 × 1.03 = 0.721 (minor adjustment + 3%).

[0029] The trust values ​​obtained after the above adjustments to all data packets are aggregated to form a "final trust value set", with each data packet corresponding to a unique final trust value.

[0030] Step 105: Collect confirmed normal electricity transaction data packets from the past 6 months, extract their final trust values, and statistically analyze their distribution (e.g., minimum 0.5, maximum 0.9, 90% of data concentrated between 0.6 and 0.8). Based on historical distribution, select the "lower limit threshold" of normal data as the isolation threshold (e.g., if 95% of normal data has a trust value ≥ 0.6, then the threshold is set to 0.6). Parameterized adjustment is also supported: for inter-provincial spot transactions (requiring higher precision), it can be increased to 0.65; for medium- to long-term transactions, it can be decreased to 0.55.

[0031] Compare the final trust value of each data packet with the isolation threshold one by one: If the final trust value is greater than or equal to the isolation threshold (e.g., 0.6), it is determined to be "normal data" and retained for further processing. If the final trust value is less than the isolation threshold, it is judged as "abnormal data" and enters the cleaning process.

[0032] Data packets with a final trust value < 0.3 (e.g., trust value 0.25) are filtered out, and their original data is checked for "serious conflicts" (e.g., the transaction volume exceeds the maximum capacity of the transmission channel by 200%, or the transaction time completely covers the maintenance period). Such data is directly marked as "seriously abnormal" and permanently removed from the data set, not included in any subsequent processing, and the reason for removal is recorded (e.g., "volume exceeds the limit by 300%, trust value 0.2"). Data packets with a final trust value between 0.3 and 0.6 (e.g., 0.55) are filtered out and marked as "slightly abnormal," and the abnormal characteristics are recorded (e.g., "transaction volume slightly exceeds the safety check value by 5%, trust value 0.55"). Such data is temporarily stored in the "abnormal data review library" and is not included in the cleaned data set, awaiting manual review (if it is confirmed to be a data entry error, it can be corrected and the trust value recalculated; if it is confirmed to be abnormal, it is removed). All data packets judged as "normal data," as well as slightly abnormal data that has been corrected to normal after manual review (if any), are summarized to form the "cleaned data set."

[0033] Based on the mutual exclusion rules of cross-provincial transactions and security verification, anomaly detection aligns with the business logic of the power market, avoiding misjudgments caused by a "one-size-fits-all" approach and improving the targeting of anomaly identification. Through a two-layer logic of "rule violation probability + dynamic adjustment of adjacent data," it balances the compliance of the data itself with its overall relevance, making the trust value more closely reflect the data's authenticity and credibility. The isolation threshold supports parameterized configuration while retaining a manual review mechanism for minor anomalies, balancing automation efficiency with the rigor of data cleaning.

[0034] In a preferred embodiment of the present invention, step 2 above, which involves extracting a time window segment of a preset optimized length from the cleaned dataset and dynamically allocating high and low weight intervals based on data integrity to generate a weighted data matrix, may include: Step 200: Extract continuous time window segments of a preset optimized length from the cleaned dataset; Step 201: Calculate the data integrity metric for each time window segment, where the integrity metric is the reciprocal of the percentage of missing data points. Step 202: Based on the comparison between the completeness quantification value and the preset completeness threshold, when the completeness quantification value reaches or exceeds the completeness threshold, a high weight coefficient greater than 1 is assigned; when the completeness quantification value is lower than the completeness threshold, a low weight coefficient less than or equal to 1 is assigned. Step 203: Apply high or low weight coefficients to the data points within the time window to generate a weighted data matrix.

[0035] In this embodiment of the invention, the "preset optimized length" of the time window is first determined. This length is set according to the sampling frequency of the power trading data and the subsequent settlement requirements (and can be parameterized). For example, if the data is collected every 15 minutes (i.e., 4 data points per hour), and the daily clearing of spot transactions needs to be summarized in hours, then the preset optimized length can be set to "1 hour" (containing 4 data points) or "1 day" (containing 96 data points).

[0036] The extraction process is as follows: Based on the starting timestamp of the cleaned dataset, the first continuous time window is extracted according to a preset length (e.g., starting from 00:00:00, all data points from 00:00:00 to 00:59:59 are extracted to form a 1-hour window). Subsequent windows are then extracted sequentially in chronological order, with the sliding step size consistent with the preset length (e.g., the first window is 00:00 to 01:00, the second window is 01:00 to 02:00, and so on), ensuring that the windows are continuous and cover the entire dataset. If the remaining data points at the end of the dataset are less than a preset length, the remaining data points are treated as a separate window (or merged with the previous window, depending on the parameter configuration), ultimately resulting in several continuous time window segments of uniform length.

[0037] Step 201: For each time window segment, calculate the complete quantification value according to the following steps: The total number of data points within the statistical window is determined based on the preset length and sampling frequency. For example, a 1-hour window (sampling every 15 minutes) has a total of 4 data points. Each data point within the window is checked for missing values ​​(e.g., null values, data marked as "invalid"). For example, if one data point is missing within a 1-hour window, the missing value is 1. The missing data point is divided by the total number of data points, i.e., percentage = missing value ÷ total number (e.g., 1 ÷ 4 = 0.25). The reciprocal of the missing data point percentage is taken, i.e., completeness measure = 1 ÷ missing percentage (e.g., 1 ÷ 0.25 = 4). The higher this value, the fewer missing data points and the higher the completeness within the window (e.g., when all data points are complete, the missing percentage is 0, and the completeness measure can be set to a maximum value, such as 1000, representing complete completeness).

[0038] Step 202: First, set the "preset integrity threshold" (which can be parameterized and set according to the data reliability requirements). For example, set the threshold to 5 (which corresponds to a missing percentage of ≤0.2, i.e., more than 80% of the data is complete).

[0039] For each time window segment, weight coefficients are assigned according to the following rules: High weight coefficient allocation: If the complete quantification value of a window is greater than or equal to the preset threshold (e.g., the complete quantification value of a window is 6 and the threshold is 5), then the window is judged to have high data integrity and is assigned a high weight coefficient greater than 1 (e.g., 1.2, 1.5, the larger the value, the higher the weight, which can be dynamically adjusted according to the complete quantification value, such as the higher the quantification value, the larger the coefficient).

[0040] Low-weight coefficient allocation: If the complete quantization value of a window is less than the preset threshold (e.g., if the complete quantization value of a window is 3 and the threshold is 5), then the window is considered to have low data integrity and is assigned a low-weight coefficient less than or equal to 1 (e.g., 0.8, 0.5; the smaller the value, the lower the weight. This can also be dynamically adjusted based on the quantization value; for example, the lower the quantization value, the smaller the coefficient). For example: with a preset threshold of 5, a complete quantization value of 8 → high weight 1.3; a quantization value of 4 → low weight 0.7.

[0041] Step 203: Apply the weighting coefficients determined in step 202 to each data point within the corresponding time window to generate a weighted data matrix. Weighted summation of individual data points: For each valid data point (non-missing value) within the window, its original value is multiplied by the weight coefficient of the window (e.g., if the original value of a data point is 100 and the window weight is 1.2, then the weighted value is 100 × 1.2 = 120). Missing data points within the window do not participate in the weighted calculation and are still marked as missing (to be fixed in subsequent steps).

[0042] Construct a weighted data matrix: Arrange the weighted data points of all time windows in chronological order to form a two-dimensional matrix (rows represent time windows, columns represent the position of each data point, and matrix elements are weighted values ​​or missing data markers). For example, 10 one-hour windows, each with 4 data points, will ultimately form a 10-row, 4-column weighted data matrix.

[0043] By dynamically assigning weights, time windows with high data integrity receive higher weight in subsequent processing, reducing interference from low-quality data (those with many missing values) and improving overall data reliability. Both time window truncation and weight allocation support parameterized configuration (such as window length, integrity threshold, and weight coefficients), flexibly adapting to the integrity requirements of data such as high-frequency spot market data and low-frequency medium-to-long-term data. The weighted data matrix preserves the temporal characteristics of the original data while differentiating data quality through weights, thus improving overall processing efficiency.

[0044] In a preferred embodiment of the present invention, step 3 above involves selecting multiple consecutive benchmark trading period scales on the trading time axis to form a benchmark sequence for the weighted data matrix, constructing trading feature vectors based on the chronological order of the timestamps in the benchmark sequence, generating global calibration coefficients by calculating the overall offset of the trading feature vectors, and applying the global calibration coefficients to calibrate the weighted data matrix as a whole, outputting the calibrated data matrix, which may include: Step 300: In the weighted data matrix, select multiple benchmark period scales that are continuously distributed on the transaction time axis to form a benchmark sequence. The number of period scales in the benchmark sequence is consistent with the inter-provincial spot daily clearing and settlement cycle. Step 301: Extract electricity price and electricity volume data for each period in the benchmark sequence according to the timestamp order, and construct a multi-dimensional transaction feature vector; Step 302: Calculate the cosine similarity deviation between the multidimensional transaction feature vector and the historical benchmark vector of the same period. Based on the cosine similarity deviation value and the preset calibration sensitivity parameter, generate a global calibration coefficient with a value range of 0.95 to 1.05. Step 303: Perform global scaling adjustment on all data points in the weighted data matrix using global calibration coefficients, and output the calibrated data matrix.

[0045] In this embodiment of the invention, the matching relationship between the “benchmark trading cycle scale” and the “inter-provincial spot daily clearing and settlement cycle” is clearly defined: the inter-provincial spot daily clearing and settlement cycle is usually a single day (24 hours), and the smallest scale is divided by hour (i.e., 1 hour is 1 cycle scale). Therefore, the “number of cycle scales of the benchmark sequence” must be consistent with the cycle (e.g., 24 consecutive hour scales).

[0046] The specific selection process is as follows: Determine the time axis range: Based on the trading time covered by the weighted data matrix (e.g., from 0:00 to 23:00 on a certain trading day), the trading time axis is divided into scales according to the smallest settlement unit (1 hour). Each scale corresponds to 1 hour of trading data (e.g., 0:00-1:00 is the first scale, 1:00-2:00 is the second scale, and so on until 23:00-24:00 is the 24th scale).

[0047] Selecting continuous benchmark scales: Select continuous and complete 24-hour scales from the time axis (matching the 24-hour daily clearing cycle), requiring that the weighted data matrix corresponding to these scales has no serious missing data (e.g., electricity price and electricity data for each scale are at least 80% complete). For example, selecting the 24-hour scale from 0:00 to 23:00 on a certain trading day constitutes a benchmark sequence containing 24 period scales.

[0048] Check whether the selected scale covers the core time period of daily clearing and settlement (such as peak electricity consumption periods of 9:00-11:00 and 15:00-19:00) to ensure that the sequence can reflect the overall characteristics of the day's transactions. If any core time period is missing, reselect it.

[0049] Step 301: For the baseline sequence determined in step 300, extract key data chronologically and construct feature vectors: Determine the data dimensions to be extracted: Select "electricity price" and "electricity volume" as core features for each period scale (because the two are the core indicators for settlement). Each period scale corresponds to 2 feature values ​​(e.g., the first scale: electricity price 0.5 yuan / kWh, electricity volume 1 million kWh).

[0050] Extraction by timestamp: Sort the 24 period scales of the baseline sequence from earliest to latest by timestamp (0:00→1:00→…→23:00), and extract the electricity price and electricity consumption data for each scale in turn. For example, extract (0.5, 100) for scale 1, (0.52, 95) for scale 2, and so on up to extract (0.48, 80) for scale 24.

[0051] Constructing a multidimensional feature vector: All extracted feature values ​​are concatenated into a one-dimensional vector in chronological order, with the dimension being "number of period scales × 2". For example, the feature vector for 24 period scales is: [0.5, 100, 0.52, 95, ..., 0.48, 80], with a total of 48 dimensions, fully reflecting the changes in electricity price and electricity consumption over time in the benchmark sequence.

[0052] Step 302, determine the two vectors to be compared: The current multidimensional transaction feature vector is the vector constructed in step 301, which contains electricity price and electricity data for 24 period scales, arranged in chronological order as 48 values ​​(e.g., [0.5 yuan, 1 million kWh, 0.52 yuan, 950,000 kWh, ..., 0.48 yuan, 800,000 kWh], denoted as vector A).

[0053] Historical benchmark vector: An average vector constructed using the same method from historical periodic data selected from the historical database (such as the past four Wednesdays of the same season and weekday). For example, first extract the feature vectors of the four historical Wednesdays, then calculate the average value of each corresponding position (e.g., the first value is the average of the first values ​​of the four historical Wednesdays), finally forming a 48-dimensional vector (denoted as vector B) with the same structure as vector A.

[0054] Calculate directional consistency (cosine similarity): Evaluate the degree of directional matching between two vectors. "Directional consistency" is measured by comparing the changing trends of two vectors (e.g., whether they rise simultaneously when electricity prices rise, or whether they fall simultaneously when electricity consumption falls), without considering the magnitude of the values ​​(e.g., the difference between the current electricity price of 0.5 yuan and the historical price of 0.48 yuan does not affect the direction judgment). The specific steps are as follows: Step 1: Calculate the dot product of the two vectors (reflecting the synergy of changes at corresponding positions). This involves multiplying the values ​​at corresponding positions in vectors A and B, then summing all the products to obtain the dot product. For example: the first value of vector A (0.5 yuan) × the first value of vector B (0.48 yuan) = 0.24; the second value of vector A (1 million kilowatt-hours) × the second value of vector B (980,000 kilowatt-hours) = 9800; and so on, calculating the products at the 3rd to 48th corresponding positions; summing all 48 products (e.g., 0.24 + 9800 + ... + 38.4) gives the total dot product (let's assume it's 56800).

[0055] Step 2: Calculate the "magnitude" of each vector (reflecting the overall "length" of the vector, related to its amplitude). The magnitude is used to eliminate the amplitude difference between the two vectors, retaining only the direction information. The calculation method is as follows: first square each value in vector A, then add all the squared results, and finally take the square root of the sum to obtain the magnitude of vector A. Calculate the magnitude of vector B in the same way (e.g., assume it is 245).

[0056] Step 3: Calculate the cosine similarity (a quantitative value of directional consistency): Divide the sum of the dot products from Step 1 by (the magnitude of vector A × the magnitude of vector B) to obtain the cosine similarity. The value ranges from -1 to 1. The closer to 1 (e.g., 0.93), the more consistent the directions of the two vectors (e.g., a high percentage of simultaneous increases and decreases in electricity prices); close to 0 (e.g., 0.1), the directions are unrelated (no correlation in trends); close to -1 (e.g., -0.8), the directions are opposite (e.g., historical decreases in electricity prices when current prices rise).

[0057] The degree of deviation in the consistency of quantification direction: The deviation value = 1 - cosine similarity, which is used to represent the magnitude of the directional difference between the current vector and the historical baseline vector. If the cosine similarity = 0.93 (highly consistent in direction), then the deviation value = 1 - 0.93 = 0.07 (small difference); if the cosine similarity = 0.6 (moderately consistent in direction), then the deviation value = 1 - 0.6 = 0.4 (large difference). The larger the deviation value, the more obvious the deviation of the current vector's trend from the historical trend.

[0058] Based on the calibration sensitivity parameters, a global calibration coefficient (0.95-1.05) is generated: Determine the calibration sensitivity parameter: This parameter is configurable (e.g., 0.5 represents "low sensitivity", 1.0 represents "high sensitivity"). The higher the sensitivity, the greater the adjustment range of the coefficient under the same deviation (i.e., more strictly close to the historical benchmark).

[0059] Calculate the initial calibration coefficient based on the deviation value and sensitivity: Initial calibration coefficient = 1 - (deviation value × sensitivity coefficient), where the "sensitivity coefficient" is a ratio set according to the sensitivity parameter (e.g., 0.3 for low sensitivity and 0.5 for high sensitivity); Example 1: Deviation value = 0.07, sensitivity = 0.5 (low), then initial coefficient = 1 - (0.07 × 0.3) = 0.979 (fine adjustment, close to 1); Example 2: Deviation value = 0.4, sensitivity = 1.0 (high), then initial coefficient = 1 - (0.4 × 0.5) = 0.8 (large adjustment range).

[0060] Limit the range of coefficients (to avoid over-adjustment): The initial calibration coefficient is forcibly constrained to be between 0.95 and 1.05: if the initial calibration coefficient is < 0.95 (such as 0.8 in Example 2), then the final calibration coefficient = 0.95; if the initial calibration coefficient is > 1.05 (such as when the deviation value is negative, i.e., the direction is opposite), then the final calibration coefficient = 1.05; if it is within the range, then the initial coefficient (such as 0.979 in Example 1) is used directly.

[0061] Step 303 employs "global scaling multiplication," where each valid data point (non-missing value) in the weighted data matrix is ​​multiplied by a global calibration coefficient to achieve overall amplitude adjustment. For example, if the global calibration coefficient is 1.02, all data points in the matrix (such as an electricity price of 0.5 yuan and an electricity consumption of 1 million kWh) are multiplied by 1.02, resulting in 0.51 yuan and 1.02 million kWh. For marked missing data points (not involved in the weighted calculation), no adjustment is made initially; the missing marker is retained for later repair. All adjusted data points are rearranged according to their original time window and period scale order to form a calibrated data matrix with the same structure as the weighted data matrix (same number of rows and columns), ensuring that the temporal characteristics and data dimensions remain unchanged, with only numerical values ​​calibrated globally through coefficients.

[0062] By comparing and calibrating the benchmark sequence with historical data of the same period, the deviations caused by timestamp discrepancies and differences in units of measurement in cross-provincial transactions are effectively resolved, ensuring that the trend and magnitude of the data are consistent in the period dimension. The period scale of the benchmark sequence matches the daily clearing and settlement period, and the calibrated matrix can directly support high-frequency (such as daily) settlement calculations, avoiding cross-period data incompatibility issues. The global calibration coefficient is limited to between 0.95 and 1.05 to avoid data distortion caused by over-adjustment.

[0063] In a preferred embodiment of the present invention, step 4 above, which involves performing linear interpolation to pre-fill missing data in the calibrated data matrix, constructing a feature matrix based on similar user groups, and using a singular value thresholding algorithm for secondary repair to output a complete dataset, may include: Step 400: Identify the timestamp positions of all missing data points based on the calibrated data matrix. At the same time, perform cluster analysis based on the load characteristic curves, historical transaction volume, and electricity price fluctuation patterns of power users in the calibrated data matrix to generate a set of similar user groups. Step 401: For each missing data point, calculate the linear interpolation base value based on the complete data points of adjacent before and after timestamps, and generate a collaborative correction value based on the average historical transaction data of users in the same group at the same timestamp in the similar user group set. The base value and the collaborative correction value are weighted and fused according to the preset weight coefficient to generate the final interpolation result of the corresponding missing point. Step 402: Fill all the final interpolation results into the corresponding positions of the calibrated data matrix to generate a pre-filled matrix. Based on the set of similar user groups, extract the electricity-electricity price feature vectors of users in the same group in the same transaction period in the calibrated data matrix to construct a multi-dimensional feature matrix. Step 403: The pre-filled matrix and the multi-dimensional feature matrix are tensor-concatenated according to the user dimension to form a joint repair matrix; Step 404: Dynamically set the singular value threshold based on the proportion of missing regions in the joint repair matrix, and perform secondary repair on the missing regions in the joint repair matrix using a low-rank decomposition approximation algorithm to output the complete dataset, specifically including: Step 4040: Calculate the proportion of missing regions in the joint repair matrix, and dynamically select the singular value threshold according to the preset missing proportion-singular value threshold mapping table. Step 4041: Using the singular value threshold as the decomposition parameter, perform singular value decomposition on the joint repair matrix to generate the left singular vector matrix, the singular value matrix, and the right singular vector matrix. Step 4042: Set the singular values ​​below the singular value threshold in the left singular vector matrix, singular value matrix, and right singular vector matrix to zero to generate the corrected singular value matrix. Step 4043: Reconstruct the matrix by performing matrix product based on the left singular vector matrix, the corrected singular value matrix, and the right singular vector matrix to generate a low-rank approximation matrix; Step 4044: Extract data points from the original missing regions in the low-rank approximation matrix, replace the missing regions in the joint repair matrix, and output the complete dataset.

[0064] In this embodiment of the invention, each element of the calibrated data matrix is ​​traversed, and all missing data points are marked (e.g., null values, special symbols marked "missing"). For each missing point, its corresponding timestamp information (including date and specific time period, such as "2024-07-01 09:00-10:00") is recorded to form a "missing point timestamp list," clearly defining the spatiotemporal location of the missing data. Three core characteristics of power users are extracted from the calibrated data matrix, and similar user groups are divided through cluster analysis: Feature extraction: Load characteristic curve: Records hourly load data for each user on typical days (such as weekdays and weekends), quantified into indicators such as "peak load ratio" (total load from 10:00-12:00 and 16:00-19:00 ÷ total daily load) and "valley load ratio" (total load from 0:00-6:00 ÷ total daily load); Historical transaction volume: Statistics on the user's average daily transaction volume and average weekly fluctuation range over the past 3 months (e.g., the difference between the maximum and minimum daily transaction volume ÷ the average daily transaction volume). Electricity price fluctuation pattern: Analyze the correlation between electricity price and electricity consumption in users' historical transactions (such as whether electricity consumption decreases when electricity price rises), and quantify it as "electricity price sensitivity" (electricity consumption change rate ÷ electricity price change rate).

[0065] Cluster analysis: The above-mentioned characteristic indicators of all users are summarized into a "user feature vector" (e.g., a user's vector is [peak hour percentage 0.4, valley hour percentage 0.2, electricity consumption fluctuation range 0.15, electricity price sensitivity -0.3]). Clustering algorithms (such as grouping by the similarity of feature vectors) are used to calculate the "distance" (such as the sum of squares of the differences in feature indicators) between any two user feature vectors. Users whose distance is less than a preset threshold are grouped into the same group. The grouping process is repeated until all users are assigned to a unique group, and finally multiple "similar user group sets" are formed (such as industrial user groups and commercial user groups, where users in the same group have highly similar load, electricity consumption, and electricity price characteristics).

[0066] Step 401: For each missing data point (timestamp t2), find its adjacent preceding complete data point (timestamp t1, data value a) and following complete data point (timestamp t3, data value b), and calculate the base value by weighting according to the time interval: If t1, t2, and t3 are consecutive time periods (e.g., t1=08:00, t2=09:00, t3=10:00, with an interval of 1 hour), then the base value = (a+b)÷2 (simple average); if the time intervals are not equal (e.g., t1=08:00, t2=09:30, t3=11:00), then weight according to distance: base value = (a×(t3-t2)+b×(t2-t1))÷(t3-t1) (the closer the point, the higher the weight).

[0067] Calculate the co-correction value: From the set of similar user groups generated in step 400, find the users in the same group as the user to whom the missing data point belongs, and extract the historical transaction data of these users at the same timestamp (t2) (such as the electricity / electricity price of the same type of day t2 in the past 4 weeks), and calculate the mean as the collaborative correction value: For example, if the missing point is the electricity consumption of user A at "Wednesday 09:00", the users in the same group are B, C and D, whose electricity consumption at Wednesday 09:00 in the past 4 weeks is 1.2 million kWh, 1.3 million kWh and 1.1 million kWh respectively, then the collaborative correction value = (1.2 million + 1.3 million + 1.1 million) ÷ 3 = 1.2 million kWh.

[0068] Weighted fusion generates the final interpolation result: The two values ​​are combined according to a preset weighting coefficient (e.g., the base value accounts for 70% and the correction value accounts for 30%), and the final interpolation result is = base value × 0.7 + collaborative correction value × 0.3; the weights can be parameterized (e.g., during periods of large data fluctuations, the weight of the correction value is increased to 40%).

[0069] Step 402: Fill the missing parts of the calibrated data matrix with all the final interpolation results calculated in step 401 according to their corresponding timestamp positions to form a "pre-filled matrix" (at this time, there may still be a small number of missing points in the matrix that are not covered by linear interpolation, such as the case where there are no adjacent points at the beginning and end of the time period).

[0070] Constructing a multidimensional feature matrix: Based on a set of similar user groups, a matrix is ​​constructed by extracting the transaction features of users in the same group. For each similar user group, electricity consumption and price data within the same transaction period (such as one week) are selected, and vectors are extracted according to the "user-time period-feature" dimension. For example, if a group has 3 users, the electricity consumption and price features of each user in 5 time periods are used to form a three-dimensional vector of "3 users × 5 time periods × 2 features". The feature vectors of all user groups are concatenated according to the user dimension to form a multi-dimensional feature matrix of "number of users × number of time periods × number of features" (reflecting the common transaction patterns of users in the same group).

[0071] Step 403, "Tensor splicing," involves merging the pre-filled matrix and the multidimensional feature matrix along the user dimension, preserving information from both: the pre-filled matrix has the dimension of "number of users × number of timestamps" (each row represents a user, and each column represents electricity / price data for a timestamp); the multidimensional feature matrix has the dimension of "number of users × number of time periods × number of features" (containing historical common features of users); during splicing, the data of each user in the pre-filled matrix is ​​aligned with the common features of that user in the feature matrix according to the timestamp, forming a joint repair matrix of "number of users × number of timestamps × (original features + common features)" (containing both individual data and group common features).

[0072] Step 4040: Count the number of all missing data points in the joint repair matrix, divide by the total number of data points in the matrix (number of rows × number of columns × number of features) to obtain the missing ratio (e.g., if there are 10,000 total data points and 500 are missing, the ratio is 5%).

[0073] Dynamically select the singular value threshold: A preset "Missing Proportion - Singular Value Threshold Mapping Table" is provided (which can be parameterized and adjusted). For example: missing proportion ≤ 5% → threshold = 0.9 (retain more features); 5% < missing proportion ≤ 10% → threshold = 0.7; missing proportion > 10% → threshold = 0.5 (filter more noise). Based on the calculated missing proportion, the corresponding singular value threshold is matched from the table (e.g., 5% corresponds to 0.9).

[0074] Step 4041, Singular Value Decomposition is the process of splitting a matrix into three special matrices: After decomposition, we obtain the "left singular vector matrix" (reflecting user-dimensional features), the "singular value matrix" (the values ​​on the diagonal represent the importance of different features in the data, with larger values ​​indicating greater importance), and the "right singular vector matrix" (reflecting timestamp-dimensional features). For example, after the joint repair matrix decomposition, the diagonal of the singular value matrix may contain values ​​such as [100, 80, 30, 5, ...], with the first two values ​​being larger (representing primary features) and the last two smaller (representing secondary features or noise).

[0075] Step 4042: Set all singular values ​​in the singular value matrix that are less than the threshold selected in step 4040 to 0: If the threshold is 0.9, all elements in the singular value matrix with values ​​<0.9 (such as 0.8, 0.5) are changed to 0, and the values ​​≥0.9 (such as 1.0, 0.95) are retained to generate the "corrected singular value matrix" (noise filtering feature).

[0076] Step 4043: The low-rank approximation matrix is ​​a matrix reconstructed from the modified matrix, which is close to the original matrix but has fewer missing parts: the reconstruction result is obtained by multiplying the left singular vector matrix by the modified singular value matrix by the right singular vector matrix; since noisy singular values ​​are filtered out, the reconstructed matrix will retain the core features of the data (such as users' transaction trends and common patterns of groups) while filling in some missing parts.

[0077] Step 4044: Extract data points from all original missing regions in the low-rank approximation matrix (i.e., the missing locations marked in step 400); replace the missing locations in the joint repair matrix with these data points to obtain the "complete dataset" (without missing points, and the data simultaneously conforms to individual trends and group commonalities).

[0078] Linear interpolation pre-filling combined with collaborative correction of similar user groups utilizes the continuity of adjacent data and references common group characteristics, reducing the error of a single method; the joint repair matrix integrates individual data and group characteristics, and the singular value thresholding algorithm retains core features by filtering noise, adapting to both scenarios with a small number of missing data and scenarios with a large number of missing data; secondary repair fills in the missing points that linear interpolation cannot cover, and the final output complete dataset conforms to the time series characteristics and group patterns of electricity trading, providing a reliable basis for settlement; In a preferred embodiment of the present invention, step 5 above, which involves inputting the complete dataset into the settlement process according to the settlement cycle, automatically triggering the cross-provincial spot daily clearing calculation and the generation of the month-end inter-provincial transmission and distribution fee settlement list, may include: The complete dataset is divided into daily settlement cycle units for inter-provincial spot trading, automatically triggering daily settlement calculations for inter-provincial spot trading volume and clearing prices, and generating daily settlement results. The complete dataset is also summarized into monthly settlement cycle units for inter-provincial transmission and distribution fees. Based on inter-provincial transmission channel metering data and transmission and distribution price parameters, an inter-provincial transmission and distribution fee settlement list containing per-kilowatt-hour allocation costs and congestion surplus allocation is automatically generated, with the daily settlement cycle units maintaining the same cycle scale as the benchmark sequence.

[0079] In this embodiment of the invention, the time range of the complete dataset is traversed, and the data is divided into "inter-provincial spot daily clearing and settlement cycle units" according to natural days (0:00-24:00). Each unit contains the core data of all inter-provincial spot transactions on that day, including: Information on the transaction entities (such as power generation companies, electricity sales companies, and provinces where electricity is purchased); Time-segmented transaction data (divided according to the period scale of the benchmark sequence, i.e., a 24-hour scale, with each scale including the declared volume, cleared volume, day-ahead clearing price, and real-time clearing price for that time period). Actual operating data (such as the actual power transmission volume and transmission channel loss rate during the period).

[0080] Ensure that the time scale of each day unit is completely consistent with the baseline sequence period scale in step 300 (both are 24-hour scales) to ensure data time sequence alignment.

[0081] Automatic daily settlement calculation: The system reads data from each daily cycle unit and automatically calculates according to the following logic: Settlement electricity confirmation: Compare the "cleared electricity" (transaction result) and "actual transmitted electricity" (metering data) of each trading entity to calculate the deviation electricity (actual transmitted electricity - cleared electricity); if the deviation is within the allowable range (e.g. ±5%), the cleared electricity will be used as the settlement electricity; if it exceeds the range, it will be included in the deviation assessment according to market rules (e.g., the excess part will be charged at 1.2 times the real-time electricity price).

[0082] Time-of-use fee calculation: For each hourly period, the basic fee for that period is calculated as "settled electricity volume × corresponding time-of-use clearing price" (e.g., if the settled electricity volume for a certain period is 1 million kWh and the clearing price is 0.5 yuan / kWh, then the basic fee = 1 million × 0.5 = 500,000 yuan); the basic fees for the 24-hour period are summed to obtain the total basic fee for the day.

[0083] Deviation fee calculation: For deviations exceeding the allowable range, the deviation fee is calculated as "deviation amount × real-time clearing price × deviation coefficient" (e.g., if the deviation amount is 100,000 kWh, the real-time price is 0.55 yuan / kWh, and the coefficient is 1.2, then the deviation fee = 100,000 × 0.55 × 1.2 = 66,000 yuan).

[0084] Daily Clearing Results Summary: The daily settlement volume, basic fees, and deviation fees (if any) of each trading entity are summarized to generate "Inter-provincial Spot Daily Clearing Results", which includes fields such as entity name, date, total settlement volume, total fees (basic fees ± deviation fees), and deviation description.

[0085] The process of aggregating data at a monthly granular level and generating an inter-provincial transmission and distribution fee settlement list involves summarizing daily data into monthly units and combining this with transmission channel data to generate the transmission and distribution fee settlement list. The specific steps are as follows: The daily settlement results are summarized into a "Provincial Transmission and Distribution Fee Settlement Cycle Unit" based on the calendar month (e.g., January 1st to January 31st). The summary includes: Monthly total transaction data: monthly total settlement volume (sum of daily settlement volume) for each trading entity, monthly total basic cost (sum of daily basic costs), and monthly total deviation cost; Inter-provincial power transmission channel metering data: Extract the actual transmitted power (time-period cumulative value) and total channel loss (calculated based on loss rate, such as 2% of the transmitted power) of each inter-provincial power transmission channel within the month from the power grid metering system. Transmission and distribution price parameters: Read the preset parameter library, including the transmission and distribution price per kilowatt-hour (e.g., 0.05 yuan / kWh), congestion surplus calculation rules (e.g., the price difference revenue generated due to transmission constraints), and sharing ratio (e.g., the total transmission and distribution fee is shared by each province according to the proportion of electricity consumption).

[0086] The system generates lists based on monthly summary data according to the following logic: Calculate the total monthly transmission and distribution fee: Total transmission and distribution fee = actual electricity transmitted through inter-provincial transmission channels × transmission and distribution price per kilowatt-hour (e.g., total monthly electricity transmission of 10 million kWh × 0.05 yuan / kWh = 500,000 yuan). The total transmission and distribution fee is allocated based on the proportion of monthly settlement electricity volume of each trading entity (or province): The fee allocated to a certain entity = total transmission and distribution fee × (monthly settlement electricity volume of the entity ÷ total monthly transmission electricity volume) (e.g., if the monthly settlement electricity volume of province A is 3 million kWh, accounting for 30%, then the allocation is 50 × 30% = 150,000 yuan).

[0087] Blocked surplus allocation calculation: Calculate monthly congestion surplus: Congestion surplus = (real-time clearing price - day-ahead clearing price) × actual transmitted electricity (e.g., if the price difference is 0.02 yuan / kWh in a certain period and the transmitted electricity is 500,000 kWh, the surplus for that period is 10,000 yuan, and the cumulative monthly surplus is 200,000 yuan). The surplus is allocated based on the "adjusted electricity consumption" (actual electricity consumption after considering deviations) of the trading entity: The amount allocated to a certain entity = monthly congestion surplus × (the entity's adjusted electricity consumption ÷ total adjusted electricity consumption).

[0088] List content integration: The electricity allocation cost and congestion surplus allocation results are summarized by subject (or province) to generate an inter-provincial transmission and distribution fee settlement list containing fields such as "subject name, monthly settlement electricity volume, transmission and distribution fee allocation amount, congestion surplus allocation amount, and total cost", and the list generation time and data source (such as metering system, daily clearing results) are marked.

[0089] By automatically segmenting data by day and month and triggering calculations, the traditional manual aggregation and accounting are replaced, reducing human error and adapting to the high-frequency settlement needs of the spot market (such as shortening the daily settlement cycle to within 2 hours). The daily settlement cycle unit and the cycle scale (24 hours) of the benchmark sequence are strictly consistent, ensuring that the daily data is aligned with the time sequence of the calibrated transaction feature vector, avoiding settlement deviations caused by data misalignment across time periods. The transmission and distribution fee settlement list automatically includes core contents such as per-kilowatt-hour allocation and congestion surplus distribution, unifying the cost accounting standards for inter-provincial transactions and reducing inter-provincial settlement disputes. Automated allocation and distribution based on actual transmitted electricity volume, transaction deviations, and other data ensure that cost bearing and revenue acquisition are consistent with the actual transaction, guaranteeing fairness for inter-provincial market participants.

[0090] like Figure 2 As shown, embodiments of the present invention also provide a power trading data cleaning and settlement system for parameterized configuration, comprising: The anomaly cleaning module is used to calculate the data trust value based on the competitive negative constraint relationship between power transaction data packets, and to clean up the abnormal data that is below the preset isolation threshold, and output the cleaned data set. The matrix generation module is used to extract time window segments of preset optimized length from the cleaned dataset, dynamically allocate high and low weight intervals according to the data integrity, and generate a weighted data matrix. The calibration module is used to select multiple consecutive benchmark trading period scales on the trading time axis to form a benchmark sequence for the weighted data matrix, construct trading feature vectors according to the chronological order of the timestamps of the benchmark sequence, generate global calibration coefficients by calculating the overall offset of the trading feature vectors, apply the global calibration coefficients to calibrate the weighted data matrix as a whole, and output the calibrated data matrix. The repair module is used to perform linear interpolation to pre-fill missing data in the calibrated data matrix, construct a feature matrix based on similar user groups and use the singular value thresholding algorithm for secondary repair, and output a complete dataset. The processing module is used to input the complete dataset into the settlement process according to the settlement cycle, and automatically trigger the cross-provincial spot daily clearing calculation and the generation of the inter-provincial transmission and distribution fee settlement list at the end of the month.

[0091] It should be noted that this system is a system corresponding to the above method. All implementation methods in the above method embodiments are applicable to this embodiment and can achieve the same technical effect.

[0092] Embodiments of the present invention also provide a computing device, including: a processor and a memory storing a computer program, wherein the computer program, when executed by the processor, performs the method described above. All implementations in the above method embodiments are applicable to this embodiment and can achieve the same technical effects.

[0093] Embodiments of the present invention also provide a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the method described above. All implementations in the above method embodiments are applicable to this embodiment and can achieve the same technical effects.

[0094] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A method for cleaning and settling power trading data for parameterized configuration, characterized in that, The method includes: Step 1: Based on the competitive negative constraint relationship between power transaction data packets, a dynamic reverse learning strategy is used to calculate the data trust value, and abnormal data below the preset isolation threshold is cleaned, and the cleaned data set is output. Step 2: For the cleaned dataset, extract time window segments of a preset optimized length, dynamically allocate high and low weight intervals according to data integrity, and generate a weighted data matrix; Step 3: For the weighted data matrix, select multiple consecutive benchmark trading cycle scales on the trading time axis to form a benchmark sequence, and construct trading feature vectors according to the chronological order of the timestamps of the benchmark sequence. Generate global calibration coefficients by calculating the overall offset of the trading feature vectors. Apply the global calibration coefficients to calibrate the weighted data matrix as a whole, and output the calibrated data matrix. Step 4: Perform linear interpolation to pre-fill missing data on the calibrated data matrix, construct a feature matrix based on similar user groups, and use the singular value thresholding algorithm for secondary repair to output the complete dataset; Step 5: Input the complete dataset into the settlement process according to the settlement cycle, which will automatically trigger the cross-provincial spot daily clearing calculation and the generation of the inter-provincial transmission and distribution fee settlement list at the end of the month.

2. The method for cleaning and settling power trading data oriented towards parameterized configuration according to claim 1, characterized in that, Based on the competitive negation constraint relationship between power transaction data packets, a dynamic reverse learning strategy is used to calculate the data trust value, and abnormal data below a preset isolation threshold is cleaned. The cleaned data set is output, including: Based on the competitive negation constraint relationship between power transaction data packets, a dynamic reverse learning strategy is adopted to calculate the data trust value. The dynamic reverse learning strategy is implemented through the following process: Identify mutually exclusive rules between cross-provincial medium- and long-term transaction results and day-ahead security verification data as competitive negation constraints; Calculate the probability of each data packet violating the competing negation constraint to generate an initial trust value; The initial trust value is dynamically adjusted based on the spatial distribution characteristics of the trust values ​​of adjacent data packets to generate the final trust value set; Data packets whose final trust value is lower than the preset isolation threshold are identified as abnormal data and cleaned, and the cleaned data set is output.

3. The method for cleaning and settling power trading data oriented towards parameterized configuration according to claim 2, characterized in that, For the cleaned dataset, time windows of preset optimized length are extracted, and high and low weight intervals are dynamically assigned based on data completeness to generate a weighted data matrix, including: For the cleaned dataset, extract continuous time window segments of a preset optimized length; Calculate the data integrity metric for each time window segment, where the integrity metric is the reciprocal of the percentage of missing data points; Based on the comparison between the completeness quantification value and the preset completeness threshold, when the completeness quantification value reaches or exceeds the completeness threshold, a high weight coefficient greater than 1 is assigned; when the completeness quantification value is lower than the completeness threshold, a low weight coefficient less than or equal to 1 is assigned. A weighted data matrix is ​​generated by applying high or low weight coefficients to data points within a time window.

4. The method and system for cleaning and settling power trading data oriented towards parameterized configuration according to claim 3, characterized in that, For the weighted data matrix, multiple consecutive benchmark trading cycle scales on the trading time axis are selected to form a benchmark sequence, and a trading feature vector is constructed according to the chronological order of the timestamps of the benchmark sequence. A global calibration coefficient is generated by calculating the overall offset of the trading feature vector. The weighted data matrix is ​​calibrated globally using global calibration coefficients, and the calibrated data matrix is ​​output, including: In the weighted data matrix, multiple benchmark period scales that are continuously distributed on the transaction time axis are selected to form a benchmark sequence. The number of period scales in the benchmark sequence is consistent with the inter-provincial spot daily clearing and settlement cycle. Electricity price and electricity volume data for each period in the benchmark sequence are extracted in time stamp order to construct a multi-dimensional transaction feature vector; Calculate the cosine similarity deviation between the multidimensional transaction feature vector and the historical benchmark vector of the same period. Based on the cosine similarity deviation value and the preset calibration sensitivity parameter, generate a global calibration coefficient with a value range of 0.95 to 1.

05. The global calibration coefficients are used to scale all data points in the weighted data matrix, and the calibrated data matrix is ​​output.

5. The method for cleaning and settling power trading data oriented towards parameterized configuration according to claim 4, characterized in that, Linear interpolation is performed on the calibrated data matrix to pre-fill missing data. A feature matrix is ​​constructed based on similar user groups, and a singular value thresholding algorithm is used for secondary data repair. The complete dataset is output, including: Based on the calibrated data matrix, the timestamp locations of all missing data points are identified. At the same time, cluster analysis is performed based on the load characteristic curves of power users, historical transaction volume, and electricity price fluctuation patterns in the calibrated data matrix to generate a set of similar user groups. For each missing data point, a linear interpolation base value is calculated based on the complete data points with adjacent preceding and following timestamps, and a collaborative correction value is generated based on the average of historical transaction data of users in the same group at the same timestamp in the similar user group set. The base value and the collaborative correction value are then weighted and fused according to a preset weight coefficient to generate the final interpolation result for the corresponding missing point. All final interpolation results are filled into the corresponding positions of the calibrated data matrix to generate a pre-filled matrix. Based on the set of similar user groups, the electricity-electricity price feature vectors of users in the same group within the same transaction period are extracted from the calibrated data matrix to construct a multi-dimensional feature matrix. The pre-filled matrix and the multi-dimensional feature matrix are tensor-concatenated according to the user dimension to form a joint repair matrix; The singular value threshold is dynamically set according to the proportion of missing regions in the joint repair matrix. The missing regions in the joint repair matrix are then repaired using a low-rank decomposition approximation algorithm, and a complete dataset is output.

6. The method for cleaning and settling power trading data oriented towards parameterized configuration according to claim 5, characterized in that, The singular value threshold is dynamically set based on the proportion of missing regions in the joint repair matrix. A low-rank decomposition approximation algorithm is used to perform secondary repair on the missing regions in the joint repair matrix, outputting a complete dataset, including: Calculate the proportion of missing regions in the joint repair matrix, and dynamically select the singular value threshold according to the preset missing proportion-singular value threshold mapping table; Using the singular value threshold as the decomposition parameter, singular value decomposition is performed on the joint repair matrix to generate a left singular vector matrix, a singular value matrix, and a right singular vector matrix. Set the singular values ​​below the singular value threshold in the left singular vector matrix, singular value matrix, and right singular vector matrix to zero to generate the corrected singular value matrix. A low-rank approximation matrix is ​​generated by matrix product reconstruction based on the left singular vector matrix, the corrected singular value matrix and the right singular vector matrix. Extract data points from the original missing regions in the low-rank approximation matrix, replace the missing regions in the joint repair matrix, and output the complete dataset.

7. The method for cleaning and settling power trading data oriented towards parameterized configuration according to claim 6, characterized in that, Inputting the complete dataset into the settlement process according to the settlement cycle will automatically trigger the daily clearing calculation of cross-provincial spot market transactions and the generation of the monthly inter-provincial transmission and distribution fee settlement list, including: The complete dataset is divided into daily settlement cycle units for inter-provincial spot trading, automatically triggering daily settlement calculations for inter-provincial spot trading volume and clearing prices, and generating daily settlement results. The complete dataset is also summarized into monthly settlement cycle units for inter-provincial transmission and distribution fees. Based on inter-provincial transmission channel metering data and transmission and distribution price parameters, an inter-provincial transmission and distribution fee settlement list containing per-kilowatt-hour allocation costs and congestion surplus allocation is automatically generated, with the daily settlement cycle units maintaining the same cycle scale as the benchmark sequence.

8. A power trading data cleaning and settlement system for parameterized configuration, the system implementing the method as described in any one of claims 1 to 7, characterized in that, include: The anomaly cleaning module is used to calculate the data trust value based on the competitive negative constraint relationship between power transaction data packets, and to clean up the abnormal data that is below the preset isolation threshold, and output the cleaned data set. The matrix generation module is used to extract time window segments of preset optimized length from the cleaned dataset, dynamically allocate high and low weight intervals according to the data integrity, and generate a weighted data matrix. The calibration module is used to select multiple consecutive benchmark trading period scales on the trading time axis to form a benchmark sequence for the weighted data matrix, construct trading feature vectors according to the chronological order of the timestamps of the benchmark sequence, generate global calibration coefficients by calculating the overall offset of the trading feature vectors, apply the global calibration coefficients to calibrate the weighted data matrix as a whole, and output the calibrated data matrix. The repair module is used to perform linear interpolation to pre-fill missing data in the calibrated data matrix, construct a feature matrix based on similar user groups and use the singular value thresholding algorithm for secondary repair, and output a complete dataset. The processing module is used to input the complete dataset into the settlement process according to the settlement cycle, and automatically trigger the cross-provincial spot daily clearing calculation and the generation of the inter-provincial transmission and distribution fee settlement list at the end of the month.

9. A computing device, characterized in that, include: One or more processors; A storage device for storing one or more programs, which, when executed by one or more processors, cause the one or more processors to implement the method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program that, when executed by a processor, implements the method as described in any one of claims 1 to 7.

Citation Information

Cited By

  • Industrial missing data generation method and system for time-sharing power supply

    CN122220708A