Processing method applied to auxiliary service settlement data
Through innovative mechanisms such as modular design and multi-table data linkage extraction, the problems of inefficient, high error rate and difficult traceability in auxiliary service settlement data processing are solved, and efficient and accurate data processing and statistical results are achieved.
Patent Information
- Application Number
- CN202510471674.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-06-20
AI Technical Summary
The existing technology is inefficient, has high error rate and difficult traceability in the processing of auxiliary service settlement data, and cannot meet the efficient operation needs of the power trading market.
Modular design integrates data extraction, calculation, diagnosis and correction functions, and introduces innovative mechanisms such as multi-table data linkage extraction and error data path tracking. Automatic processing is achieved through data extraction modules, data calculation modules, data self-diagnosis modules, error data traceability modules and data correction modules.
It significantly improves data processing efficiency, reduces time consumption, ensures data accuracy and accuracy, solves the statistical chaos caused by inconsistent subclasses, and realizes quality management of the entire life cycle of data.
Smart Images

Figure CN120182004A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data processing, and particularly relates to a data processing method applied to auxiliary service settlement data. Background Art
[0002] With the deepening of the power market reform, the power trading system has gradually formed a multi-level market structure covering medium- and long-term trading, spot trading, and auxiliary service trading; in this context, as an important link to ensure the stable operation of the power grid, the data processing scale of auxiliary service settlement shows exponential growth. Auxiliary service settlement involves a large number of complex calculation scenarios, including the cost accounting of various service types such as frequency modulation, reserve capacity, and reactive power compensation. The data dimension covers multiple levels such as the on-grid power, per-kWh price, assessment, and compensation of different power plants of power generation enterprises. The traditional data processing method highly relies on manual operations. Staff need to repeatedly switch between multiple spreadsheets and complete data sorting and accounting through manual queries, copy-pasting, formula calculations, etc. This processing mode can be maintained when the data volume is small, but with the increasing richness of power trading varieties and the refinement of data granularity, its drawbacks are gradually exposed, which are specifically manifested in the following aspects: 1. Low data processing efficiency: In the prior art, the statistics and calculation of auxiliary service settlement data need to be operated across multiple sheets. For example, the monthly auxiliary service assessment cost accounting needs to be associated with various assessment sheets such as the primary frequency modulation service and generation plan of power plants. However, there is no automated mechanism for data association between different sheets, resulting in a large amount of time-consuming manual matching by staff. Statistical data shows that in a single settlement task, about 40% of the time is spent on data searching and format conversion, and the efficiency of the core calculation link is seriously dragged down; 2. Poor data classification consistency: Power trading data often has problems with inconsistent sub-classification quantities due to market rule adjustments or enterprise declaration differences. For example, new assessment items may be added to the auxiliary service settlement data in different months. Such differences lead to problems such as field misalignment and calculation logic confusion during data aggregation, further causing settlement result deviations. The prior art lacks a dynamic classification adaptation mechanism and cannot automatically identify and unify the data hierarchy. Eventually, manual intervention is required for correction, increasing the operation complexity and error risk.
[0003] 3. Prone to errors in manual processing of high-precision data: The settlement of ancillary services involves a large amount of high-precision numerical calculations, and problems such as digital misrecording and decimal point misalignment are very likely to occur during manual input or modification. For example, in an actual case of a provincial power trading center, the deviation of a single settlement amount exceeded one million yuan due to manual input errors. Although some enterprises adopt basic formula verification functions, they can only detect simple logical errors and lack effective means to detect hidden errors in complex calculation links.
[0004] 4. Low efficiency in error tracing and correction: When the settlement result is abnormal, the existing technology needs to locate the problem by tracing the original data line by line, which takes an extremely long time. For example, in a certain settlement, due to the wrong version of the price parameter table, the entire volume of data was abnormal, and it took the staff 3 days to locate a certain cell in a specific table. In addition, the corrected data needs to re-execute the complete calculation process, lacking a local data update and incremental calculation mechanism, which further prolongs the processing cycle.
[0005] Although there are some data processing tools in the existing technology, their functions are highly fragmented and cannot meet the comprehensive requirements for data integrity, calculation accuracy, and process automation in the scenario of ancillary service settlement in the power industry. For example, although conventional database scripts can achieve batch queries, they lack the ability to adapt to the special structure of power trading data; although commercial BI tools support visual analysis, their calculation engines are difficult to handle high-precision operations with six decimal places, and they cannot embed customized error diagnosis logic.
[0006] Therefore, it is necessary to design a processing method for ancillary service settlement data to solve the above technical problems. Summary of the Invention
[0007] The technical problem to be solved by the present invention is to provide a processing method for ancillary service settlement data, which integrates data extraction, calculation, diagnosis, and correction functions through modular design, and at the same time introduces innovative mechanisms such as multi-table data linkage extraction and error data path tracing, so as to fundamentally solve the problems of low efficiency, high error rate, and difficult traceability in the existing technology, and provide technical support for the efficient operation of the power trading market.
[0008] To achieve the above technical effects, the technical solution adopted by the present invention is: A processing method for ancillary service settlement data includes the following steps: S1, input initial data, where the initial data includes ancillary service settlement data in the power trading market; S2, through a data extraction module, extract a target data set from the initial data based on preset column name and row name rules; S3, perform numerical calculations on the extracted target data set through a data calculation module and process the empty set data therein; S4. Determine, via the data self-diagnosis module, whether there is an error in the calculation result that exceeds the preset threshold. If so, jump to step S5; otherwise, jump to step S6; S5. Use the error data tracing module to locate the coordinates of the problem data according to the data classification and calculation path; S6. Correct the problem data via the data correction module; S7. Output the corrected data to the target table in a preset format via the data output module.
[0009] Preferably, the data extraction module includes: A column name positioning unit for matching the column names in the table according to a regular expression; A row name association unit for associating the row names with the column names through a multi-dimensional index; A dynamic extraction unit for generating a target data set according to the matching result.
[0010] Preferably, the processing of the empty set data in step S3 includes: Replacing the null value with a preset default value, or generating a filling value based on the interpolation algorithm of adjacent data.
[0011] Preferably, the error determination method of the data self-diagnosis module is: Statistically set the threshold according to historical data. If the deviation of the calculated data exceeds the threshold, it is determined as abnormal data.
[0012] Preferably, the error data tracing module also executes: Construct a tree structure of the data calculation path and record the input-output relationship of each step of the calculation; Use the reverse tracing algorithm to trace back from the abnormal data node to the original input node to locate the error source.
[0013] Furthermore, the execution process of the error data tracing module includes: Calculation path modeling: Based on the operation log of the data calculation module, construct a tree-like topological structure with the initial data as the root node, the intermediate calculation results as the branch nodes, and the final output as the leaf nodes. Each node records the following metadata: The data source table and coordinates; The formulas or algorithms used in the calculation, including linear interpolation and weighted average; The list of dependent parent nodes; Reverse tracing algorithm: Starting from the abnormal leaf node, trace back layer by layer to the parent node; Store all the error data coordinates in the error set matrix, traverse all the nodes in the error set matrix, and determine the corresponding error data categories; Intelligent repair unit: Automatically correct the data calculation for the error data with large differences.
[0014] Preferably, the data correction module includes: A rule base that stores preset correction rules and associated field constraint conditions; An automatic correction unit that batch replaces or recalculates formulas for problem data according to the rule base.
[0015] Furthermore, the data correction module further includes: A multi-level rule base architecture: A basic rule layer that defines data format constraints; A correction strategy selector that matches the best correction method according to the error type: A correction effect verification loop: Re-execute steps S3 to S4 on the corrected data. If there are still anomalies, activate the manual review interface; Preferably, the data output module supports: Multi-table linked output, automatically generating associated tables according to data classification; A data formatting unit that rounds the calculation result to two decimal places and adds a unit identifier.
[0016] Preferably, the initial data integrates the following sources: Data such as various assessments, assessment refunds, compensations, and compensation allocations in the auxiliary service transaction data of different power plants; The data extraction module performs tagged classification on data from different sources and establishes a cross-table association index.
[0017] Preferably, a computer program for executing the processing method for auxiliary service settlement data described in claims 1 to 9; the program includes: A data input interface that supports direct import of CSV, Excel, and databases; Parallel computing to accelerate the processing efficiency of the data calculation module; The beneficial effects of the present invention are: 1. Improve data processing efficiency and reduce time consumption: This method integrates traditional scattered operation processes into an automated processing link through modular design, significantly improving data processing efficiency; the data extraction module uses regular expression matching and multi-dimensional index association technology to automatically identify target column names and row names in the table; the data calculation module introduces the MapReduce framework to split massive calculation tasks into subtasks for parallel processing; through a preset rule engine, the system automatically triggers the flow between modules.
[0018] 2. Solve the problem of quickly extracting valid data from massive data: This method breaks through the limitations of traditional manual screening through an innovative data extraction mechanism. In view of the characteristics of multi-source heterogeneous power trading data, the data extraction module automatically parses the file structure and establishes a dynamic index library; adopts feature filtering technology based on business rules to effectively eliminate redundant data; and automatically associates relevant data scattered in multiple tables through primary key intelligent matching technologies such as power generation enterprise codes and time intervals. 3. Solve the statistical chaos problem caused by inconsistent sub-classifications in the same type of data: This method completely solves the calculation errors caused by sub-classification differences through dynamic classification adaptation and data standardization technologies: for the addition and deletion of fields caused by changes in market rules, the system automatically matches the old and new fields through a version control mechanism; built-in power industry data cleaning rules such as unit unified conversion and null value filling strategies prevent statistical errors caused by format differences from the source.
[0019] 4. Build data diagnosis, error tracing, and correction functions to ensure data correctness: This method realizes the whole life cycle quality management of data through a closed-loop error correction mechanism: In the data self-diagnosis module, an alarm is immediately issued when the deviation between the calculated data sum and the actual data sum exceeds the threshold; the error data tracing module can locate the atomic-level error source through the calculation path topology; the data correction module integrates a multi-level rule library to support automatic repair. BRIEF DESCRIPTION OF THE DRAWINGS Figure 1 It is a schematic diagram of the process framework of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0020] Example 1: As Figure 1 shown, a processing method applied to ancillary service settlement data includes the following steps: S1. Input initial data, which includes ancillary service settlement data of the power trading market; S2. Through the data extraction module, based on preset column name and row name rules, extract a target data set from the initial data; S3. Numerically calculate the extracted target data set through the data calculation module and process the empty set data therein; S4. Through the data self-diagnosis module, determine whether there is an error exceeding the preset threshold in the calculation result. If so, jump to step S5; otherwise, jump to step S6; S5. Through the error data tracing module, locate the coordinates of the problem data according to the data classification and calculation path; S6. Correct the problem data through the data correction module; S7. Output the corrected data to the target table in a preset format through the data output module.
[0021] Preferably, the data extraction module includes: A column name positioning unit for matching the column names in the table according to a regular expression; A row name association unit for associating row names with column names through multi-dimensional indexing; A dynamic extraction unit for generating a target data set according to the matching result.
[0022] Preferably, the processing of the empty set data in step S3 includes: Replacing the null value with a preset default value, or generating a filling value based on an interpolation algorithm of adjacent data.
[0023] Preferably, the error determination method of the data self-diagnosis module is: Setting a threshold according to historical data statistics. If the calculated data deviation exceeds the threshold, the data is determined to be abnormal data.
[0024] Preferably, the error data traceability module further performs: Constructing a tree structure of the data calculation path and recording the input-output relationship of each calculation step; Through a reverse tracing algorithm, tracing back from the abnormal data node to the original input node to locate the error source.
[0025] Furthermore, the execution process of the error data traceability module includes: Calculation path modeling: Based on the operation log of the data calculation module, constructing a tree-like topological structure with the initial data as the root node, the intermediate calculation results as the branch nodes, and the final output as the leaf nodes, where each node records the following metadata: The data source table and coordinates; The formulas or algorithms used in the calculation, including linear interpolation and weighted average; The list of dependent parent nodes; Reverse tracing algorithm: Starting from the abnormal leaf node, tracing back layer by layer to the parent node; storing all the error data coordinates in the error set matrix, traversing all the nodes in the error set matrix, and determining the corresponding error data category; Intelligent repair unit: Automatically correcting the data calculation for the error data with large differences.
[0026] Preferably, the data correction module includes: A rule library for storing preset correction rules and associated field constraint conditions; An automatic correction unit for batch replacing or formula recalculation of the problem data according to the rule library.
[0027] Furthermore, the data correction module further includes: Multi-level rule library architecture: Basic rule layer: Define data format constraints, such as numerical range and number of decimal places; Correction strategy selector, which matches the best correction method according to the error type: Single-point data anomaly: Invoke Lagrange interpolation method or replace it with the median of the same type of data; Correction effect verification loop: Re-execute steps S3 to S4 on the corrected data. If there are still anomalies, activate the manual review interface; Preferably, the data output module supports: Multi-table linked output, automatically generating associated tables according to data classification; Data formatting unit, rounding the calculation result to two decimal places and adding unit identifiers.
[0028] Preferably, the initial data integrates the following sources: Data such as various assessments, assessment refunds, compensations, and compensation sharing in the auxiliary service transaction data of different power plants; The data extraction module classifies the data from different sources with tags and establishes a cross-table association index.
[0029] Preferably, a computer program for executing a processing method for auxiliary service settlement data described in claims 1 to 9; the program includes: Data input interface, supporting direct import of CSV, Excel, and database connections; Parallel computing to accelerate the processing efficiency of the data calculation module; Embodiment 2: To enable efficient processing of data, a corresponding program is designed, as described in Table 1 below: Table 1: Pseudo-code of a program for efficient processing of business data;
[0030] The above embodiments are only the preferred technical solutions of the present invention and should not be regarded as limitations to the present invention. The embodiments and features in the embodiments of this application can be arbitrarily combined with each other without conflict. The protection scope of the present invention should be the technical solutions recorded in the claims, including equivalent replacement solutions of the technical features in the technical solutions recorded in the claims. That is, equivalent replacement improvements within this scope are also within the protection scope of the present invention.
Claims
1. A method for processing auxiliary service settlement data, characterized in that: The following steps are involved: S1, input initial data, which includes ancillary service settlement data of the power trading market; S2, extracting the target data set from the initial data based on the preset column name and row name rules through the data extraction module; S3, performing numerical calculation on the extracted target data set through the data calculation module, and processing the empty set data therein; S4, judging whether the calculation result has an error exceeding a preset threshold through the data self-diagnosis module, if yes, jump to step S5, otherwise jump to step S6; S5, through the error data tracing module, locate the coordinates of the problem data according to data classification and calculation path; S6, correcting the problematic data through a data correction module; S7, outputting the corrected data to the target table in a preset format through the data output module.
2. A method for processing auxiliary service settlement data according to claim 1, characterized in that: The data extraction module comprises: Column name locator, used to match column names in a table based on regular expressions; A row name association unit is used to associate row names with column names through a multidimensional index; The dynamic extraction unit generates a target data set according to the matching results.
3. A method for processing auxiliary service settlement data according to claim 1, characterized in that: The processing of the empty set data in step S3 includes: Replaces null values with preset default values or generates fill values based on an interpolation algorithm of neighboring data.
4. The method for processing auxiliary service settlement data according to claim 1, characterized in that: The error determination method of the data self-diagnosis module is: The threshold is set based on historical data statistics. If the calculated data deviation exceeds the threshold, it is judged as abnormal data.
5. The method for processing auxiliary service settlement data according to claim 1, characterized in that: The error data tracing module also performs: Build a tree structure of the data calculation path and record the input and output relationship of each step of the calculation; Through the reverse tracing algorithm, trace back from the abnormal data node to the original input node to locate the source of the error.
6. A method for processing auxiliary service settlement data according to claim 1, characterized in that: The data correction module comprises: The rule base stores preset correction rules and associated field constraints; The automatic correction unit performs batch replacement of problem data or recalculation of formulas based on the rule base.
7. A method for processing auxiliary service settlement data according to claim 1, characterized in that: The data output module supports: Multiple tables are linked for output, and related tables are automatically generated based on data classification.
8. A method for processing auxiliary service settlement data according to claim 7, characterized in that: The data output module also supports: The data formatting unit rounds the calculation result to two decimal places and adds a unit identifier.
9. A method for processing auxiliary service settlement data according to claim 1, characterized in that: The initial data was compiled from the following sources: The data on various assessments, assessment returns, compensation, compensation sharing, etc. in the auxiliary service transaction data of different power plants; the data extraction module labels and classifies the data from different sources and establishes a cross-table association index.
10. A computer program, characterized in that Used to execute a method for processing auxiliary service settlement data as described in claims 1 to 9; the program comprises: Data input interface, supports CSV, Excel and database direct import; Parallel computing engine, using MapReduce framework to accelerate the processing efficiency of data computing modules; The log tracking unit records the coordinates of error data, correction records and output paths.