Data integration and storage method and system for new energy grid-connected evaluation
Through modular data acquisition, preprocessing and integration processes, the problems of diversified data sources and low integration efficiency of new energy stations and power grid equipment are solved, efficient data integration and storage are achieved, and the evaluation accuracy and intelligence level of the power system are improved.
Patent Information
- Application Number
- CN202510180173.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-07-22
AI Technical Summary
Traditional power grid data acquisition methods lack coverage of real-time power generation data and network connection point information of new energy, resulting in low data integration efficiency, making it difficult to achieve efficient evaluation of the maximum access capacity of new energy and intelligent development of power systems.
Modular data acquisition, preprocessing, calibration and fusion processes are adopted to collect real-time data from multiple data sources, and data fusion and storage are carried out through the alignment and connection of device ID and network connection point information, combining power grid topological relationships and clustering algorithms to ensure data accuracy and consistency.
It has achieved efficient integration of distributed new energy data and power grid equipment operation data, improved data accuracy and real-timeness, provided strong support for power system planning and operation, and promoted the intelligent development of the power grid.
Smart Images

Figure CN120354061A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of power system data analysis, and particularly to a data integration and storage method and system for new energy grid connection assessment. Background Art
[0002] With the rapid development of new energy technology, its proportion in the power grid has been increasing year by year, and the operating characteristics of new energy power stations (such as wind power and photovoltaic) after grid connection have gradually become the focus of attention of the power system. However, the traditional power grid data acquisition method has insufficient coverage of important parameters such as real-time power generation data and grid connection point information of new energy, and it is difficult to take into account the complexity of large-scale new energy access in data processing, resulting in limitations in the accuracy and real-time performance of grid connection assessment.
[0003] The assessment of the maximum new energy access capacity involves real-time data of distributed energy sources such as wind power and photovoltaic and the operating constraints of grid equipment, and multi-dimensional and multi-type data need to be processed. In the actual operation of the power grid, problems such as multi-source data acquisition, unstable data quality, mismatched spatio-temporal scales, large data volume, and inconsistent formats are often faced. These problems make traditional data integration methods usually rely on a single processing mechanism, resulting in cumbersome and inefficient data cleaning, calibration, and fusion processes. Specifically, how to unify the spatio-temporal scales of different data sources, improve the stability of data quality, process the efficiency of massive data, and achieve format unification are the key problems that need to be solved urgently.
[0004] Therefore, researching an efficient data integration method can not only improve the accuracy of the maximum new energy access capacity assessment, but also promote the development of the power system in terms of intelligence and automation, and further improve the planning and operation management level of the power system. Summary of the Invention
[0005] In view of the above existing problems, the present invention is proposed.
[0006] Therefore, the present invention provides a data integration and storage method and system for new energy grid connection assessment, mainly solving two problems: ① For problems such as diverse data sources, mismatched spatio-temporal scales, inconsistent formats, and complex data structures of new energy power stations and grid equipment data, how to efficiently integrate data from different data sources to meet the real-time assessment requirements; ② For new energy data and grid equipment operation data that need to be processed and analyzed in multiple links, how to ensure the consistency and efficiency of data during transmission and storage to support the accurate assessment of new energy grid connection margin.
[0007] To solve the above technical problems, the present invention provides the following technical solutions:
[0008] In a first aspect, the present invention provides a data integration and storage method for new energy grid connection assessment, including: collecting real-time data and key information from multiple data sources, and classifying the collected data; preprocessing the classified data to establish a data synchronization and update mechanism; calibrating and consistency processing the preprocessed data, implementing data fusion through a connection algorithm, and verifying the fusion effect; aligning and connecting the data through device ID and grid connection point information, constructing a logical association based on the power grid topology relationship, and deeply fusing and storing the new energy output data and power grid operation data within the region in combination with a clustering algorithm to achieve data integration and dynamic assessment.
[0009] As a preferred solution of the data integration and storage method for new energy grid connection assessment according to the present invention, wherein: the collecting of real-time data and key information from multiple data sources includes:
[0010] Collecting the real-time power generation data of distributed new energy and the operation status of power grid equipment and power grid topology relationship information from the dispatching and production management system, on-site monitoring system, meteorological monitoring station, historical data management system, and new energy power station information management system; generating meteorological prediction information based on the real-time meteorological data of the meteorological monitoring station and the historical meteorological data of the historical data management system; generating load prediction information in combination with the dispatching and production management system and the historical data management system; generating power generation prediction information based on the new energy power station information management system and meteorological prediction data;
[0011] According to the data characteristics, the collected data is divided into static data and dynamic data;
[0012] Establish a standardized data collection mechanism, and realize the real-time access of multi-source data through a distributed collection module and a communication protocol.
[0013] As a preferred solution of the data integration and storage method for new energy grid connection assessment according to the present invention, wherein: the preprocessing of the classified data to establish a data synchronization and update mechanism includes:
[0014] For the data from new energy power stations, conventional units, and power grid equipment, use an ETL conversion tool to convert it into a unified XML data format, and at the same time perform data cleaning and loading operations to remove redundant or invalid information;
[0015] Adopt GPS time synchronization technology to ensure that each data source is based on a unified time reference. For the required real-time dynamic data, it is updated at the minute level, the required historical data is obtained in batches at the day level, and the static data is obtained once.
[0016] The collected data is subjected to noise elimination using an improved weighted median filtering method, and the missing data is complemented using a weighted linear interpolation method based on time weights.
[0017] As a preferred embodiment of the data integration and storage method for new energy grid connection assessment according to the present invention, wherein: the preprocessing of the classified data and the establishment of a data synchronization and update mechanism further include:
[0018] The noise elimination of the collected data using an improved weighted median filtering method is as follows:
[0019] According to the change rates of the real-time power generation data of new energy power stations and the operating status data of grid equipment, the size of the sliding window is dynamically adjusted to adapt to the data fluctuation characteristics;
[0020] The data points within the sliding window are weighted, and the weights are assigned according to time characteristics, and the weighted median is calculated;
[0021] The weighted median is used as the current data point to replace the abnormal points or noise values in the original data;
[0022] Output the real-time power generation data of new energy power stations and the operating status data of grid equipment after dynamic window adjustment and weighted median filtering;
[0023] The complementing of missing data using a weighted linear interpolation method based on time weights is as follows:
[0024] Detect the missing points in the real-time power generation data or load data of new energy power stations;
[0025] For the missing points, use a weighted linear interpolation method based on time weights to complement;
[0026] Output the complemented data sequence for subsequent analysis and processing.
[0027] As a preferred embodiment of the data integration and storage method for new energy grid connection assessment according to the present invention, wherein: the calibration and consistency processing of the preprocessed data include:
[0028] Specific data calibration rules are formulated for different data types, and the data calibration rules are as follows:
[0029] Use the percentage of the rated capacity as the calibration benchmark. For the real-time power generation data, calibrate according to the percentage of the rated capacity of the equipment; the operating status of the equipment is calibrated by status codes, and the specific statuses of "normal", "fault", and "maintenance" are defined;
[0030] Set reasonable calibration ranges for various types of data to ensure that the power generation data and load data are within the physical range, and subsequent data beyond the reasonable range is regarded as abnormal and recorded for processing;
[0031] Design a calibration formula according to the data characteristics;
[0032] By setting the update frequency and validity verification rules of real-time data, regularly checking the device capacity data, marking and correcting abnormal data, and ensuring the consistency between dynamic data and static data, the standardization and accuracy of data processing are achieved. The consistency processing rules are as follows:
[0033] Set the update frequency of real-time power generation data; set validity verification rules, mark abnormal data by detecting whether the real-time power generation data exceeds the preset range, and record and process it;
[0034] Regularly check the device capacity data to ensure that it conforms to the actual operating state of the device; for the device data with capacity changes, update it to the database in a timely manner and verify its accuracy;
[0035] Mark the data exceeding the preset range as abnormal data and trigger the verification program.
[0036] As a preferred solution of the data integration and storage method for new energy grid connection assessment described in the present invention, wherein: the alignment and connection of data through the device ID and connection point information, and the construction of logical associations based on the power grid topology relationship include:
[0037] Use the device ID and connection point information as the primary key to align and deeply connect the data of new energy power stations and power grid devices;
[0038] Adopt the shortest path algorithm to construct the logical associations between data based on the power grid topology relationship, model the main transformers, key lines, partitions, and load node elements in the power grid to form the topology diagram of the power grid, and associate the corresponding power grid device data with the device ID according to the power grid topology relationship.
[0039] As a preferred solution of the data integration and storage method for new energy grid connection assessment described in the present invention, wherein: the deep fusion and storage of new energy output data and power grid operation data in the region by combining the clustering algorithm include:
[0040] Through the K-means clustering algorithm, fuse the distributed new energy output data and power grid operation data;
[0041] After data fusion, check the consistency of different data sources in the time dimension, verify whether the timestamps of the data of each device are consistent, and ensure the synchronization of multi-source data; by comparing with historical data or standard data, ensure the accuracy of the fused data; if data inconsistency is found, linearly interpolate and correct the time for the mismatched data according to the trend of adjacent time points. When there are multiple data sources, preferentially select the data source with higher historical accuracy, mark the conflict data that cannot be resolved, and prompt the user for further processing;
[0042] Store the integrated data in a unified storage structure in a mode of first stratifying by region and then by time within the region, and establish an indexing mechanism for the stored data to support fast query and data access.
[0043] In a second aspect, the present invention provides a data integration and storage system for new energy grid connection assessment, including:
[0044] A data collection and access module, configured to collect real-time data and key information from multiple data sources and classify the collected data;
[0045] A data preprocessing module, configured to preprocess the classified data and establish a data synchronization and update mechanism;
[0046] A data calibration and integration module, configured to calibrate and perform consistency processing on the preprocessed data, implement data fusion through a connection algorithm, and verify the fusion effect;
[0047] A data connection and storage module, configured to align and connect data through device IDs and grid connection point information, construct logical associations based on the power grid topology relationship, and deeply fuse and store the new energy output data and power grid operation data within the region in combination with a clustering algorithm to achieve data integration and dynamic assessment.
[0048] In a third aspect, the present invention provides an electronic device, including:
[0049] A memory and a processor;
[0050] The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the data integration and storage method for new energy grid connection assessment are implemented.
[0051] In a fourth aspect, the present invention provides a computer-readable storage medium, which stores computer-executable instructions. When the computer-executable instructions are executed by a processor, the steps of the data integration and storage method for new energy grid connection assessment are implemented.
[0052] Compared with the prior art, the beneficial effects of the present invention are as follows: The present invention provides a data integration and storage method and system for new energy grid connection assessment, aiming to solve the problems of diverse data sources, low integration efficiency, and difficulty in efficiently supporting grid connection assessment for new energy power stations and grid equipment. Through the method provided by the present invention, a modular data acquisition, preprocessing, calibration, and fusion process is adopted to achieve the efficient integration of distributed new energy data and grid equipment operation data, ensuring the accuracy, consistency, and real-time nature of the data, and being able to provide strong data support for the planning and operation of the power system, promoting the intelligent development of the power grid. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.
[0054] Figure 1 It is a schematic diagram of the overall process logic of the data integration and storage method for new energy grid connection assessment according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0055] In order to make the above objects, features, and advantages of the present invention more obvious and understandable, the following will provide a detailed description of the specific embodiments of the present invention with reference to the drawings of the specification. Obviously, the described embodiments are some embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0056] Embodiment 1
[0057] Refer to Figure 1 In an embodiment of the present invention, a data integration and storage method for new energy grid connection assessment is provided, aiming to solve the problems of diverse data sources, low integration efficiency, and inability to efficiently support new energy grid connection assessment for new energy power stations and grid equipment. This method is a systematic solution that realizes the efficient integration and management of multi-source heterogeneous data through data acquisition, data preprocessing, data calibration, data connection fusion, and storage, providing accurate and real-time data support for the assessment of the maximum access capacity of new energy. As Figure 1 shown, it specifically includes the following steps:
[0058] S100: Collect real-time data and key information from multiple data sources and classify the collected data;
[0059] S200: Preprocess the classified data and establish a data synchronization and update mechanism;
[0060] S300: Calibrate and perform consistency processing on the preprocessed data, implement data fusion through a connection algorithm, and verify the fusion effect;
[0061] S400: Align and connect the data through device IDs and connection point information, construct logical associations based on the power grid topology relationship, and deeply fuse and store the new energy output data and power grid operation data within the region in combination with a clustering algorithm to achieve data integration and dynamic assessment.
[0062] It should be noted that the present invention provides a method and system for data integration and storage for new energy grid connection assessment, aiming to solve the problems of diverse data sources, low integration efficiency, and difficulty in efficiently supporting grid connection assessment for new energy power stations and power grid equipment. Through the method provided by the present invention, a modular data acquisition, preprocessing, calibration, and fusion process is adopted to achieve the efficient integration of distributed new energy data and power grid equipment operation data, ensuring the accuracy, consistency, and real-time nature of the data, and being able to provide strong data support for the planning and operation of the power system, promoting the intelligent development of the power grid.
[0063] In the embodiment of the present application, the above step S100 includes the following sub-steps A1 to A3:
[0064] In A1: Collect the real-time power generation data of distributed new energy (such as wind power, photovoltaic) and the operation status of power grid equipment (such as main transformers, lines), and power grid topology relationship information from the dispatching and production management system, on-site monitoring system, meteorological monitoring station, historical data management system, and new energy power station information management system; Generate meteorological prediction information based on the real-time meteorological data (such as wind speed, wind direction, light intensity, temperature, humidity, etc.) of the meteorological monitoring station and the historical meteorological data of the historical data management system; Generate load prediction information in combination with the dispatching and production management system and the historical data management system; Generate power generation prediction information based on the new energy power station information management system and the meteorological prediction data;
[0065] In A2: Classify the collected data into static data and dynamic data according to the data characteristics;
[0066] In A3: Establish a standardized data acquisition mechanism to achieve the real-time access of multi-source data through distributed acquisition modules and communication protocols.
[0067] In an optional embodiment, the operation status of the power grid equipment specifically includes:
[0068] Operating status of the main transformer: including real-time load rate (i.e., the ratio of current load to rated capacity), winding temperature, voltage status (primary and secondary side voltage values), and operating mode (such as parallel operation, standby status);
[0069] Operating status of the line: including real-time active and reactive power of the line, line current magnitude, line temperature, and commissioning and fault status of the line (such as open circuit, short circuit, tripping);
[0070] Operating status of circuit breakers and switches: whether the line is in a conducting state, whether there is overcurrent, short circuit, etc.;
[0071] The above operating statuses need to be within a reasonable range to ensure safe operation.
[0072] In an optional embodiment, the static data and dynamic data are respectively:
[0073] Static data includes: basic parameters of grid equipment (such as main transformer capacity, line rated capacity, equipment rated current, grid topology), equipment operation limit conditions (such as voltage limits, frequency range, safety margin threshold), layout information of new energy power stations (such as grid connection point location, power station rated capacity), historical statistical characteristics (such as annual average power generation, equipment life assessment results), etc.;
[0074] Dynamic data includes: real-time generation data of distributed new energy (such as wind power, photovoltaic) (such as active power, reactive power, generation current and voltage), real-time operating status of grid equipment (such as main transformer load rate, line current, equipment temperature), real-time meteorological information (such as wind speed, light intensity, air temperature), prediction data (such as meteorological prediction information, load prediction information, generation prediction information), etc.
[0075] In an optional embodiment, the data acquisition mechanism is:
[0076] Adopt the HTTP protocol, and through the API interfaces of the dispatching and production management system and the on-site monitoring system, collect information such as grid topology, load, and new energy generation in real time; through the Modbus protocol, use meteorological monitoring equipment to obtain real-time meteorological data such as wind speed and light intensity; directly extract historical data of load and generation from the historical data management system;
[0077] Data acquisition uses an automated data acquisition module, and combines on-site equipment (sensors, intelligent terminals, etc.) and remote server interfaces to ensure real-time data transmission.
[0078] It should be noted that the above step S100 collects real-time data and key information from multiple data sources and classifies them to ensure the comprehensiveness and accuracy of the data, providing a high-quality data foundation for subsequent processing, which is conducive to improving the efficiency and accuracy of the entire data processing process.
[0079] In the embodiment of the present application, the above step S200 includes the following sub-steps B1 to B3:
[0080] In B1: For the data from new energy power stations, conventional units, and grid equipment, use an ETL (Extract, Transform, Load) conversion tool to convert it into a unified XML data format for interoperability and compatibility during the data processing process. At the same time, perform data cleaning and loading operations to remove redundant or invalid information and ensure data integrity and consistency;
[0081] In B2: Adopt GPS time synchronization technology to ensure that each data source is based on a unified time reference. Update the required real-time dynamic data at the minute level, batch obtain the required historical data at the daily level, and obtain the static data once;
[0082] In B3: Use an improved weighted median filtering method to eliminate noise from the collected data, and use a weighted linear interpolation method based on time weights to fill in the missing data.
[0083] Specifically, using an improved weighted median filtering method to eliminate noise from the collected data includes:
[0084] Dynamically adjust the size of the sliding window according to the change rate of the real-time power generation data of the new energy power station and the operation status data of the grid equipment. When the change rate is large, reduce the window; when the change rate is small, increase the window to adapt to the data fluctuation characteristics;
[0085] Define the calculation formula of the data change rate as:
[0086]
[0087] where, x t and x t-1 represent the data values at the current and previous moments, Δt is the time interval, and adjust the window length W according to the size of R t :
[0088]
[0089] where, W min and W max are the minimum and maximum window lengths respectively, and α and β are adjustment parameters.
[0090] Assign weights ω i to the data points within the sliding window:
[0091]
[0092] Among them, t i is the time of the i-th data point, and t c is the time at the center point of the current window, and σ is to control the weight range. The weights are assigned according to the time characteristics, and the weighted median is calculated, (ω1·x1, ω2·x2....., ω n ·x n ) Take the weighted median as the current data point x c , and replace the abnormal points or noise values in the original data:
[0093] Output the real-time power generation data of the new energy power station and the operation status data of the power grid equipment after dynamic window adjustment and weighted median filtering;
[0094] Specifically, the method of using weighted linear interpolation based on time weights to complete the missing data includes:
[0095] Detect the missing points in the real-time power generation data or load data of the new energy power station;
[0096] For the missing points, use the weighted linear interpolation method based on time weights to complete them, and the calculation formula is:
[0097]
[0098]
[0099] Among them, ω1 and ω2 represent weights;
[0100] Output the completed data sequence for subsequent analysis and processing.
[0101] It should be noted that the above step S200 preprocesses the classified data, establishes a data synchronization and update mechanism, ensures the timeliness and consistency of the data, provides stable and reliable data support for subsequent data fusion and processing, and improves the efficiency and accuracy of the overall data processing process.
[0102] In the embodiment of the present application, the above step S300 includes the following sub-steps C1 to C5:
[0103] In C1: For different data types, including wind farm data, photovoltaic power station data, traditional power grid equipment data, etc., specific data calibration rules are formulated respectively;
[0104] Specifically, the data calibration rules include:
[0105] Use the percentage of the rated capacity as the calibration benchmark, and calibrate the real-time power generation data according to the percentage of the rated capacity of the equipment; design a calibration formula based on the data characteristics. For example, the real-time power generation is calculated as "calibration value = (real-time power generation / rated capacity) * 100%".
[0106] The operating status of the equipment is calibrated by status codes, defining "normal" (status code: 1), "fault" (status code: 2), "maintenance" (status code: 3), "out of service" (status code: 4), "early warning" (status code: 5); the equipment status is updated in real time through the on-site monitoring system. If the operating status of the equipment changes (such as a fault, maintenance, or outage), the system automatically updates the corresponding status code.
[0107] Set reasonable calibration ranges for various types of data to ensure that the power generation data and load data are within the physical range. Subsequently, data beyond the reasonable range is regarded as abnormal and recorded for processing.
[0108] Exemplarily, the rated capacity of wind power (WA);
[0109] Calibration standard: Set the rated capacity of the wind farm as 100%, and the actual output is expressed as a percentage of the rated capacity.
[0110] Calibration range: 0% - 100%.
[0111] Wind power real-time power generation data (WB):
[0112] Calibration standard: Calibrate the real-time power generation data according to the percentage of the rated capacity of the wind farm.
[0113] Calibration formula: Calibration value = (real-time power generation / rated capacity) * 100%.
[0114] Calibration range: 0% - 100%.
[0115] In C2: Standardize and ensure the accuracy of data processing by setting the update frequency and validity verification rules of real-time data, regularly checking the equipment capacity data, marking and correcting abnormal data, and ensuring the consistency between dynamic data and static data;
[0116] Specifically, the consistency processing rules are as follows:
[0117] Set the update frequency of the real-time power generation data; set the validity verification rules, mark and process abnormal data by detecting whether the real-time power generation data exceeds the preset range;
[0118] Regularly check the equipment capacity data to ensure its consistency with the actual operating status of the equipment; for the equipment data with capacity changes, update it to the database in a timely manner and verify its accuracy;
[0119] Mark the data beyond the preset range as abnormal data and trigger the verification program.
[0120] In C3: Select the method based on device ID to determine the correlation relationship of data from different sources;
[0121] In C4: Use a stream processing framework (such as Apache Kafka, Flink) to implement the real-time processing and efficient fusion of the associated data in the above steps, and regularly update the fused data to the target database to ensure the integrity and real-time nature of the data;
[0122] In C5: Conduct data quality inspection on the fused data, verify the consistency between the fused data and the original data, and use statistical methods to evaluate and make necessary adjustments to the accuracy of data fusion to ensure the reliability and consistency of the fused data.
[0123] It should be noted that the above step S300 ensures the accuracy and reliability of the data, effectively improves the data quality, and lays a solid foundation for the subsequent in-depth integration and dynamic evaluation of data based on the power grid topology relationship.
[0124] In the embodiment of the present application, the above step S400 includes the following sub-steps D1 to D5:
[0125] In D1: Use the device ID and the grid connection point information as the primary key to align and deeply connect the data of the new energy power station and the power grid equipment;
[0126] In D2: Adopt the shortest path algorithm to construct the logical association between data based on the power grid topology relationship (including main transformer, line information, etc.);
[0127] In D3: Through the K-means algorithm, fuse the distributed new energy output data and the power grid operation data;
[0128] In D4: After data fusion, establish a set of data consistency verification methods to verify the fusion result;
[0129] In D5: Store the integrated data in a unified storage structure in a mode of first stratifying by region and then stratifying by time within the region, and establish an efficient indexing mechanism for the stored data to support fast query and data access.
[0130] Specifically, to ensure data consistency between different data sources, first, the device IDs and connection point information of each new energy power station and grid equipment are used as primary keys for data alignment. Through standardized device identifiers and connection point information, data from different sources (such as wind farms, photovoltaic farms, main transformers, lines, etc.) are uniformly associated, providing a basis for subsequent data fusion. The data sources include real-time power generation data of new energy power stations (such as wind power and photovoltaic power) and operation data of grid equipment (such as main transformers, lines, voltage, current, etc.).
[0131] Specifically, the shortest path algorithm is adopted to construct the logical association between data based on the grid topology relationship. The main transformers, key lines, partitions, and load node elements in the grid are modeled to form the topology map of the grid. According to the grid topology relationship, the corresponding grid equipment data is associated with the device ID to realize the organic combination of data in each part of the grid system.
[0132] Specifically, the K-means clustering algorithm is used to fuse the distributed new energy output data and grid operation data, including:
[0133] According to the geographical location and characteristics of the devices (such as the wind speed characteristics of wind farms and the light intensity of photovoltaic farms), as well as the operating load characteristics of the grid, multiple regions are divided. Each region represents a cluster of one or more devices. During the clustering process, the output data of the devices and the grid operation data will be divided by region to ensure that the data within the region is similar in terms of geography and function.
[0134] Within each region, the K-means algorithm is used to deeply fuse the output data of various devices and the grid operation data to generate region-level data features. This process will enable the data characteristics within each region to be fully integrated within the same cluster, improving the relevance of the data and the accuracy of the analysis.
[0135] After clustering, each region data set obtained includes comprehensive data such as wind power output, photovoltaic power generation, device operation status, and grid load.
[0136] Specifically, after data fusion, it is necessary to ensure that the fusion results of each data source are consistent and there is no information loss or data conflict. For this reason, a set of data consistency verification methods is established to verify the integrity, accuracy, and consistency of the fusion results. The consistency verification methods are as follows:
[0137] Check the consistency of different data sources in the time dimension, verify whether the timestamps of each device data are consistent, and ensure the synchronization of multi-source data.
[0138] Compare with historical data or standard data to ensure that the fused data is accurate and error-free.
[0139] Use a rule engine to detect possible data conflicts, such as contradictions between device data (e.g., the real-time power generation of a wind farm does not match the grid load), and correct them in a timely manner.
[0140] If data inconsistencies are found, they are processed by the following methods:
[0141] a. For mismatched data, linearly interpolate and correct the time based on the trends at adjacent time points.
[0142] b. When there are multiple data sources, preferentially select the data source with a higher historical accuracy rate.
[0143] c. Mark the conflict data that cannot be resolved and prompt the user for further processing.
[0144] It should be noted that the above step S400 aligns and connects data through the device ID and the connection point information, constructs a logical association based on the power grid topology relationship, and uses a clustering algorithm to achieve deep integration of new energy output data and power grid operation data, which not only improves the accuracy and efficiency of data integration, but also supports the dynamic assessment and optimal operation of the power grid within the region.
[0145] Embodiment 2
[0146] In this embodiment, a data integration and storage system for new energy grid connection assessment is provided, including:
[0147] A data collection and access module, which is used to collect real-time data and key information from multiple data sources and classify the collected data;
[0148] A data preprocessing module, which is used to preprocess the classified data and establish a data synchronization and update mechanism;
[0149] A data calibration and integration module, which is used to calibrate and perform consistency processing on the preprocessed data, implement data fusion through a connection algorithm, and verify the fusion effect;
[0150] A data connection and storage module, which is used to align and connect data through the device ID and the connection point information, construct a logical association based on the power grid topology relationship, and perform deep integration and storage of new energy output data and power grid operation data within the region in combination with a clustering algorithm to achieve data integration and dynamic assessment.
[0151] It should be noted that the technical solution of the system for data integration and storage for new energy grid connection assessment belongs to the same concept as the technical solution of the above-mentioned method for data integration and storage for new energy grid connection assessment. For the details not described in detail in the technical solution of the system for data integration and storage for new energy grid connection assessment in this embodiment, reference can be made to the description of the technical solution of the above-mentioned method for data integration and storage for new energy grid connection assessment.
[0152] The above-mentioned unit modules can be embedded in the processor in the computer device in hardware form or be independent of the processor, or can be stored in the memory in the computer device in software form, so as to facilitate the processor to call and execute the operations corresponding to the above modules.
[0153] This embodiment also provides an electronic device, which includes a processor, a memory, a communication interface, a display screen and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a carrier network, NFC (Near Field Communication) or other technologies. When the computer program is executed by the processor, it realizes a method for data integration and storage for new energy grid connection assessment. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covered on the display screen, or a button, a trackball or a touchpad set on the shell of the computer device, or an external keyboard, a touchpad or a mouse, etc.
[0154] This embodiment also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by the processor, it realizes the method proposed in the above embodiment.
[0155] The storage medium proposed in this embodiment and the method proposed in the above embodiment belong to the same inventive concept. The technical details not described in detail in this embodiment can be referred to the above embodiment, and this embodiment has the same beneficial effects as the above embodiment.
[0156] From the above description of the embodiments, those skilled in the art can clearly understand that the present invention can be implemented by means of software and necessary general-purpose hardware. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation. Based on this understanding, the technical solution of the present invention, in essence, or the part that makes a contribution to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as a floppy disk, read-only memory (ROM), random access memory (RAM), flash memory, hard disk, or optical disc of a computer, etc., including several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute the method of the embodiments of the present invention.
[0157] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not restrictive. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered by the scope of the claims of the present invention.
[0158] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, system, or computer program product. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The solutions in the embodiments of the present application can be implemented in various computer languages.
[0159] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the flows and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0160] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to operate in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including instruction means that implement the function specified in one or more of the procedures Figure 1 or blocks Figure 1 specified in one or more of the procedures and / or blocks.
[0161] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the function specified in one or more of the procedures Figure 1 or blocks Figure 1 specified in one or more of the procedures and / or blocks.
[0162] Although the preferred embodiments of the present application have been described, additional changes and modifications can be made to these embodiments by those skilled in the art once they learn of the basic inventive concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications that fall within the scope of the present application.
[0163] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these modifications and variations.
Claims
1. A data integration and storage method for new energy grid connection assessment, characterized in that, Including: Collecting real-time data and key information from multiple data sources and classifying the collected data; Preprocessing the classified data and establishing a data synchronization and update mechanism; Calibrating and processing the consistency of the preprocessed data, implementing data fusion through a connection algorithm, and verifying the fusion effect; Aligning and connecting the data through device ID and grid connection point information, constructing logical associations based on the power grid topology relationship, and deeply fusing and storing the new energy output data and power grid operation data within the region in combination with a clustering algorithm to achieve data integration and dynamic assessment.
2. The data integration and storage method for new energy grid connection assessment according to claim 1, wherein The collecting real-time data and key information from multiple data sources includes: Collecting the real-time power generation data of distributed new energy and the operation status of power grid equipment and power grid topology relationship information from the dispatching and production management system, on-site monitoring system, meteorological monitoring station, historical data management system, and new energy power station information management system; generating meteorological prediction information based on the real-time meteorological data of the meteorological monitoring station and the historical meteorological data of the historical data management system; generating load prediction information in combination with the dispatching and production management system and the historical data management system; generating power generation prediction information based on the new energy power station information management system and meteorological prediction data; Classifying the collected data into static data and dynamic data according to data characteristics; Establishing a standardized data collection mechanism to achieve real-time access of multi-source data through distributed collection modules and communication protocols.
3. The data integration and storage method for new energy grid connection assessment according to claim 2, characterized in that Preprocessing the classified data and establishing a data synchronization and update mechanism includes: For the data from new energy power stations, conventional units, and power grid equipment, using an ETL conversion tool to convert it into a unified XML data format, and at the same time performing data cleaning and loading operations to remove redundant or invalid information; Adopting GPS time synchronization technology to ensure that each data source is based on a unified time reference, updating the required real-time dynamic data at the minute level, batch obtaining the required historical data at the day level, and obtaining static data once; Eliminating noise from the collected data using an improved weighted median filtering method, and complementing missing data using a weighted linear interpolation method based on time weights.
4. The data integration and storage method for new energy grid connection assessment according to claim 3, wherein Preprocessing the classified data and establishing a data synchronization and update mechanism also includes: The eliminating noise from the collected data using an improved weighted median filtering method is: Dynamically adjusting the size of the sliding window according to the change rate of the real-time power generation data of new energy power stations and the operation status data of power grid equipment to adapt to the data fluctuation characteristics; Performing weighted processing on the data points within the sliding window, distributing the weights according to time characteristics, and calculating the weighted median; Taking the weighted median as the current data point to replace the abnormal points or noise values in the original data; Outputting the real-time power generation data of new energy power stations and the operation status data of power grid equipment after dynamic window adjustment and weighted median filtering processing; The complementing missing data using a weighted linear interpolation method based on time weights is: Detecting the missing points in the real-time power generation data or load data of new energy power stations; Complementing the missing points using a weighted linear interpolation method based on time weights; Output the completed data sequence for subsequent analysis and processing.
5. The data integration and storage method for new energy grid connection assessment according to claim 4, characterized in that The calibration and consistency processing of the preprocessed data includes: Specific data calibration rules are formulated for different data types. The data calibration rules are as follows: Use the percentage of the rated capacity as the calibration benchmark. Calibrate the real-time power generation data according to the percentage of the rated capacity of the equipment. The operating status of the equipment is calibrated by status codes, and the specific statuses of "normal", "fault", and "maintenance" are defined. Set reasonable calibration ranges for various types of data to ensure that the power generation data and load data are within the physical range. Subsequently, data outside the reasonable range is regarded as abnormal and recorded for processing. Design calibration formulas according to the data characteristics. By setting the update frequency and validity verification rules of real-time data, regularly checking the equipment capacity data, marking and correcting abnormal data, and ensuring the consistency between dynamic data and static data, the standardization and accuracy of data processing are achieved. The consistency processing rules are as follows: Set the update frequency of real-time power generation data; set validity verification rules. By detecting whether the real-time power generation data exceeds the preset range, mark abnormal data and record and process it. Regularly check the equipment capacity data to ensure its consistency with the actual operating status of the equipment; for the equipment data with capacity changes, update it to the database in a timely manner and verify its accuracy. Mark the data outside the preset range as abnormal data and trigger the verification program.
6. The data integration and storage method for new energy grid connection assessment according to claim 5, characterized in that The alignment and connection of data through the equipment ID and the connection point information, and the construction of logical associations based on the power grid topology include: Use the equipment ID and the connection point information as the primary key to align and deeply connect the data of new energy power stations and power grid equipment. Adopt the shortest path algorithm to construct the logical associations between data based on the power grid topology. Model the main transformers, key lines, partitions, and load node elements in the power grid to form the topology diagram of the power grid, and associate the corresponding power grid equipment data with the equipment ID according to the power grid topology.
7. The data integration and storage method for new energy grid connection assessment according to claim 6, characterized in that, The deep fusion and storage of the new energy output data and the power grid operation data in the region by combining the clustering algorithm include: Fuse the distributed new energy output data and the power grid operation data through the K-means clustering algorithm. After data fusion, check the consistency of different data sources in the time dimension, verify whether the timestamps of each equipment data are consistent, and ensure the synchronization of multi-source data; by comparing with historical data or standard data, ensure that the fused data is accurate. If data inconsistency is found, linearly interpolate and correct the time for the mismatched data according to the trend of adjacent time points. When there are multiple data sources, preferentially select the data source with higher historical accuracy, mark the conflict data that cannot be resolved, and prompt the user for further processing. Store the integrated data in a unified storage structure in a mode of first stratifying by region and then stratifying by time within the region, and establish an indexing mechanism for the stored data to support fast query and data access.
8. A system applying the data integration and storage method for new energy grid connection assessment as described in any one of claims 1 to 7, characterized in that Include: The data acquisition and access module is used to collect real-time data and key information from multiple data sources and classify the collected data. A data preprocessing module, which is used to preprocess the classified data and establish a data synchronization and update mechanism; A data calibration and integration module, which is used to calibrate and perform consistency processing on the preprocessed data, implement data fusion through a connection algorithm, and verify the fusion effect; A data connection and storage module, which is used to align and connect data through device IDs and connection point information, construct logical associations based on the power grid topology, and deeply fuse and store the new energy output data and power grid operation data in the region in combination with a clustering algorithm to achieve data integration and dynamic assessment.
9. An electronic device, comprising a memory and a processor, characterized in that: The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the method described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having computer-executable instructions stored thereon, characterized in that: When the computer-executable instructions are executed by the processor, the steps of the method described in any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Photovoltaic power station operation data statistical method and system
CN121880422A
A method and system for statistical analysis of photovoltaic power plant operation data
CN121880422B
New energy station simulation modeling data management system
CN121997703A