A method and system for collecting and storing SCADA measurement data based on stream computing
By employing a stream computing-based approach, efficient acquisition and storage of SCADA measurement data were achieved, solving the problems of large data volume, diverse data types, and low value density in power systems. This approach enables redundant storage of multi-source data and online data quality monitoring, supporting real-time decision-making in power systems.
Patent Information
- Application Number
- CN202110412906.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-04-16
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2041-04-16
AI Technical Summary
The power system generates massive amounts of data of various types with low value density and requires rapid analysis and processing. Existing technologies struggle to efficiently collect and store SCADA measurement data, particularly in the areas of redundancy removal and data quality monitoring under multi-source data.
A stream computing-based approach is adopted to collect SCADA measurement data at preset intervals through an online monitoring data acquisition device, convert it into cross-sectional data files and save them, store them in a distributed file system using a stream processing service, and establish an index table for data monitoring, thereby achieving redundant storage of multi-source data and online data quality monitoring.
It enables efficient acquisition and storage of SCADA measurement data, meets the performance requirements of online monitoring and real-time data processing of power equipment, ensures data security and integrity, and supports diversified decision-making in the power system.
Smart Images

Figure CN113204534B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of power system simulation and stability analysis, and more particularly, to a SCADA measurement data collection and storage method and system based on stream computing. BACKGROUND
[0002] In recent years, with the development of China's power system towards high informationization and automation, the collection and use of power data are becoming more and more widespread, and higher requirements are put forward for the management of power equipment and facility data, user data, planning data, etc.
[0003] According to the different sources of data, smart grid data can be divided into two categories: one is the internal data of the power grid, mainly from power information collection system, distribution management system, equipment detection and monitoring system, etc.; the other is external data, mainly from geographic information system, public service department, etc.; according to the different data content and collection time nodes, smart grid data can also be divided into static data and stream data: static data refers to basic data, including planning design data (such as power equipment location, loop, construction time, manufacturer, etc.), power grid resource data (such as generator excitation system model, station primary wiring diagram and related information); the other is state data, that is, stream data, including equipment operation log data, monitoring data, etc. As for the data in the power grid, there are also the characteristics of big data: ① huge data volume: with the rapid construction of power enterprise informationization and the completion of smart power system, the growth rate of power data will far exceed the expectation of power enterprises, according to statistics, the basic data and equipment state operation online monitoring data in each link of the power system have jumped from TB level (Terabyte, trillion bytes) to PB level (Petabyte, thousand trillion bytes); ② multiple data types: power grid data are widely distributed and have many types, including real-time data, historical data, text data, multimedia data, etc., and the frequency and performance requirements of query and processing of various data are different; ③ low value density: in the state detection of transmission and transformation equipment, most of the collected data are normal data, only a small amount of abnormal data, and abnormal data is an important basis for condition-based maintenance; ④ fast analysis and processing speed: the processing performance requirement of online state data is much higher than that of offline data, and a large amount of data needs to be compared and processed in a very short time to support decision making.
[0004] The power generation, transformation, transmission and consumption system are complex systems containing a large amount of information, and through data collection, structured processing, correlation and fusion processing, the relevant information can be integrated to the maximum extent, thereby providing power system decision makers with a diversified, panoramic and full operation track basic data reflecting the power grid, as a basis for decision making.
[0005] The smart grid dispatching control system (D5000 system) realizes the horizontal integration and vertical penetration of the power grid dispatching business, the model sharing and integration based on the CIM / E, CIM / G standards and model splicing technology, the vertical realization of the coordinated control of the national, regional and provincial dispatching businesses, and the support of the whole-network sharing of the real-time data, real-time pictures and application functions. Therefore, it is particularly important to improve the information processing and intelligent decision-making ability of the power system by collecting the online data of the D5000 system, constructing an integrated data center and a computing analysis and decision-making platform suitable for massive data, integrating the power grid data resources by using the data architecture technology such as the data warehouse, analyzing the information and mining the potential value of the data resources. SUMMARY
[0006] In view of the above problems, the application provides a SCADA measurement data acquisition and storage method based on stream computing, which comprises the following steps:
[0007] An online monitoring data acquisition device is deployed in a target area of a power system, SCADA measurement data are acquired at a preset time interval, the SCADA measurement data are converted into section data files, and the section data files are saved;
[0008] The saved section data files are collected to a local device according to stream computing, the section data files saved to the local device are saved to a distributed file system in a preset storage format through a stream processing service;
[0009] The acquisition condition and the quality of the data of the section data files stored in the distributed file system are monitored.
[0010] Optionally, the preset time interval is 5-15 minutes.
[0011] Optionally, the preset storage format is a structure of the acquisition time and the acquisition place.
[0012] Optionally, after the section data files are saved to the distributed file system, an index table is established, acquisition information is established by using the index table, and the acquisition information comprises the acquisition time, the saving path and file information.
[0013] The application further provides a SCADA measurement data acquisition and storage system based on stream computing, which comprises the following steps:
[0014] A data acquisition unit is configured to deploy an online monitoring data acquisition device in a target area of a power system, acquire SCADA measurement data at a preset time interval, convert the SCADA measurement data into section data files, and save the section data files;
[0015] The distributed storage unit saves the saved cross-section data file according to the stream calculation collection to the local, saves the cross-section data file saved to the local to the distributed file system through the stream processing service, and saves to the distributed file system in a preset storage format.
[0016] The data monitoring unit monitors the data collection situation and the data quality of the cross-section data file stored in the distributed file system.
[0017] Optionally, the preset time interval is 5-15 minutes.
[0018] Optionally, the preset storage format is the structure of the collection time and the collection place.
[0019] Optionally, after the cross-section data file is saved to the distributed file system, an index table is established, and collection information is established by using the index table, the collection information including the collection time, the saving path and the file information.
[0020] The present application realizes the SCADA measurement data collection in the stream calculation mode, realizes the de-redundancy storage under the multi-source data, saves the measurement time stamp to the PSDB distributed file system, and monitors the running situation of the collection service and the collection online data quality through the online monitoring service. BRIEF DESCRIPTION OF DRAWINGS
[0021] Figure 1 The flow chart of the method of the present application;
[0022] Figure 2 The flow chart of the D5000 online measurement data stream processing and the related application of the present application;
[0023] Figure 3 The physical architecture diagram of the D5000 online measurement data stream processing of the present application;
[0024] Figure 4 The distributed file server storage logic structure of the D5000 online measurement data of the present application;
[0025] Figure 5 The structure diagram of the system of the present application. DETAILED DESCRIPTION
[0026] Exemplary embodiments of the present application will now be described with reference to the accompanying drawings, however, the present application can be implemented in many different forms and is not limited to the embodiments described herein, and these embodiments are provided to thoroughly and completely disclose the present application and to fully convey the scope of the present application to those skilled in the art. The terms used in the exemplary embodiments represented in the accompanying drawings are not limitations of the present application. In the accompanying drawings, the same elements / elements are denoted by the same reference numerals.
[0027] The terms (including scientific and technical terms) used herein, unless otherwise defined, have the ordinary meaning understood by one of ordinary skill in the art. In addition, it is to be understood that terms defined in a commonly used dictionary should be interpreted as having a meaning that is consistent with their meaning in the context of the relevant art and not be interpreted in an idealized or overly formal sense unless expressly so defined herein.
[0028] The present application provides a SCADA measurement data acquisition and storage method based on stream computing, as shown in the formula (I): Figure 1 The present application provides a SCADA measurement data acquisition and storage method based on stream computing, as shown in the formula (I):
[0029] The online monitoring data acquisition device is deployed in the target area of the power system, and the SCADA measurement data is collected at a preset time interval, and the SCADA measurement data is converted into a cross-section data file, and the cross-section data file is saved;
[0030] The saved cross-section data file is collected according to the stream computing to the local, and the cross-section data file saved to the local is saved to the distributed file system in a preset storage format through the stream processing service;
[0031] The cross-section data file stored in the distributed file system is monitored for data collection and data quality.
[0032] The preset time interval is 5-15 minutes.
[0033] The preset storage format is the structure of the collection time and the collection place.
[0034] After the cross-section data file is saved to the distributed file system, an index table is established, and the collection information is established by the index table, the collection information including the collection time, the saving path and the file information.
[0035] The present application will be further described in conjunction with the embodiments:
[0036] The stream computing framework is mainly used for stream data processing, and at present, it is mainly based on open source Apache Storm and Spark Streaming. Unlike Hadoop MapReduce and Spark batch computing framework, it has the characteristics of event triggering and short response time, and the event triggering and response time can reach s level, even ms level.
[0037] The application considers the service characteristics of D5000 online data acquisition, independently develops the functions of the stream computing acquisition part, is deployed in the III area, monitors the data on the multi-source monitoring server through the SFTP service, and realizes the quasi-real-time data acquisition and storage. The overall architecture of the system can be divided into three parts: online data acquisition, distributed data storage, and online data monitoring. Experimental tests show that the overall processing delay is controlled in the s level, which can meet the performance requirements of the power equipment online monitoring and real-time data processing. The business process is as shown in Figure 2 The method of the application comprises:
[0038] 1 online data acquisition;
[0039] The physical architecture of the power grid online data acquisition system is as shown in Figure 3 The D5000 national network online monitoring data acquisition machine is deployed in the I area, the SCADA measurement data are automatically calculated to generate the section data file every 5 minutes or 15 minutes, the real-time running condition data are saved in the text file with the QS suffix, and the scheduling unique name and the primary key of the equipment are saved in the text file with the ID suffix.
[0040] The SCADA measurement data are saved to the D5000 online measurement data cluster in the III area through the physical isolation device, and are externally published. In order to ensure the stable operation of the system, the D5000 system adopts the multi-machine hot standby mode, that is, the QS and ID files are stored in multiple servers. Each server may save the complete acquisition data of the State Grid, or may save part of the data, and may update a set of QS or ID files every 5 minutes, or may update the data of several days at a time.
[0041] The PSDB stream service is deployed in the III area, the data in the D5000 online measurement data cluster in the III area are collected to the distributed file of the PSDB system in the stream processing mode, and the collection results are recorded.
[0042] The online monitoring service of the PSDB system reads the collection results, and displays the running conditions of the collection service and the quality of the online data collection.
[0043] 2, SCADA measurement data acquisition;
[0044] Stream processing;
[0045] Through the D5000 online measurement data node server SFTP service, through the remote file operation interface, the folder file (QS and ID file) change on multiple D5000 online service nodes is monitored, if there is a new file, the remote file is collected to the local. Although the QS file is generated once every 5 minutes, the transmission time is random, and multiple D5000 services each may be a random group or multiple groups, and the monitoring time should not be too long, and the default is 1 minute. The monitoring time window supports free control, which can be set to seconds or minutes.
[0046] The stream computing service sequentially scans each D5000 service, each scanning task starts a sub-process, when the monitoring file is obtained, the QS or ID timestamp is the primary key, and the record in the PSDB system distributed file service is compared, if it does not exist, it is collected and stored. If it exists and the local file is legal (the file size is reasonable, the structured device monitoring parameter), delete the file on the remote server, complete a server scanning and collection.
[0047] Multiple collection processes write to a file library at the same time, or other application operations on the same local timestamp file (such as a data warehouse backup process is reading), the first process that occupies the file locks the file, and the concurrent processes do not process in this window period, if the local target file is operable next time window period, the collection file legality and the target file legality are judged, the latest and largest version is updated to the local file, and the remote server file is deleted after the update is successful.
[0048] In this way, the local monitoring data is guaranteed to only retain a set of monitoring data in the distributed file system, and the historical records on the remote server are deleted "with collection", avoiding the hard disk from being blown up. In order to ensure the safety and integrity of the collected data, the remote acquisition of measurement data and the deletion of the measurement data are separated in each collection process. Many reasons can cause the collected measurement data to be incomplete, which may be due to network reasons, or the data version inconsistency in multiple D5000 file services. The measurement data of the same timestamp in each D5000 file server needs to be compared, and the latest version is taken as the final version in the local.
[0049] At the same time, in order to ensure the safety of the data, the remote D5000 file server supports caching historical data for three days, that is, after the data collection is completed, only the records of the previous 72 hours are deleted.
[0050] 3. Distributed file storage
[0051] The data on the multiple online monitoring file servers is collected into the PSDB distributed file system through a stream processing service, the data is saved in a file form according to a monitoring collection time point, according to a logical hierarchical structure of years, months and days, and the storage logical structure of the D5000 online measurement data distributed file server is as shown in Figure 4 The distributed file storage is designed based on a cloud platform, historical backup is performed on original monitoring files, and the safety and global uniqueness of collected data are ensured.
[0052] After each collection is successful, an online data collection index table in a relational database is updated, and the collection information of the online data, such as a collection time, a distributed file server saving path and file information (size and creation time), is recorded in the index table.
[0053] Some business functions in the PSDB are operated based on the table, high-IO and high-concurrency access to the file server is avoided, and the online data monitoring function is taken as an example.
[0054] 4. Online monitoring;
[0055] The online monitoring service of the PSDB system monitors the running condition of the collection service and the quality of collected online data, displays the collection condition of online data of the current month in a month and day calendar mode by reading the online monitoring index table in real time, and prompts the collection condition of SCADA measurement data of the current month through the depth of color (green and red). If the collection is complete, the color is green, if the data is missing, the color will be gradually changed from green, the mouse is hovered over the table of the day, and a prompt box is popped up to display the collection condition of the data of the day.
[0056] The collection condition of the data of a day, such as a collection time, a distributed file server saving path and file information (size and creation time), is queried and displayed through a time point, and the data saved in the distributed file system can be downloaded to the local.
[0057] The application further provides a SCADA measurement data collection and storage system 200 based on stream calculation, as shown in Figure 5 The application further provides a SCADA measurement data collection and storage system 200 based on stream calculation, as shown in
[0058] The data collection unit 201 deploys an online monitoring data collection device in a target area of a power system, collects SCADA measurement data at a preset time interval, converts the SCADA measurement data into section data files, and saves the section data files;
[0059] The distributed storage unit 202 collects the saved section data files to the local according to stream calculation, saves the section data files saved to the local to the distributed file system through a stream processing service in a preset storage format.
[0060] The data monitoring unit 203 monitors the data collection condition and data quality of the cross-section data file stored in the distributed file system.
[0061] The preset time interval is 5-15 minutes.
[0062] The preset storage format is the structure of the collection time and collection location.
[0063] The cross-section data file is saved to the distributed file system, and an index table is established, and the collection information is established by the index table, and the collection information includes the collection time, the saving path and the file information.
[0064] The application realizes the SCADA measurement data collection in the stream computing mode, realizes the de-redundancy storage under the multi-source data, saves the measurement time stamp into the PSDB distributed file system, and monitors the running condition of the collection service and the collection online data quality through the online monitoring service.
[0065] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system or a computer program product. Therefore, the present application can adopt a completely hardware embodiment, a completely software embodiment or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer usable storage media containing computer usable program codes (including but not limited to disk memory, CD-ROM, optical memory, etc.). The solutions in the embodiments of the present application can be implemented in various computer languages, for example, object-oriented programming language Java and direct interpretation script language JavaScript, etc.
[0066] The present application is described with reference to flowcharts and / or block diagrams according to the methods, devices (systems) and computer program products of the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of the flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device produce a device for implementing the functions specified in the flowcharts and / or block diagrams. Figure 1 The device for implementing the functions specified in one flow or multiple flows and / or blocks. Figure 1 The device for implementing the functions specified in one block or multiple blocks.
[0067] These computer program instructions can also be stored in a computer readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer readable memory produce an article of manufacture including instructions which implement the Figure 1 function specified in the flow or flows and / or blocks Figure 1 of the block or blocks.
[0068] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions that are executed on the computer or other programmable apparatus provide steps for implementing the Figure 1 function specified in the flow or flows and / or blocks Figure 1 of the block or blocks.
[0069] While the preferred embodiments of the application have been described, additional variations and modifications can be made to the embodiments by those skilled in the art once they learn of the basic inventive concepts. Therefore, it is intended that the appended claims shall cover all such modifications and variations as fall within the true spirit and scope of the application.
[0070] Obviously, numerous modifications and variations of the present application are possible in light of the above teachings. It is therefore to be understood that within the scope of the appended claims and their equivalents, the application can be practiced otherwise than as specifically described.
Claims
1. A method for acquiring and storing SCADA measurement data based on stream computing, the method comprising: Online monitoring data acquisition devices are deployed in the target area of the power system to collect SCADA measurement data at preset time intervals, convert the SCADA measurement data into cross-sectional data files, and save the cross-sectional data files. The saved cross-sectional data files are collected locally using stream computing. The saved cross-sectional data files are then saved to the distributed file system in a preset storage format using the stream processing service. Monitor the data collection and quality of cross-sectional data files stored in the distributed file system; SCADA measurement data acquisition includes: stream processing, specifically: By using the SFTP service of the D5000 online measurement data node server and through the remote file operation interface, changes in folder files on multiple D5000 online service nodes can be monitored. If a new file is added, the remote file is collected and transferred to the local machine. Although the QS file is generated once every 5 minutes, the transmission time is random, and each of the multiple D5000 services may have one or more random sets. The monitoring time should not be too long. The default is 1 minute. The monitoring time window can be freely controlled and can be set to the second level or the minute level. The stream computing service sequentially scans each D5000 service. Each scan task starts a subprocess. After obtaining the monitoring file, the QS or ID timestamp is used as the primary key. The process is compared with the record in the PSDB system's distributed file service. If the file does not exist, it is collected and stored. If the file exists and the local file is legal, the file is deleted from the remote server. This completes the scanning and collection of one server. When multiple data collection processes write to a single file library simultaneously, or when other applications operate on the same local timestamp file, the first process to occupy the file locks the file. Concurrent processes do not process the file during this window period. In the next window period, if the local target file is operable, the validity of the data collection file and the target file are checked. The local file is updated with the latest and largest version. If the update is successful, the file on the remote server is deleted. This process is repeated to ensure that only one set of monitoring data is retained in the distributed file system, while historical records on the remote server are deleted as they are collected to avoid filling up the hard drive. To ensure the security and integrity of the collected data, the acquisition and deletion of measurement data are separated for each collection process. Many factors can cause the collected measurement data to be incomplete, such as network issues or inconsistent data versions on multiple remote D5000 file servers. It is necessary to compare the measurement data at the same timestamp on each D5000 file server and take the latest version as the final local version. Meanwhile, to ensure data security, it supports caching three days of historical data on a remote D5000 file server, meaning that only records from the 72 hours prior to the current time are deleted after data collection is completed.
2. The method according to claim 1, wherein the preset time interval is 5-15 minutes.
3. The method according to claim 1, wherein the preset storage format is a structure of acquisition time and acquisition location.
4. The method according to claim 1, wherein after the cross-sectional data file is saved to a distributed file system, an index table is established, and the collected information is established using the index table, the collected information including: Collection time, save path and file information.
5. A SCADA measurement data acquisition and storage system based on stream computing, the system comprising: The data acquisition unit deploys the online monitoring data acquisition device in the target area of the power system, collects SCADA measurement data at preset time intervals, converts the SCADA measurement data into cross-sectional data files, and saves the cross-sectional data files. The distributed storage unit collects the saved cross-sectional data files locally using stream computing, and then saves the locally stored cross-sectional data files to the distributed file system in a preset storage format through the stream processing service. The data monitoring unit monitors the data collection and quality of cross-sectional data files stored in the distributed file system. SCADA measurement data acquisition includes: stream processing, specifically: By using the SFTP service of the D5000 online measurement data node server and through the remote file operation interface, changes in folder files on multiple D5000 online service nodes can be monitored. If a new file is added, the remote file is collected and transferred to the local machine. Although the QS file is generated once every 5 minutes, the transmission time is random, and each of the multiple D5000 services may have one or more random sets. The monitoring time should not be too long. The default is 1 minute. The monitoring time window can be freely controlled and can be set to the second level or the minute level. The stream computing service sequentially scans each D5000 service. Each scan task starts a subprocess. After obtaining the monitoring file, the QS or ID timestamp is used as the primary key. The process is compared with the record in the PSDB system's distributed file service. If the file does not exist, it is collected and stored. If the file exists and the local file is legal, the file is deleted from the remote server. This completes the scanning and collection of one server. When multiple data collection processes write to a single file library simultaneously, or when other applications operate on the same local timestamp file, the first process to occupy the file locks the file. Concurrent processes do not process the file during this window period. In the next window period, if the local target file is operable, the validity of the data collection file and the target file are checked. The local file is updated with the latest and largest version. If the update is successful, the file on the remote server is deleted. This process is repeated to ensure that only one set of monitoring data is retained in the distributed file system, while historical records on the remote server are deleted as they are collected to avoid filling up the hard drive. To ensure the security and integrity of the collected data, the acquisition and deletion of measurement data are separated for each collection process. Many factors can cause the collected measurement data to be incomplete, such as network issues or inconsistent data versions on multiple remote D5000 file servers. It is necessary to compare the measurement data at the same timestamp on each D5000 file server and take the latest version as the final local version. Meanwhile, to ensure data security, it supports caching three days of historical data on a remote D5000 file server, meaning that only records from the 72 hours prior to the current time are deleted after data collection is completed.
6. In the system according to claim 5, the preset time interval is 5-15 minutes.
7. The system according to claim 5, wherein the preset storage format is a structure of acquisition time and acquisition location.
8. In the system according to claim 5, after the cross-sectional data file is saved to the distributed file system, an index table is established, and the collected information is established using the index table, the collected information including: Collection time, save path and file information.
Citation Information
Patent Citations
Power utilization information acquisition system and method based on big data technology
CN106651633A
Scheduling real-time section data generation method and system
CN108874859A