Data Detection Method and Device for Batch Streaming Fusion of Multi-source Time-series Data in Power Grid Regulation
By adopting data detection methods and devices that integrate multi-source time series data batch flow in the power grid regulation system, the incremental, re-transfer and historical data are uniformly processed, and the detection and calculation components are dynamically triggered, the problem of the inability to unified data detection and processing architecture in the prior art is solved, and the real-time and stability of data detection and calculation are improved.
Patent Information
- Application Number
- CN202410159276.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-02-04
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2044-02-04
AI Technical Summary
When the existing power grid control system processes multi-source timing data, there are differences in the processing methods, granularity and scheduling timing of incremental, re-transfer and historical data, resulting in the inability to unified data detection and processing architecture, occupying a large amount of resources and increasing operation and maintenance work costs.
Using a data detection method and device for controlling the fusion of multi-source timing data batch flows by the power grid, the service system data, dimension feature information and data time scale information are encapsulated into service analysis messages and sent to the data detection engine through the data collection module. The data detection engine determines the triggering method based on the data time scale information, processes the business system data, triggers the detection calculation component to detect it, and obtains the detection detailed data.
It realizes the unified processing of incremental, re-transfer and historical data, dynamically triggers detection and calculation components, supports real-time calculation and detection of streaming data, and can regularly batch the existing historical data, improves the real-time and stability of data detection and calculation, and simplifies the logic of data detection and development and operation and maintenance.
Smart Images

Figure CN118069643B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of power systems, and in particular to a data detection method and device for batch-stream fusion of multi-source time-series data in power grid regulation and control. Background Art
[0002] In the actual operation of the power grid regulation and control system and the process of data acquisition, the regulation and control operation data comes from multiple systems, and there are often time-series data of multiple time scales such as increment, retransmission, and history in the regulation and control operation data. The data types are diverse and the characteristics are various. Increment data, retransmission data mainly in the form of dynamic data streams and massive long-cycle historical stock data have certain differences in processing methods, granularity, and scheduling times, which often lead to the inability to fully unify the data detection and processing architecture. Currently, the detection of increment data and retransmission data mainly in the form of dynamic data streams and the detection of historical stock data are usually split into two different independent architectures for separate processing, occupying a large amount of resources and increasing the operation and maintenance work cost. Summary of the Invention
[0003] Aiming at solving the problem that the differences in processing methods, granularity, and scheduling times of increment data, retransmission data mainly in the form of dynamic data streams and massive long-cycle historical stock data cannot be unified in the data detection and processing architecture, the present invention provides a data detection method and device for batch-stream fusion of multi-source time-series data in power grid regulation and control.
[0004] In the first aspect, the present invention provides a data detection method for batch-stream fusion of multi-source time-series data in power grid regulation and control, and the method includes:
[0005] Collect data from multiple business systems at the source data end, and send the business system data to the data aggregation module according to the data type;
[0006] The data aggregation module receives the business system data, encapsulates the business system data, the dimension feature information of the business system data, and the data time scale information into a business analysis message, and sends the business analysis message to the data detection engine;
[0007] The data detection engine receives and parses the business analysis message to obtain the dimension feature information and data time scale information of the business system data, determines the triggering method according to the data time scale information, processes the business system data according to the triggering method to obtain the scheduling operation data, and triggers the detection calculation component to detect the scheduling operation data to obtain the detection detail data.
[0008] Based on the above technical solution, further, the data aggregation module receives the business system data, encapsulates the business system data, the dimensional feature information of the business system data, and the data time scale information into a business analysis message, and sends the business analysis message to the data detection engine, specifically including:
[0009] The data aggregation module determines the data time scale information of the business system data according to the topic of the message received for encapsulating the business system data, and the data time scale information includes incremental data type, reissued data type, or stock historical data type;
[0010] After the data aggregation module filters the data in the business system data whose primary key does not conform to the specification, it determines whether the primary key of the business system data is the same as the primary key of the business data already stored in the data block;
[0011] If they are the same, the business system data is stored at the position of the already stored business data with the same primary key in the data block and the already stored business data is overwritten;
[0012] Otherwise, the business system data is stored in the data block;
[0013] The data aggregation module encapsulates the business system data, the dimensional feature information of the business system data, the data time scale information, and the trigger mode into a business analysis message, and sends the business analysis message to the data detection engine.
[0014] Based on the above technical solution, further, the data detection engine receives and parses the business analysis message to obtain the dimensional feature information and data time scale information of the business system data, determines the trigger mode according to the data time scale information, processes the business system data according to the trigger mode to obtain scheduling operation data, and triggers the detection calculation component to detect the scheduling operation data to obtain detection detail data, specifically including:
[0015] When the data time scale information is incremental data type or reissued data type, and the trigger mode is real-time processing, the data detection engine processes the business analysis message in real time to obtain the scheduling operation data, and triggers the detection calculation component to perform detection through the scheduling operation data to obtain the detection detail data;
[0016] When the data time scale information is stock historical data type, and the trigger mode is scheduled task processing, the data detection engine starts a scheduled task to process the business analysis message to obtain the scheduling operation data, and triggers the detection calculation component to perform detection through the scheduling operation data to obtain the detection detail data.
[0017] Based on the above technical solution, further, when the data time scale information is of the incremental data type or the reissued data type, and the triggering method is real-time processing, the data detection engine processes the service analysis message in real time to obtain the scheduling operation data, and triggers the detection calculation component to perform detection through the scheduling operation data to obtain the detection detail data, specifically including:
[0018] The data detection engine determines whether the service analysis message is of the incremental data type or the reissued data type according to the data time scale information;
[0019] When the service analysis message is of the incremental data type, the data detection engine partitions the incremental data read from the service analysis message according to the dimension feature information;
[0020] The computing nodes of each data partition concurrently perform data preprocessing on the assigned incremental data to obtain the scheduling operation data, and call the detection calculation component to calculate the scheduling operation data to obtain a first aggregated calculation result;
[0021] The data detection engine triggers the computing nodes of the data partition to concurrently call the detection calculation component to perform a second aggregated calculation on the first aggregated calculation result according to the first aggregated calculation results of all the incremental data to obtain the detection detail data;
[0022] Among them, when the data detection engine is in the process of processing and calculating, it records the status information of each link. When an abnormal status link appears, it restarts the calculation of the task corresponding to the abnormal status link;
[0023] When the service analysis message is of the reissued data type, the data detection engine reads the original data corresponding to the reissued data from the database, and adds the reissued data to the original data to obtain the scheduling operation data;
[0024] Call the detection calculation component to calculate the scheduling operation data to obtain the detection detail data.
[0025] Based on the above technical solution, further, when the data time scale information is of the stock historical data type, and the triggering method is scheduled task processing, the data detection engine starts a scheduled task to process the service analysis message to obtain the scheduling operation data, and triggers the detection calculation component to perform detection through the scheduling operation data to obtain the detection detail data, specifically including:
[0026] The data detection engine sets a scheduled scheduling task for the stock historical data in the service analysis message according to the dimension feature information;
[0027] The timing scheduling task uses the stock historical data as the scheduling operation data, and calls the detection and calculation component to calculate the scheduling operation data to obtain the detection detail data.
[0028] Based on the above technical solution, further, the method further includes:
[0029] When the data detection engine's real-time processing of the service analysis message and the start of the timing task to process the service analysis message are triggered simultaneously, execute or suspend the calculation task according to a preset scheduling strategy, where the preset scheduling strategy includes weight distribution according to the priority of the task.
[0030] Based on the above technical solution, further, the triggering of the detection and calculation component to detect the processed scheduling operation data to obtain detection detail data specifically includes:
[0031] The detection and calculation component starts sub-detection components according to detection dimension identifiers, and the sub-detection components include an integrity sub-detection component, an accuracy sub-detection component, and a timeliness sub-detection component;
[0032] The integrity sub-detection component detects the integrity of the operation data collection type, the integrity of the stored data records, the start time of data storage, and the data storage period of the scheduling operation data to obtain first detection result data;
[0033] The accuracy sub-detection component detects quota rules, summation rules, mutual exclusion rules, comparison rules, and trend rules for the scheduling operation data to obtain second detection result data;
[0034] The timeliness sub-detection component detects the data transmission period of the scheduling operation data to obtain third detection result data;
[0035] According to the first detection result data, the second detection result data, and / or the third detection result data, the detection detail data is obtained.
[0036] In a second aspect, the present invention further provides a data detection device for batch-stream fusion of multi-source time-series data in power grid regulation and control. The device includes a source data end, a data aggregation module, and a data detection engine. The data detection engine includes a detection and calculation component:
[0037] The source data end is used to collect data from multiple business systems and send the business system data to the message bus according to the data type;
[0038] The data aggregation module is used to receive the business system data from the message bus, encapsulate the business system data, the dimensional feature information of the business system data, and the data time scale information into a business analysis message, and send the business analysis message to the data detection engine;
[0039] The data detection engine is used to receive and parse the business analysis message to obtain the dimensional feature information and data time scale information of the business system data, determine the triggering method according to the data time scale information, process the business system data according to the triggering method to obtain scheduling operation data, and trigger the detection calculation component;
[0040] The detection calculation component is used to detect the scheduling operation data to obtain detection detail data.
[0041] Based on the above technical solution, further, the data aggregation module is specifically used to determine the data time scale information of the business system data according to the topic of the message encapsulating the business system data received, and the data time scale information includes incremental data type, reissue data type, or inventory historical data type;
[0042] The data aggregation module is specifically used to filter the data in the business system data whose primary key does not conform to the specification, and then determine whether the primary key of the business system data is the same as the primary key of the business data already stored in the data block;
[0043] If they are the same, store the business system data at the position of the already stored business data with the same primary key in the data block and overwrite the already stored business data;
[0044] Otherwise, store the business system data in the data block;
[0045] The data aggregation module is specifically used to encapsulate the business system data, the dimensional feature information of the business system data, and the data time scale information into a business analysis message, and send the business analysis message to the data detection engine.
[0046] Based on the above technical solution, further, when the data time scale information is incremental data type or reissue data type, and the triggering method is real-time processing, the data detection engine processes the business analysis message in real time to obtain the scheduling operation data, and triggers the detection calculation component to perform detection through the scheduling operation data to obtain the detection detail data;
[0047] When the data time scale information is of the stock historical data type and the triggering method is timed task processing, the data detection engine starts a timed task to process the service analysis message, obtains the scheduling operation data, and triggers the detection calculation component to perform detection through the scheduling operation data to obtain the detection detail data.
[0048] Based on the above technical solution, further, the data detection engine is specifically configured to determine whether the service analysis message is of the incremental data type or the resending data type according to the data time scale information;
[0049] When the service analysis message is of the incremental data type, the data detection engine is specifically configured to partition the incremental data read from the service analysis message according to the dimension feature information;
[0050] The computing nodes of each data partition concurrently perform data preprocessing on the allocated incremental data to obtain scheduling operation data, and call the detection calculation component to calculate the scheduling operation data to obtain a primary aggregation calculation result;
[0051] The data detection engine is specifically configured to trigger the computing nodes of the data partition to concurrently call the detection calculation component to perform secondary aggregation calculation on the primary aggregation calculation result according to the primary aggregation calculation result of all the incremental data to obtain the detection detail data;
[0052] Among them, the data detection engine is specifically configured to record the status information of each link during the processing and calculation process, and when an abnormal status link appears, restart the calculation of the task corresponding to the abnormal status link;
[0053] When the service analysis message is of the resending data type, the data detection engine is specifically configured to read the original data corresponding to the resending data from the database, and add the resending data to the original data to obtain scheduling operation data;
[0054] The data detection engine is specifically configured to call the detection calculation component to calculate the scheduling operation data to obtain the detection detail data.
[0055] Based on the above technical solution, further, the data detection engine is specifically configured to set a timed scheduling task for the stock historical data in the service analysis message according to the dimension feature information;
[0056] The timed scheduling task uses the stock historical data as the scheduling operation data, and calls the detection calculation component to calculate the scheduling operation data to obtain the detection detail data.
[0057] Based on the above technical solution, further, when the data detection engine is specifically used to process the service analysis message in real time and start a timed task to process the service analysis message at the same time, the execution or suspension of the calculation task is performed according to a preset scheduling strategy, where the preset scheduling strategy includes weight allocation according to the priority of the task.
[0058] Based on the above technical solution, further, the detection and calculation component is specifically used to start the sub-detection component according to the detection dimension identifier, and the sub-detection component includes an integrity sub-detection component, an accuracy sub-detection component, and a timeliness sub-detection component;
[0059] The integrity sub-detection component is used to detect the integrity of the operation data aggregation type, the integrity of the stored data records, the start time of data storage, and the data storage period of the scheduling operation data, and obtain the first detection result data;
[0060] The accuracy sub-detection component is used to detect the quota rule, the summation rule, the mutual exclusion rule, the comparison rule, and the trend rule of the scheduling operation data, and obtain the second detection result data;
[0061] The timeliness sub-detection component is used to detect the data transmission period of the scheduling operation data, and obtain the third detection result data;
[0062] The detection and calculation component is specifically used to obtain the detection detail data according to the first detection result data, the second detection result data, and / or the third detection result data, and send the detection detail data to the data storage module.
[0063] In a third aspect, the present invention also provides a computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, it implements the data detection method for batch-stream fusion of multi-source time-series data in power grid regulation and control described in any one of the first aspects.
[0064] In a fourth aspect, the present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the data detection method for batch-stream fusion of multi-source time-series data in power grid regulation and control described in any one of the first aspects.
[0065] The data detection method and device for batch-stream fusion of multi-source time-series data in power grid regulation provided by the present invention include: the source data end sends service system data to the message bus respectively, and the data aggregation module encapsulates the service system data, the dimension feature information of the service system data, and the data time scale information into a service analysis message, and sends the service analysis message to the data detection engine; the data detection engine determines the triggering mode according to the data time scale information, processes the service system data according to the triggering mode, and sends the processed scheduling operation data to the detection calculation component, the detection calculation component starts the corresponding detection calculation component according to the scheduling operation data to obtain detection detail data, and sends the detection detail data to the data storage module, and the data storage module establishes a data multi-dimensional index for the detection detail data according to the service time, data source, data type, and detection dimension, constructs a database storage primary key through the data multi-dimensional index, and stores the detection detail data. The present invention realizes unified data detection logic based on the big data stream and batch computing framework, constructs a set of data detection components applicable to stream data and batch data, and through setting a reasonable processing mechanism and scheduling strategy for batch-stream fusion, dynamically processes multi-time scale regulation operation data such as increment, resending, and history, adaptively triggers the detection calculation component, supports both real-time calculation and detection of stream data, ensures the quality of newly added data in real-time detection, and at the same time can periodically perform batch processing on the stock historical data, without splitting the stream data detection and batch data detection into two different independent architectures for separate processing, improves the real-time performance and stability of data detection calculation, and at the same time simplifies the working logic of data detection development and operation and maintenance, and solves the problems of high tuning labor cost and low resource utilization rate. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] The specification drawings forming a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:
[0067] Figure 1 is a schematic flow chart of the data detection method for batch-stream fusion of multi-source time-series data in power grid regulation provided by an embodiment of the present invention;
[0068] Figure 2 is a schematic diagram of the data detection system framework of the data detection method for batch-stream fusion of multi-source time-series data in power grid regulation provided by another embodiment of the present invention;
[0069] Figure 3 is a schematic flow chart of the data detection method for batch-stream fusion of multi-source time-series data in power grid regulation provided by another embodiment of the present invention.
[0070] Figure 4It is a schematic diagram of a streaming parallel computing method for multi-source partitioning and breakpoint resumption of time-series data in the data detection method for batch-stream fusion of multi-source time-series data for power grid regulation and control provided by another embodiment of the present invention;
[0071] Figure 5 It is a module schematic diagram of a data detection device for batch-stream fusion of multi-source time-series data for power grid regulation and control provided by another embodiment of the present invention. Specific embodiments
[0072] The present invention will be described in detail below with reference to the drawings and in conjunction with embodiments. It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments may be combined with each other.
[0073] The following detailed descriptions are all exemplary descriptions, aiming to provide further detailed descriptions of the present invention. Unless otherwise specified, all technical terms adopted by the present invention have the same meaning as commonly understood by those of ordinary skill in the art to which the present invention pertains. The terms used in the present invention are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention.
[0074] The following will be combined with the attached Figure 1 , and the data detection method for batch-stream fusion of multi-source time-series data for power grid regulation and control provided by the embodiments of the present invention will be described, including the following steps:
[0075] S1. The source data end collects data from multiple business systems and sends the business system data to the data aggregation module according to the data type;
[0076] Specifically, the source data end can be dispersed at each business system end or on the server side. The source data end collects data from multiple business systems and sends the business system data to the data aggregation module through the message bus respectively.
[0077] S2. The data aggregation module receives the business system data, encapsulates the business system data, the dimension feature information of the business system data, and the data time scale information into a business analysis message, and sends the business analysis message to the data detection engine;
[0078] Specifically, the data aggregation module can receive the business system data through the message bus or by other means, which is not limited in the present invention.
[0079] S3. The data detection engine receives and parses the business analysis message to obtain the dimension feature information and data time scale information of the business system data, determines the triggering mode according to the data time scale information, processes the business system data according to the triggering mode to obtain the processed dispatching operation data, and triggers the detection calculation component to detect the processed dispatching operation data to obtain the detection detailed data.
[0080] Specifically, the data detection engine sends the detection detail data to the data storage module. The data storage module establishes a multi-dimensional data index for the detection detail data according to the service time, data source, data type, and detection dimension, constructs a database storage primary key through the multi-dimensional data index, and stores the detection detail data.
[0081] Based on the above embodiments, further, step S2 specifically includes:
[0082] The data aggregation module determines the data time scale information of the business system data according to the topic of the message encapsulating the business system data received. The data time scale information includes incremental data type, reissued data type, or stock historical data type;
[0083] The data aggregation module filters the data in the business system data with non-compliant primary keys and then determines whether the primary key of the business system data is the same as the primary key of the business data already stored in the data block;
[0084] If they are the same, the business system data is stored at the position of the already stored business data with the same primary key in the data block and the already stored business data is overwritten;
[0085] Otherwise, the business system data is stored in the data block;
[0086] The data aggregation module encapsulates the business system data, the dimension feature information of the business system data, and the data time scale information into a business analysis message and sends the business analysis message to the data detection engine.
[0087] Based on the above embodiments, further, step S3 specifically includes:
[0088] S31. When the data time scale information is incremental data type or reissued data type and the triggering method is real-time processing, the data detection engine processes the business analysis message in real time to obtain the scheduling operation data, and triggers the detection calculation component to perform detection through the scheduling operation data to obtain the detection detail data;
[0089] S32. When the data time scale information is stock historical data type and the triggering method is scheduled task processing, the data detection engine starts a scheduled task to process the business analysis message, obtains the scheduling operation data, and triggers the detection calculation component to perform detection through the scheduling operation data to obtain the detection detail data.
[0090] Based on the above embodiments, further, step S31 specifically includes:
[0091] The data detection engine determines whether the service analysis message is an incremental data type or a retransmission data type according to the data time scale information;
[0092] When the service analysis message is an incremental data type, the data detection engine partitions the incremental data read from the service analysis message according to the dimension feature information;
[0093] The computing nodes of each data partition concurrently perform data preprocessing on the assigned incremental data to obtain scheduled operation data, and after calling the detection computing component to calculate the scheduled operation data, a primary aggregation calculation result is obtained;
[0094] The data detection engine triggers the computing nodes of the data partition to concurrently call the detection computing component to perform secondary aggregation calculation on the primary aggregation calculation result according to the primary aggregation calculation result of all the incremental data, and the detection detail data is obtained;
[0095] Among them, when the data detection engine is in the process of processing and calculating, it records the status information of each link, and when an abnormal status link appears, it restarts the calculation of the task corresponding to the abnormal status link;
[0096] When the service analysis message is a retransmission data type, the data detection engine reads the original data corresponding to the retransmission data from the database, and adds the retransmission data to the original data to obtain scheduled operation data;
[0097] Call the detection computing component to calculate the scheduled operation data to obtain the detection detail data.
[0098] Based on the above embodiment, further, step S32 specifically includes:
[0099] The data detection engine sets a timed scheduling task for the stock historical data in the service analysis message according to the dimension feature information;
[0100] The timed scheduling task uses the stock historical data as the scheduled operation data, and calls the detection computing component to calculate the scheduled operation data to obtain the detection detail data.
[0101] Based on the above embodiment, further, step S3 further includes:
[0102] When the data detection engine is triggered to process the service analysis message in real time and start the timed task to process the service analysis message at the same time, the execution or suspension of the calculation task is performed according to the preset scheduling policy, where the preset scheduling policy includes weight assignment according to the priority of the task.
[0103] Based on the above embodiments, further, in step S3, the trigger detection calculation component detects the processed scheduling operation data to obtain detection detail data, which specifically includes:
[0104] The detection calculation component starts the sub-detection components according to the detection dimension identifier, and the sub-detection components include an integrity sub-detection component, an accuracy sub-detection component, and a timeliness sub-detection component;
[0105] The integrity sub-detection component detects the integrity of the operation data aggregation type, the integrity of the stored data records, the start time of data storage, and the data storage period of the scheduling operation data to obtain the first detection result data;
[0106] The accuracy sub-detection component detects the quota rule, the summation rule, the mutual exclusion rule, the comparison rule, and the trend rule of the scheduling operation data to obtain the second detection result data;
[0107] The timeliness sub-detection component detects the data transmission period of the scheduling operation data to obtain the third detection result data;
[0108] According to the first detection result data, the second detection result data, and / or the third detection result data, the detection detail data is obtained.
[0109] It should be understood that the method of the present invention first receives, through the data aggregation module, incremental and supplementary recruitment data packets of multiple business systems and multiple types sent by the source data end in real time, analyzes and extracts the original data in the packets and dimension feature information such as the primary key, data object, data source, data type, time, and data value, writes the original data and dimension feature information into a data block for storage, and at the same time forwards the incremental and supplementary recruitment original data and dimension features to the data detection engine in real time to trigger the detection processing flow.
[0110] The data detection engine provides a unified detection calculation component, covering data missing value and data accuracy detection. Through the adaptive interaction between the batch flow processing module and the detection calculation component, the engine dynamically processes operation data regulated by multiple time scales such as increment, supplement, and history. For incremental data and supplementary data, a streaming parallel computing processing method based on multi-source partitioning of time series data is proposed, and real-time calculation is performed through the stream processing module. For historical stock data, regular calculation is performed through the batch processing module.
[0111] This embodiment realizes unified data detection logic based on big data streams and batch computing frameworks, constructs a set of data detection components applicable to both streaming data and batch data, and through setting a reasonable processing mechanism and scheduling strategy for batch-stream fusion, dynamically processes operation data with multiple time scales such as incremental, resupplied, and historical data, adaptively triggers the detection computing components, supports both real-time computing and detection of streaming data to ensure the quality of newly added data in real-time detection, and at the same time can periodically perform batch processing on the existing historical data, without splitting the streaming data detection and batch data detection into two different independent architectures for separate processing, improving the real-time performance and stability of data detection computing, while simplifying the working logic of data detection development and operation and maintenance, and solving problems such as high optimization labor costs and low resource utilization rates.
[0112] The following will combine the attached Figures 2 to 4 Describe the data detection of batch-stream fusion of multi-source time-series data for power grid regulation provided by the embodiments of the present invention, including the following steps:
[0113] Step 1: Parse and collect multi-system and multi-type massive real-time power grid regulation data, extract and forward the parsed original data and dimension feature information such as primary key, data source, data type, time, and data value.
[0114] Step 2: The data detection engine identifies the input data time scale information and dimension feature information, and uses the processing mechanism and scheduling strategy of batch-stream fusion to dynamically process operation data with multiple time scales such as incremental, resupplied, and historical data, and adaptively triggers the detection computing components.
[0115] Step 3: The detection computing components detect the scheduled operation data, and can selectively start computing components such as data missing value and data accuracy detection to perform detection computing on the input data.
[0116] Step 4: Store the detailed result data formed by the detection and perform multi-dimensional statistical analysis to form a multi-dimensional portrait of data quality, and perform visual dynamic display.
[0117] In step 1, parse and collect multi-system and multi-type massive real-time power grid regulation data, extract and forward the parsed original data and dimension feature information such as primary key, data object, data source, data type, time, and data value.
[0118] The source data end encapsulates the multi-business system data collected into a message and sends it to the regulation big data platform through the message bus.
[0119] The data sent from the source data end includes various message types such as incremental data and retransmission data. Incremental data is the new data generated by the dispatching business system every day, and retransmission data is the historical data retransmitted. Their message topics and keys are different. The data aggregation program of the regulation big data platform receives and parses the message data sent from the source data end in real time, and initially filters out the data that does not conform to the primary key specification. On the one hand, the aggregated original data is stored, and the newly aggregated data will overwrite the historical data with the same primary key to ensure data uniqueness; at the same time, the original data and the dimension feature information such as the extracted primary key, data source, data type, time, and data value are encapsulated into a message to drive business analysis and sent to the data detection engine to drive the scheduling and execution of the data detection service.
[0120] In step 2, the data detection engine identifies the input data time scale information and dimension feature information, and uses the batch-stream fusion processing mechanism and scheduling strategy to dynamically process the multi-time scale regulation operation data such as incremental, retransmission, and historical data, and adaptively triggers the detection calculation component.
[0121] There are two ways to trigger the data detection engine. One is data stream trigger, which mainly performs real-time processing on the incremental data and retransmission data sent by data aggregation; the other is timed task trigger, which mainly performs timed processing on the stored historical data in the big data platform.
[0122] For the processing method of data stream trigger, the data detection engine receives the data sent by data aggregation, and judges whether it is incremental data or retransmission data according to the input data time scale information.
[0123] If it is incremental data, data preprocessing is performed through the streaming parallel computing method of multi-source partitioning and breakpoint resumption of time series data. The specific implementation method is as follows:
[0124] (1) The concurrent read incremental data is partitioned according to dimensions such as data object, data source, and data type. Each partition calculation node concurrently performs data preprocessing on the data, and calls the data detection calculation component to calculate the data, realizing a primary aggregation calculation of the dimension data;
[0125] (2) For the initially calculated data results, secondary aggregation is performed again according to partitions to realize the calculation of data detection results by day. By adopting the multi-source partitioning method, multi-partition data parallel computing can be realized, improving the performance of data reading, writing, etc.; at the same time, during the processing and calculation of each link of the data stream, the state of each link will be recorded. If the data stream breakpoint in a certain link causes the task to restart the calculation, instead of restarting multiple tasks, only the task of the abnormal state link is restarted for calculation, reducing the multi-task restart calculation caused by the data stream breakpoint, effectively saving computing resources and improving computing efficiency.
[0126] If it is resubmitted data, there are slight differences in the acquisition of data participating in detection and the processing of calculation results compared with incremental data. Resubmitted data is often historical data that is resubmitted, and there is a situation where only partial time period data within a day is resubmitted. It is difficult to analyze the characteristics of one-day time series data from the data stream. Therefore, it is necessary to read historical data from the database to obtain one-day data, and then form a complete data stream before calling the data detection and calculation component to calculate the data.
[0127] For the calculated results, it is necessary to update the historical data detection results stored in the database according to the dimensional feature information such as the primary key, data object, data source, data type, time, and data value to ensure the uniqueness and validity of the detection results.
[0128] For the triggering method of scheduled tasks, first, schedule tasks for the existing historical data according to time, data object, data type, and verification dimension, extract the existing historical data at regular intervals, and call the data detection and calculation component to calculate the data to achieve batch detection of the existing historical data.
[0129] When the data stream and scheduled tasks are triggered simultaneously, the batch-stream fusion calculation engine will dynamically collect task operation parameters, execute and suspend calculation tasks according to the set scheduling strategy. The engine will assign weight to the priorities of tasks, and tasks with higher weights will be triggered first to optimize the coordination between batch and stream tasks and prevent calculation task conflicts.
[0130] In step 3, the detection and calculation component detects the scheduled operation data, and can selectively start calculation components such as data missing value and data accuracy detection to detect and calculate the input data.
[0131] In the detection and calculation component, integrity, accuracy, and timeliness rule components are preset according to detection dimensions and rules, providing refined data detection rule items and solutions for each type of quality problem, and constructing a configurable and easily extensible detection rule library for multi-source time series data of the power grid, as shown in Table 1.
[0132]
[0133] Table 1 Detection Rule Library for Multi-source Time Series Data of the Power Grid
[0134] When calling the detection and calculation component, various rule calculation components such as data integrity, accuracy, and timeliness can be selectively started by passing identification parameters. The detection and calculation component obtains the rules and threshold parameters in the detection rule library according to the identification parameters to detect the data.
[0135] The detection and calculation component can be oriented to multi-scenario services such as high-precision prediction of multi-time-scale loads and analysis of the power supply-demand balance of the entire network. It analyzes the data requirements and characteristics of different business scenarios, and dynamically, intelligently, and adaptively adjusts data detection rules, thresholds, etc. to meet the diverse and personalized data detection requirements of grid regulation services.
[0136] In step 4, the detailed result data formed by detection is stored and subjected to multi-dimensional statistical analysis to form a multi-dimensional portrait of data quality, and a visual dynamic display is performed.
[0137] The detailed result data calculated in steps 3 and 4 will be stored in the data block. For resupplied data and historical stock data, it is necessary to iteratively update the result data calculated historically to ensure that the data stored in the database is the latest result data.
[0138] The detection result statistics module performs multi-dimensional statistics on the detailed results, including but not limited to time dimensions such as weekly, monthly, and annual, data object dimensions, and dimensions of each dispatching agency, etc. It uses forms such as bar charts and pie charts to analyze the rankings of each dispatching agency and the proportion of problems, etc. For the daily quality situation of each data asset, a long-time-scale panoramic analysis and display are carried out through a data calendar, providing a fast, business-driven multi-dimensional interactive analysis and query service for the detection results of grid regulation data, assisting various business applications to understand the quality of the data used from multiple angles, and providing high-quality panoramic data support for in-depth mining of the value of grid data.
[0139] As Figure 2 shown in the data detection system framework for batch-stream fusion of multi-source time-series data for grid regulation, specifically including:
[0140] The source data end module collects data from multiple business systems and encapsulates the data into messages of different themes according to different data types and sends them to the message bus in real time. Incremental data is the new data generated by the dispatching business system every day, and resupplied data is the historical data re-sent and uploaded.
[0141] The data aggregation module receives and parses the message data sent from the source data end in real time, and initially filters out the data that does not conform to the primary key specification. On the one hand, the aggregated raw data is stored, and the newly aggregated data will overwrite the historical data with the same primary key to ensure data uniqueness; at the same time, the raw data and the extracted dimension feature information such as primary key, data source, data type, time, and data value are encapsulated into messages to drive business analysis and sent to the data detection engine to drive the scheduling and execution of data detection services.
[0142] The data detection engine module identifies the input data time scale information and dimensional feature information, and uses the batch-stream fusion processing mechanism and scheduling strategy to dynamically process the incremental, supplementary, historical and other multi-time scale regulated operation data, adaptively trigger the detection calculation component to perform data detection, and statistically analyze and display the detection results.
[0143] The data storage module establishes a multi-dimensional index of data through key information such as business time, data source, data type, and detection dimension, constructs a reasonable database storage primary key to store the detection result data. For better analysis and display, it regularly synchronizes the detailed result data to the MPP database, combines the model data to support the analysis and display of the detection results, and stores the statistical result data as needed.
[0144] As Figure 3 shown in the schematic diagram of the detection method process for batch-stream fusion of multi-source time-series data of power grid regulation and control, it includes the following steps:
[0145] 1) For the data collected and forwarded by the input data, identify the input data time scale information, dimensional feature information, and trigger mode. If the trigger mode is data stream trigger, then execute step 2) to perform data time scale information discrimination. Otherwise, if the trigger mode is a scheduled task mode, then prioritize task scheduling, then obtain the stock historical data that needs to be detected and calculated from the database, and then execute step 5).
[0146] 2) Discriminate the data time scale information. If it is incremental data, then execute step 4) to perform data preprocessing. Otherwise, if it is supplementary data, execute step 3) to obtain historical data.
[0147] 3) Obtain historical data. According to the primary keys such as the incoming time, data object, data source, and data type, read the stock historical data in the data block, and fuse it with the supplementary data to form a time-series data set for a complete day, and then execute step 4).
[0148] 4) Data preprocessing. Use the streaming parallel computing method of multi-source partitioning and breakpoint resumption of time-series data to perform data preprocessing to form a data set to be detected.
[0149] 5) Call the detection calculation component, and start the detection calculation component by passing parameters such as time, data object, data source, data type, and verification dimension to perform data quality detection.
[0150] 6) Details of the detection calculation results. The detection calculation component verifies and generates details of the detection calculation results such as data problems and stores them.
[0151] 7) Statistical analysis of the detection results. Perform multi-dimensional statistical analysis on the detailed results to form data for long-time scale, multi-dimensional panoramic analysis and display, and store it as needed.
[0152] 8) Display of detection results, presenting the detailed results and multi-dimensional statistical analysis results, providing a fast, business-driven multi-dimensional interactive analysis and query service for grid regulation data detection results, and assisting various business applications to understand the quality of the data used from multiple perspectives.
[0153] Such as Figure 4 The streaming parallel computing method for multi-source partitioning and breakpoint resuming of time-series data shown in the figure includes:
[0154] The collected data is partitioned according to dimensions such as data object, data source, and data type. Each partition computing node concurrently reads and parses the data, then forwards the time-series data according to the partition, and triggers the detection computing engine. At the same time, the collected data is also stored in the database in full volume. During the processing and computing of each link in the data stream, the status of each link is recorded, and checkpoints are set to monitor events. If a data stream breakpoint in a certain link causes the task to restart the calculation, it is not necessary to restart multiple tasks, but only restart the task of the abnormal status link, reducing the multi-task restart calculation caused by the data stream breakpoint, effectively saving computing resources and improving computing efficiency.
[0155] The present invention partitions the collected data according to dimensions such as data object, data source, and data type through multi-source partitioning technology. Each partition computing node concurrently reads and parses the data, then forwards the time-series data according to the partition, and triggers the detection computing engine. At the same time, the collected data is also stored in the database in full volume. During the processing and computing of each link in the data stream, the status of each link is recorded, and checkpoints are set to monitor events. If a data stream breakpoint in a certain link causes the task to restart the calculation, it is not necessary to restart multiple tasks, but only restart the task of the abnormal status link, reducing the multi-task restart calculation caused by the data stream breakpoint, effectively saving computing resources and improving computing efficiency.
[0156] The present invention proposes an adaptive processing mechanism and scheduling strategy for batch-stream fusion of multi-time-scale regulation operation data such as increment, supplement, and history, identifies the input data time-scale information and dimension feature information, dynamically processes the multi-time-scale regulation operation data such as increment, supplement, and history, and adaptively triggers the detection computing component. When the data stream and the timing task are triggered simultaneously, the batch-stream fusion computing engine will dynamically collect the task operation parameters, execute and suspend the computing tasks according to the set scheduling strategy. The engine will assign weights to the priorities of the tasks, and the tasks with higher weights will be triggered first, optimizing the coordination between batch-stream tasks and preventing computing task conflicts.
[0157] The present invention proposes a data detection method based on batch-stream fusion, which presets integrity, accuracy, and timeliness rule components according to detection dimensions and detection rules, provides data detection solutions for each type of quality problem, constructs a configurable and easily extensible detection rule library for multi-source time-series data of the power grid, and at the same time, for multi-scenario services such as high-precision prediction of load at multiple time scales and analysis of power supply and demand balance across the network, analyzes the data requirements and characteristics of different business scenarios, and can dynamically, intelligently, and adaptively adjust data detection rules, thresholds, etc., to meet the diverse and personalized data detection requirements of power grid regulation services.
[0158] The following will be combined with the attached Figure 5 , to illustrate the data detection device for batch-stream fusion of multi-source time-series data for power grid regulation provided by the embodiments of the present invention. The device includes a source data end, a data aggregation module, and a data detection engine. The data detection engine includes a detection calculation component:
[0159] The source data end is used to collect data from multiple business systems and send the business system data to the data aggregation module according to the data type;
[0160] The data aggregation module is used to receive the business system data, encapsulate the business system data, the dimension feature information of the business system data, and the data time scale information into a business analysis message, and send the business analysis message to the data detection engine;
[0161] The data detection engine is used to receive and parse the business analysis message to obtain the dimension feature information and data time scale information of the business system data, determine the trigger mode according to the data time scale information, process the business system data according to the trigger mode to obtain the processed dispatching operation data, and trigger the detection calculation component;
[0162] The detection calculation component is used to detect the processed dispatching operation data to obtain detection detail data.
[0163] Specifically, the device further includes a data storage module; the data detection engine sends the detection detail data to the data storage module;
[0164] The data storage module is used to establish a multi-dimensional index for the detection detail data according to business time, data source, data type, and detection dimension, construct a database storage primary key through the multi-dimensional index, and store the detection detail data.
[0165] Based on the above embodiments, further, the data aggregation module is specifically configured to determine the data time scale information of the business system data according to the topic of the message encapsulating the business system data, where the data time scale information includes incremental data type, reissued data type, or stock historical data type;
[0166] The data aggregation module is specifically configured to filter the data in the business system data whose primary key does not conform to the specification, and then determine whether the primary key of the business system data is the same as the primary key of the business data already stored in the data block;
[0167] If they are the same, store the business system data at the position of the already stored business data with the same primary key in the data block and overwrite the already stored business data;
[0168] Otherwise, store the business system data in the data block;
[0169] The data aggregation module is specifically configured to encapsulate the business system data, the dimension feature information of the business system data, and the data time scale information into a business analysis message, and send the business analysis message to the data detection engine.
[0170] Based on the above embodiments, further, when the data time scale information is incremental data type or reissued data type, and the triggering method is real-time processing, the data detection engine processes the business analysis message in real time to obtain the scheduling operation data, and triggers the detection calculation component to perform detection through the scheduling operation data to obtain the detection detail data;
[0171] When the data time scale information is stock historical data type, and the triggering method is scheduled task processing, the data detection engine starts a scheduled task to process the business analysis message to obtain the scheduling operation data, and triggers the detection calculation component to perform detection through the scheduling operation data to obtain the detection detail data.
[0172] Based on the above embodiments, further, the data detection engine is specifically configured to determine whether the business analysis message is incremental data type or reissued data type according to the data time scale information;
[0173] When the business analysis message is incremental data type, the data detection engine is specifically configured to partition the incremental data read from the business analysis message according to the dimension feature information;
[0174] The computing nodes of each data partition concurrently perform data preprocessing on the assigned incremental data to obtain the scheduling operation data, and call the detection calculation component to calculate the scheduling operation data to obtain a primary aggregation calculation result;
[0175] Specifically, the data detection engine is configured to trigger the computing nodes of the data partition to concurrently call the detection computing component to perform secondary aggregation calculation on the result of the primary aggregation calculation of all the incremental data, so as to obtain the detailed detection data;
[0176] Specifically, during the processing and calculation, the data detection engine is configured to record the status information of each link. When an abnormal status link occurs, restart the calculation of the task corresponding to the abnormal status link;
[0177] When the service analysis message is of the retransmission data type, the data detection engine is specifically configured to read the original data corresponding to the retransmission data from the database, and add the retransmission data to the original data to obtain the scheduling operation data;
[0178] The data detection engine is specifically configured to call the detection computing component to calculate the scheduling operation data, so as to obtain the detailed detection data.
[0179] Based on the above embodiments, further, the data detection engine is specifically configured to set a timing scheduling task for the stock historical data in the service analysis message according to the dimension feature information;
[0180] The timing scheduling task takes the stock historical data as the scheduling operation data, and calls the detection computing component to calculate the scheduling operation data, so as to obtain the detailed detection data.
[0181] Based on the above embodiments, further, when the data detection engine is specifically configured to process the service analysis message in real time and start a timing task to process the service analysis message at the same time, execute or suspend the calculation task according to a preset scheduling policy, where the preset scheduling policy includes weight distribution according to the priority of the task.
[0182] Based on the above embodiments, further, the detection computing component is specifically configured to start the sub-detection component according to the detection dimension identifier, and the sub-detection component includes an integrity sub-detection component, an accuracy sub-detection component, and a timeliness sub-detection component;
[0183] The integrity sub-detection component is configured to detect the integrity of the operation data aggregation type, the integrity of the stored data records, the start time of data storage, and the data storage period of the scheduling operation data, so as to obtain the first detection result data;
[0184] The accuracy sub-detection component is configured to detect the quota rule, the summation rule, the mutual exclusion rule, the comparison rule, and the trend rule of the scheduling operation data, so as to obtain the second detection result data;
[0185] The timeliness detection component is used to detect the data transmission period of the scheduling operation data, and obtain the third detection result data;
[0186] The detection and calculation component is specifically used to obtain the detection detail data according to the first detection result data, the second detection result data, and / or the third detection result data, and send the detection detail data to the data storage module.
[0187] This embodiment realizes a unified data detection logic based on big data streams and batch computing frameworks, constructs a set of data detection components applicable to stream data and batch data, and through setting a reasonable processing mechanism and scheduling strategy for batch-stream fusion, dynamically processes operation data with multiple time scales such as increment, resending, and history, adaptively triggers the detection and calculation component, supports both real-time calculation and detection of stream data, ensures the quality of newly added data in real-time detection, and at the same time can perform batch processing on the stock historical data regularly. There is no need to split the stream data detection and batch data detection into two different independent architectures for separate processing, which improves the real-time performance and stability of data detection and calculation, and simplifies the working logic of data detection development and operation and maintenance, and solves problems such as high optimization labor costs and low resource utilization rates.
[0188] In addition, an embodiment of the present invention includes a computer device, including a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, it implements the data detection method for batch-stream fusion of multi-source time-series data in power grid regulation described in any one of the above technical solutions.
[0189] An embodiment of the present invention further includes a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the data detection method for batch-stream fusion of multi-source time-series data in power grid regulation described in any one of the above technical solutions.
[0190] As is known by technical common sense, the present invention can be implemented by other embodiments that do not deviate from its spiritual essence or essential features. Therefore, the above-disclosed embodiments are illustrative in all aspects and are not exclusive. All changes within the scope of the present invention or equivalent to the present invention are included in the present invention.
[0191] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0192] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in one or more flows Figure 1 one or more flows and / or blocks Figure 1 or means for implementing the functions specified in one or more blocks.
[0193] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means for implementing the functions specified in one or more flows Figure 1 one or more flows and / or blocks Figure 1 or means for implementing the functions specified in one or more blocks.
[0194] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more flows Figure 1 one or more flows and / or blocks Figure 1 or means for implementing the functions specified in one or more blocks.
[0195] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: the specific implementation manners of the present invention can still be modified or equivalently replaced, and any modification or equivalent replacement without departing from the spirit and scope of the present invention shall be covered by the protection scope of the claims of the present invention.
Claims
1. A data detection method for batch-stream fusion of multi-source time series data for power grid control, characterized in that: The method comprises: The source data end collects data from multiple business systems and sends the business system data to the data collection module according to the data type; The data collection module receives the business system data, encapsulates the business system data, the dimensional feature information of the business system data and the data time scale information into a business analysis message, and sends the business analysis message to the data detection engine; The data detection engine receives and parses the business analysis message to obtain the dimensional feature information and data time scale information of the business system data, determines the trigger mode according to the data time scale information, processes the business system data according to the trigger mode to obtain the scheduling operation data, and triggers the detection calculation component to detect the scheduling operation data to obtain the detection detailed data: when the data time scale information is an incremental data type or a reissued data type, and the trigger mode is real-time processing, the data detection engine processes the business analysis message in real time to obtain the scheduling operation data, triggers the detection calculation component to perform detection through the scheduling operation data, and obtains the detection detailed data; when the data time scale information is a stock historical data type, and the trigger mode is scheduled task processing, the data detection engine starts a scheduled task to process the business analysis message to obtain the scheduling operation data, triggers the detection calculation component to perform detection through the scheduling operation data, and obtains the detection detailed data.
2. The method according to claim 1, characterized in that The data collection module receives the business system data, encapsulates the business system data, the dimensional feature information of the business system data, and the data time scale information into a business analysis message, and sends the business analysis message to the data detection engine, specifically including: The data collection module determines data time scale information of the business system data according to the subject of the message encapsulating the business system data received, wherein the data time scale information includes an incremental data type, a reissued data type or a stock historical data type; After filtering the data whose primary key does not conform to the specification in the business system data, the data collection module determines whether the primary key of the business system data is the same as the primary key of the business data stored in the data block; If they are the same, the business system data is stored in the data block at the location of the stored business data with the same primary key and the stored business data is overwritten; Otherwise, storing the business system data in the data block; The data collection module encapsulates the business system data, the dimensional feature information of the business system data, the data time scale information and the triggering method into a business analysis message, and sends the business analysis message to the data detection engine.
3. The method according to claim 2, characterized in that When the data time scale information is an incremental data type or a reissued data type, and the trigger mode is real-time processing, the data detection engine processes the business analysis message in real time to obtain the scheduling operation data, and triggers the detection calculation component to perform detection through the scheduling operation data to obtain the detection detail data, specifically including: The data detection engine determines, based on the data time scale information, whether the service analysis message is an incremental data type or a supplementary data type; When the business analysis message is of an incremental data type, the data detection engine partitions the incremental data read from the business analysis message according to the dimensional feature information; The computing nodes of each data partition concurrently preprocess the allocated incremental data to obtain scheduling operation data, and after calling the detection computing component to calculate the scheduling operation data, obtain an aggregated computing result; The data detection engine triggers the computing nodes of the data partition to concurrently call the detection calculation component to perform secondary aggregation calculation on the first aggregation calculation result according to the first aggregation calculation result of all the incremental data, so as to obtain the detection detailed data; Among them, when the data detection engine is in the process of processing and calculation, the status information of each link is recorded, and when an abnormal state link occurs, the task corresponding to the abnormal state link is restarted for calculation; When the service analysis message is of the supplementary data type, the data detection engine reads the original data corresponding to the supplementary data from the database, and obtains the scheduling operation data after adding the supplementary data to the original data; The detection calculation component is called to calculate the scheduling operation data to obtain the detection detailed data.
4. The method according to claim 2, characterized in that: When the data time scale information is a stock historical data type and the trigger mode is scheduled task processing, the data detection engine starts the scheduled task to process the business analysis message to obtain the scheduling operation data, and triggers the detection calculation component to perform detection through the scheduling operation data to obtain the detection detail data, specifically including: The data detection engine sets a timed scheduling task for the stock historical data in the business analysis message according to the dimensional feature information; The scheduled scheduling task uses the inventory historical data as the scheduling operation data, calls the detection calculation component to calculate the scheduling operation data, and obtains the detection detailed data.
5. The method according to claim 2, characterized in that: The method further comprises: When the data detection engine processes the business analysis message in real time and starts the scheduled task to process the business analysis message and is triggered at the same time, the computing task is executed or suspended according to a preset scheduling strategy, wherein the preset scheduling strategy includes weight allocation according to the priority of the task.
6. The method according to claim 1, characterized in that The trigger detection calculation component detects the processed scheduling operation data to obtain detection detailed data, which specifically includes: The detection calculation component initiates a sub-detection component according to the detection dimension identification, and the sub-detection component includes an integrity sub-detection component, an accuracy sub-detection component, and a timeliness sub-detection component; The integrity sub-detection component detects the completeness of the operation data collection type, the completeness of the storage data record, the data storage start time and the data storage period of the scheduling operation data to obtain first detection result data; The accuracy sub-detection component detects the scheduling operation data according to the limit rule, the addition rule, the mutual exclusion rule, the comparison rule and the trend rule to obtain second detection result data; The timeliness sub-detection component detects the data transmission period of the scheduling operation data to obtain third detection result data; The detection detail data is obtained according to the first detection result data, the second detection result data and / or the third detection result data.
7. A data detection device for batch-stream fusion of multi-source time series data for power grid control, characterized in that: The device comprises a source data terminal, a data collection module and a data detection engine, wherein the data detection engine comprises a detection calculation component: The source data terminal is used to collect data from multiple business systems and send the business system data to the data collection module according to the data type; The data collection module is used to receive the business system data from the message bus, encapsulate the business system data, the dimensional feature information of the business system data and the data time scale information into a business analysis message, and send the business analysis message to the data detection engine; The data detection engine is used to receive and parse the business analysis message to obtain the dimensional feature information and data time scale information of the business system data, determine the trigger mode according to the data time scale information, process the business system data according to the trigger mode, obtain the scheduling operation data, and trigger the detection calculation component; The detection and calculation component is used to detect the scheduling operation data to obtain the detection detailed data: when the data time scale information is an incremental data type or a reissued data type, and the trigger mode is real-time processing, the data detection engine processes the business analysis message in real time to obtain the scheduling operation data, and triggers the detection and calculation component to perform detection through the scheduling operation data to obtain the detection detailed data; When the data time scale information is the type of stock historical data and the trigger method is scheduled task processing, the data detection engine starts the scheduled task to process the business analysis message to obtain the scheduling operation data, and triggers the detection calculation component to perform detection through the scheduling operation data to obtain the detection detail data.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the data detection method for batch-stream fusion of multi-source time series data for power grid regulation described in any one of claims 1 to 6 is implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the data detection method for batch-stream fusion of multi-source time series data for power grid regulation described in any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Dynamic resource scheduling method for space-time big data stream processing engine
CN115510100A
Custom rule engine system and judgment method based on equipment inspection work order and equipment data analysis
CN116048926A