Data panorama acquisition architecture based on full-period data center and 5G network
Through distributed architecture and data cleaning technology, panoramic collection and control of multiple types of data in the integrated energy system are achieved, solving the problem of data isolation in existing technologies, improving the system's data sharing and analysis capabilities, and reducing construction and maintenance costs.
Patent Information
- Application Number
- CN202510679319.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-26
- Publication Date
- 2025-09-23
AI Technical Summary
Existing technologies have failed to effectively integrate multiple types of data such as electricity, gas, cooling, heat, and clean energy, resulting in the construction of integrated energy systems at the perception layer, network layer, and platform layer being limited to the power system. Other energy networks such as cooling, heat, and gas have not been incorporated. There is a lack of panoramic data collection and control technology, making it difficult to support the expansion of the application layer and the improvement of comprehensive energy efficiency.
The distributed data panoramic acquisition architecture includes the device acquisition layer, communication layer, master station layer, storage layer and business layer. The device acquisition layer collects data in real time, the communication layer transmits data, the master station layer processes and stores data, and the business layer performs data cleaning and prediction. Fuzzy rules and neural networks are combined to perform data repair, thus achieving unified data conversion and secure transmission.
It realizes the panoramic integration of source, grid, load and storage data, reduces system monitoring costs, supports the sharing and analysis of multiple types of data, provides a flexible data interface and low maintenance costs, and realizes the seamless splicing of different equipment models and the accuracy of data cleaning.
Smart Images

Figure CN120692138A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a data panoramic acquisition architecture, and specifically to a data panoramic acquisition architecture based on a full-cycle data center and 5G network, which analyzes the needs of data panoramic acquisition and control technology, analyzes the multiple types of data sources involved in each link, determines the scope and content of data acquisition and control, clarifies information such as the source of data, acquisition frequency, and interaction method, and reduces system construction costs. Background Art
[0002] With the rapid development of information technology, data centers and 5G networks have become critical infrastructure supporting the operation of modern society. A comprehensive data collection system, capable of real-time, accurate, and comprehensive collection of data on energy usage and production, including electricity, gas, cooling, heat, clean energy, supply networks, various load types, and energy storage, provides a foundation for system analysis and decision-making. Data collection can also provide a better understanding of energy management. The regulation of data centers and 5G base stations remains largely limited to the complementary nature of multiple energy sources on both the supply and demand sides. Due to a lack of top-level mechanism design and a lack of universal information models, collection and metering equipment, and control systems, the development of integrated energy systems at the perception, network, and platform levels remains confined to the power system, failing to incorporate other energy networks such as cooling, heating, and gas. The platform and hub attributes of the power grid remain largely unrecognized, and the operational and control mechanisms across the power source, grid, load, and storage links remain largely unconnected. Furthermore, the key software and hardware support for the perception, network, and platform layers underlying energy management within data centers and 5G base stations is relatively weak, making it difficult to support the development and expansion of the application layer. Therefore, there is an urgent need to explore a data panoramic collection and control technology based on full-cycle data centers and 5G networks to solve the problem of lack of technical support for comprehensive energy regulation of source-grid-load-storage interaction, improve the system's comprehensive energy efficiency and clean energy consumption level, and support the construction and development of new information infrastructure. Summary of the Invention
[0003] In response to the above problems, the main purpose of the present invention is to provide a data panoramic acquisition architecture based on a full-cycle data center and 5G network, which analyzes the needs of data panoramic acquisition and control technology, analyzes the multiple types of data sources involved in each link, determines the scope and content of data acquisition and control, clarifies information such as the source of data, acquisition frequency, and interaction method, and reduces system construction costs.
[0004] The present invention solves the above technical problems through the following solution: a data panoramic acquisition architecture based on a full-cycle data center and 5G network, which is a distributed architecture and includes a device acquisition layer, a communication layer, a master station layer, a storage layer and a business layer.
[0005] The equipment acquisition layer is the acquisition terminal, responsible for collecting and reporting data; the communication layer is responsible for communication between the equipment acquisition layer and the master station layer; the master station layer is responsible for communicating with the acquisition device, summarizing and collecting data, and processing the data and storing it in the warehouse; the storage layer is responsible for the storage system to collect measurement data and archival structured data; the business layer implements data cleaning preprocessing and situation prediction functions, and realizes the upload of measurement data and the download of control instructions from the equipment acquisition end through edge computing control technology.
[0006] In a specific implementation example of the present invention, the device collection layer collects data in real time; the collected data is ensured to correspond one-to-one with the actual status data, the daily loss rate of the collected data is no higher than 0.1%, and the data security is ensured by the link encryption protocol; it supports data storage above PB level, and uses a non-relational database for real-time device message storage, and this type of database supports mixed row and column storage.
[0007] In a specific implementation example of the present invention, the communication layer uses network components to manage various network services and is responsible for receiving / sending messages; uses a device gateway to connect devices and push device messages to the event bus; uses the event bus to forward data within the process, forwarding data to the database or pushing messages.
[0008] In a specific implementation example of the present invention, the device acquisition layer and the communication layer involve industrial Ethernet and device Internet of Things communication protocols.
[0009] The equipment collection layer and communication layer also include the integration of data models: the communication object server can access subsystems of various different data models, and the communication object server integrates these different data models to unify the comprehensive energy system information model.
[0010] The equipment acquisition layer and communication layer also include the extension of the data model: the communication object server expands the data model of the online monitoring field according to the extension principle of the IEC61850 standard.
[0011] The device acquisition layer and communication layer also include data model conversion: the communication object server converts the data model as needed, including: conversion between models based on the IEC61850 standard and models based on traditional linear point tables; conversion between models based on the IEC61850 standard and models based on the IEC61970 standard.
[0012] The device acquisition layer and communication layer also include the conversion of communication protocols: the communication object server converts the general communication protocol outside the station according to different data models: models and interfaces based on the IEC61850 standard and models and interfaces based on the IEC61970 standard.
[0013] The equipment acquisition layer and communication layer also need to collect the wind turbine measurement data together with the video data after image processing.
[0014] In a specific implementation example of the present invention, the master station layer mainly includes five modules: front-end service, buffer, stream computing, data storage, and business management, as follows:
[0015] (1) The front-end service implements the parsing and upward forwarding of uplink messages, as well as the encapsulation and issuance of downlink terminal control commands, and has the characteristics of high concurrency, strong timing, and large-scale data processing.
[0016] (2) The buffer is responsible for buffering the large batches of message data forwarded by the front-end service, shielding the instantaneous pressure of data peaks on subsequent programs, and distributing the data to the stream computing component for processing according to the strategy; the buffer adopts a message queue and is implemented based on the message queue component of the big data platform.
[0017] (3) The stream computing module is responsible for performing real-time computing on the collected data distributed by the buffer, realizing data preprocessing, real-time alarm, and event processing; the stream computing function is implemented based on the stream computing component of the big data platform.
[0018] (4) The data storage module realizes the storage of measurement data and archival data into the analysis domain and the temporary storage of short-term measurement data.
[0019] (5) The business management module is responsible for managing the main station system.
[0020] In a specific implementation example of the present invention, the storage layer includes two modes: a row storage scheme to store device data and a column storage scheme to store device data.
[0021] In a specific implementation example of the present invention, a row-based storage solution stores device data, and each attribute value of the device is saved as an index record. Its typical application scenario is that the device only reports a part of the attributes each time, and supports reading part of the attribute data at the same time; a column-based storage solution stores device data, with one attribute as a column and one attribute message as an index record. It is suitable for scenarios where the device reports all attribute values each time.
[0022] In a specific implementation example of the present invention, the business layer specifically includes: an integrated energy system data cleaning process, and multi-source data repair combining fuzzy rules and neural networks.
[0023] (I) Integrated energy system data cleaning process: To identify and repair abnormal data caused by sensor failure, the following steps are included:
[0024] Step 1: Time series data construction;
[0025] After preprocessing the data, relevant analysis and calculation are performed on the preprocessed data. Based on the data list during preprocessing, the data is imported into the memory space using a program to form a calculation matrix or list data format. Based on the preprocessing results, the horizontal and vertical marks of each data are used to extract it from the total table for direct calculation, which can index not only each data element, but also each other type of element. In order to facilitate subsequent indicator calculation and feature analysis, based on these preprocessed data, a time series is first constructed using a program.
[0026] The so-called time series data, that is, data with a time stamp, is necessary data for studying the trend of each variable changing over time, and is also necessary data for calculating parameters such as volatility and load rate. Therefore, it is necessary to construct time series data for each variable within the research time scale based on historical data; the voltage of each topological node, the power at the I end of each line, the active output of each generator and other information are all regarded as a variable. A file of power grid source data records the information at this section at a time section, so a file cannot construct time series data. It is only a storage file for the variable data of the entire network at a time section. Therefore, to construct a time series, it is necessary to traverse all QS files; each file only provides a value for the construction of a sequence of a specific variable, which needs to be extracted in sequence from the data files continuously recorded by the power grid. Assuming that the data at the first time point is taken out from the first file, the data at the second time point needs to be extracted from the second file, and so on.
[0027] Through the above steps, the massive historical data of the integrated energy system can be converted into time series data that can be included in program calculation and analysis for further mining.
[0028] Step 2: Association rule mining: Use the generalized association rules (GAR) algorithm to analyze the support and confidence between the data sequences of the integrated energy system and mine the data sequences with higher correlation.
[0029] Step 3: Abnormal data detection: Use the abnormal data detection technology based on the sliding time window to detect each data sequence of the integrated energy system. If abnormal data exists, perform abnormal data repair in step 4.
[0030] Step 4: Abnormal data repair: When there is an associated sequence of abnormal data, and only one data in the associated sequence can be abnormal within a time period, the abnormal data is repaired by combining fuzzy rules and neural network multi-source data repair technology.
[0031] (1) If the abnormal data does not have other data sequences with a high correlation, it is impossible to confirm whether the abnormal data is caused by sensor failure or equipment abnormality. In this case, data cleaning is not performed, and an "abnormal cause to be determined" warning signal is issued. The operation and maintenance personnel and inspection personnel manually determine the cause of the abnormal data.
[0032] (2) If abnormal data appears in multiple data sequences with a high degree of correlation at a similar time, it is determined that the abnormal data is caused by equipment abnormality. At this time, data cleaning is not performed and an "equipment abnormality" warning signal is issued, and maintenance personnel can perform maintenance work on the equipment.
[0033] (3) If only a single data sequence among multiple data sequences with a high degree of correlation produces an anomaly, it is determined that the abnormal data is caused by a sensor failure. At this time, these data sequences are input into the multi-source data repair technology that combines fuzzy rules and neural networks for repair.
[0034] (2) Multi-source data repair combining fuzzy rules and neural networks includes the following steps:
[0035] Step 1: Obtain fuzzy business rules through business research and rewrite them in the form of first-order logic; the language description of the operating rules should be fuzzy rules.
[0036] Step 2: Design and train the indicator rule satisfaction classifier offline.
[0037] Design an indicator rule satisfaction classifier, whose structure is a neural network, divided into three parts: input part, indicator part, and output part. Its structure is as follows:
[0038] (1) The input part is all the data of the integrated energy system at the same time.
[0039] (2) The indicator part is the category possibility of all indicators involved in the fuzzy business rules in step 1.
[0040] (3) The output part is the degree of violation of each fuzzy rule. The neural network design from the indicator part to the output part is designed according to the first-order logic representation of each fuzzy rule. The number of neurons in the output part is the number of business rules; the details are as follows:
[0041] The intersection operation is designed to connect all neurons involved in the intersection operation to a new neuron with a connection coefficient of 1 and a bias of -n, where n is the number of neurons involved in the intersection operation.
[0042] (4) The loss function of the trainer is the sum of the squares of the prediction results, and training is performed using the gradient descent method. Since the above loss function and neuron calculations are composed of sigmoid functions and basic operation symbols, they can be differentiated, and the gradient descent method is feasible.
[0043] The training data is the historical non-abnormal data of the integrated energy system. The output result mark value of the data at each moment is set to 0, which means that the violation of the fuzzy rules by the non-abnormal data is 0.
[0044] Step 3: Use the wavelet neural network with momentum online to predict the data to be repaired using non-abnormal data that has a high correlation with the data to be repaired.
[0045] The predicted result replaces the abnormal data and is input into the offline trained indicator rule satisfaction classifier. If the value of the loss function is lower than the set threshold, the predicted result is used as the repair value; if the value of the loss function is higher than the set threshold, it is determined that the abnormal data is caused by equipment abnormality. At this time, no data cleaning is performed, and an "equipment abnormality" warning signal is issued, allowing maintenance personnel to perform equipment maintenance work.
[0046] The positive progress of the present invention is that the data panoramic collection architecture based on the full-cycle data center and 5G network provided by the present invention has the following advantages:
[0047] (1) It realizes the organic integration of steady-state, transient and dynamic data of source, grid, load and storage, as well as panoramic data of equipment status, images and operating conditions, which facilitates the sharing of various system resources and reduces system monitoring costs.
[0048] (2) Convert all collected data into data objects of the integrated energy system information model.
[0049] (3) The seamless connection between the secondary equipment model following IEC61850 SCL and the primary equipment model following IEC61970 CIM is achieved.
[0050] (4) A remote service interface based on the integrated energy system information model is provided, and third-party analysis programs can easily obtain and analyze the collected data, providing convenience for related application work, flexible deployment, and low maintenance costs. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 This is a flow chart of the overall architecture design of the data acquisition solution in this invention.
[0052] Figure 2 Schematic diagram of the system control topology in the present invention.
[0053] Figure 3 This is a schematic diagram of the network topology relationship of the equipment collection in the present invention.
[0054] Figure 4 The overall process logic diagram constructed for the time series in this invention.
[0055] Figure 5 This is a flow chart of the integrated energy system data cleaning process in the present invention.
[0056] Figure 6 This is a specific schematic diagram of the intersection operation in the multi-source data repair technology combining fuzzy rules and neural networks in the present invention.
[0057] Figure 7 This is a schematic diagram of the multi-source data repair technology that combines fuzzy rules and neural networks in the present invention and performs specific operations.
[0058] Figure 8 This is a non-computational schematic diagram of the multi-source data repair technology that combines fuzzy rules and neural networks in the present invention.
[0059] Figure 9 This is a specific flow chart of the multi-source data repair technology combining fuzzy rules and neural networks in the present invention. DETAILED DESCRIPTION
[0060] The preferred embodiments of the present invention are given below in conjunction with the accompanying drawings to illustrate the technical solutions of the present invention in detail.
[0061] The application principle of the present invention will be further described below with reference to the accompanying drawings and specific embodiments.
[0062] Figure 1 This is a flow chart of the overall architecture design of the data acquisition solution in the present invention. Figure 1 As shown, this invention targets a specific data collection architecture design, opting for a distributed architecture. The overall architecture is primarily divided into a device collection layer, a communication layer, a master station layer, a storage layer, and a business layer. The architecture achieves the goal of "collect once, store in one place, and use everywhere." Collected data is received, pushed, and processed in real time via a gateway before ultimately being stored in a database using a combination of Redis, Elasticsearch, and MySQL.
[0063] The device collection layer is mainly composed of collection terminals (i.e. sensors), which are responsible for collecting and reporting data. The data collection has real-time, accuracy, stability, security and data storage requirements, as follows:
[0064] (1) Real-time requirements: For real-time data collection, the required data is collected with a collection cycle of 1 second; for data with a slower update rate, such as equipment power consumption data, a collection cycle of 15 minutes can be used.
[0065] (2) Accuracy requirements: The collected data must correspond one-to-one with the actual status data to ensure the accuracy of the transmitted data.
[0066] (3) Stability requirement: The daily loss rate of collected data shall not exceed 0.1%.
[0067] (4) Security requirements: Ensure data security through link encryption protocols, etc.
[0068] (5) Data storage requirements: Support data storage above PB level, and use non-relational databases for real-time device message storage. Such databases also support mixed row and column storage.
[0069] The communication layer is responsible for the communication between the collection layer device and the main station. The communication layer uses network components to manage various network services (MQTT, TCP, etc.), and is only responsible for receiving / sending messages, and does not process logic; uses network components to manage various network services (MQTT, TCP, etc.), and is only responsible for receiving / sending messages, and does not process logic; uses device gateways to connect devices and push device messages to the event bus; uses the event bus for data forwarding within the process, and can forward data to the database or push messages on WeChat, etc.
[0070] The communication basically adopts IEEE802.15.4, namely ZigBEE.
[0071] The device acquisition layer and communication layer involve technologies such as industrial Ethernet and device IoT communication protocols. To meet the needs of the model, the following functions need to be implemented:
[0072] (1) Data model integration: The communication object server can access subsystems of various data models. The communication object server integrates these different data models to unify the integrated energy system information model;
[0073] (2) Data model expansion: In the integrated energy system information model, the field of substation online monitoring is rarely involved. The communication object server expands the data model of the online monitoring field according to the extension principle of the IEC61850 standard;
[0074] (3) Data model conversion: The communication partner server can convert the data model as needed, including conversion between the IEC61850 standard model and the traditional linear point table model; conversion between the IEC61850 standard model and the IEC61970 standard model;
[0075] (4) Communication protocol conversion: The communication object server can convert the general communication protocol outside the station according to different data models: models and interfaces based on the IEC61850 standard and models and interfaces based on the IEC61970 standard.
[0076] (5) In addition, the wind turbine measurement data needs to be collected together with the video data after image processing. Figure 3 It shows the device collection network topology diagram.
[0077] The master station layer is responsible for communicating with the acquisition device, aggregating and collecting data, processing the data (especially image data and video data) and storing it in the database. The application master station layer mainly includes five modules: front-end service, buffer, stream computing, data storage, and business management. The details are as follows:
[0078] (1) The front-end service implements the parsing and upward forwarding of uplink messages, as well as the encapsulation and issuance of downlink terminal control commands, and has the characteristics of high concurrency, strong timing, and large-scale data processing.
[0079] (2) The buffer is responsible for buffering large amounts of message data forwarded by the front-end service, shielding the instantaneous pressure of data surges on subsequent programs, and distributing data to the stream computing component for processing according to the strategy. The buffer uses a message queue and is implemented based on the message queue component of the big data platform.
[0080] (3) The stream computing module is responsible for real-time computing of the collected data distributed by the buffer, mainly realizing business functions such as data preprocessing (including images and videos), real-time alarms, and event processing. The stream computing function is implemented based on the stream computing component of the big data platform.
[0081] (4) The data storage module realizes the storage of measurement data and archival data into the analysis domain and the temporary storage of short-term measurement data.
[0082] (5) The business management module is responsible for managing the main station system.
[0083] The storage layer is responsible for storing the measurement data and archival structured data collected by the system.
[0084] (1) The row storage solution stores device data. Each attribute value of the device is saved as an index record. Its typical application scenario is when the device only reports a part of the attributes at a time and supports reading part of the attribute data. The advantage is that it can almost meet the attribute data storage requirements in any scenario. The disadvantage is that when the number of device attributes is large, the data volume increases exponentially, which may affect performance.
[0085] (2) The columnar storage solution stores device data, with one attribute as a column and one attribute message as an index record. This solution is suitable for scenarios where the device reports all attribute values every time. The advantage is that the system performance is higher when there are many attributes and the device reports all attributes every time. The disadvantage is that the device must report all attributes.
[0086] The business layer implements functions such as data cleaning and preprocessing, situation prediction, and uses edge computing control technology to upload measurement data from the equipment collection end and transmit control instructions. The system control topology relationship is as follows: Figure 2 shown.
[0087] The three key technologies provided by the present invention are as follows:
[0088] 1: Dynamic loading protocol processing technology.
[0089] This technology is a key function of data acquisition. The protocol processing module of data acquisition can dynamically load and unload protocols according to the configuration and supports online operation, that is, it can add, delete and modify new protocols without compiling or restarting.
[0090] Third-party developers can implement protocol extension functionality by compiling the corresponding protocol processing code based on the interface specifications, forming a dynamic link library. The protocol processing module for data acquisition can dynamically load and unload protocols based on configuration, supporting online operation, that is, adding, deleting, and modifying new protocols without compiling or restarting the system.
[0091] 2: Security control technology based on software and hardware encryption.
[0092] This technology is the main external data interface of the smart grid dispatching and control system basic platform, responsible for the system's external data exchange.
[0093] In order to enhance the security of the system control command process, a security authentication mechanism is added to each link of the processing and transmission process in the system control process.
[0094] When data acquisition receives control commands from the application through the secure message bus, it must first undergo security authentication, then convert the specific content into a protocol message, and then correctly send the information to the plant station through a secure encrypted tunnel.
[0095] Similarly, the control command feedback from the plant side received through the encrypted tunnel is processed through data collection and the results are returned to the relevant applications through the secure message bus.
[0096] 3: Multi-machine load balancing and multi-source data processing technology.
[0097] This technology involves the collection management module assigning collection tasks to each normally functioning collection server based on the load balancing principle. If a collection server experiences an abnormality, the collection management module can automatically assign its collection tasks to other collection servers to ensure that data is not lost.
[0098] In order to solve the above-mentioned problems existing in the existing technology (i.e., conventional data-driven methods have high requirements on data quality and quantity, rule-driven methods require clear rules, and integrated energy systems can often only provide fuzzy rules), a data cleaning and repair method for integrated energy systems that can improve the accuracy and reliability of data cleaning is provided, which considers combining fuzzy rules and neural networks.
[0099] The raw data of the integrated energy system inevitably has problems such as missing, redundancy, conflict, error, and anomaly. It must be preprocessed to eliminate useless data, retain useful data, and adjust the format to make it time series data that is easily accepted by the program.
[0100] Therefore, the present invention implements a data cleaning method that combines rules and machine learning, combining the rules in the integrated energy system with machine learning technologies such as neural networks to achieve safe machine learning and complete the processing of abnormal data.
[0101] "Rules" are logical rules in the form of "if..., then...", with clear semantics, which can describe the objective laws or domain concepts implied by data distribution.
[0102] Rule learning involves learning a set of rules from training data that can be used to distinguish unseen instances. There have been many studies in rule learning (sequential covering, pruning, first-order inductive learners, and inductive logic programming). However, these methods use Boolean logic and are not suitable for large-scale problems.
[0103] Instead of interpreting clauses using Boolean logic, Probabilistic Soft Logic (PSL) interprets clauses using Lukasiewicz logic, which extends Boolean logic to the continuous interval [0, 1]. Extending truth values to a continuous domain allows them to represent ambiguous concepts, as they are often neither completely true nor completely false.
[0104] The logical operators of PSL include ∧ (conjunction), ∨ (disjunction) and (negation), the logical operation is as follows:
[0105] x1∧x2=max{x1+x2-1,0}
[0106] x1∨x2=max{x1+x2,1}
[0107]
[0108] The specific technical solution of the present invention consists of an integrated energy system data cleaning process and a multi-source data repair technology that combines fuzzy rules and neural networks. The specific process is as follows:
[0109] 1. The integrated energy system data cleaning process aims to identify and repair abnormal data caused by sensor failures, including the following steps:
[0110] Step 1: Time series data construction.
[0111] After preprocessing the data, correlation analysis and calculations can be performed based on the preprocessed data. Based on the preprocessed data list, the program can be used to import it into memory space, forming a matrix or list data format that can be calculated. Based on the preprocessed results, the horizontal and vertical labels of each data point can be used to extract it from the total table and perform direct calculations. Not only can it index into each data element, but it can also index into every other type of element. To facilitate subsequent indicator calculations and feature analysis, the program first constructs a time series based on this preprocessed data.
[0112] The so-called time series data, that is, data with a time stamp, is necessary data for studying the trend of each variable changing over time, and is also necessary data for calculating parameters such as volatility and load rate. Therefore, it is necessary to construct the time series data of each variable within the time scale of the study based on historical data. The voltage of each topological node, the power at the I end of each line, the active output of each generator and other information are all taken as a variable. A file of power grid source data records the information at this section at a time section, so a file cannot construct time series data. It is only a storage file for the variable data of the entire network at a time section. Therefore, to construct a time series, it is necessary to traverse all QS files. Each file only provides a value for the construction of a sequence of a specific variable, which needs to be extracted in sequence from the data files continuously recorded by the power grid. Assuming that the data at the first time point is taken out from the first file, the data at the second time point needs to be extracted from the second file, and so on. The overall process logic diagram of time series construction is as follows: Figure 4 shown.
[0113] Through the above steps, the massive historical data of the integrated energy system can be converted into time series data that can be included in program calculation and analysis for further mining.
[0114] Step 2: Association rule mining: Use the generalized association rules (GAR) algorithm to analyze the support and confidence between the data sequences of the integrated energy system and mine data sequences with high correlation.
[0115] Step 3: Abnormal data detection. Use the abnormal data detection technology based on the sliding time window to detect each data sequence of the integrated energy system. If abnormal data exists, proceed to step 4 to repair the abnormal data.
[0116] Step 4: Abnormal data repair. When the abnormal data has an associated sequence, and the associated sequence can only have one abnormal data in a sequence within a time period, the abnormal data is repaired by combining fuzzy rules and neural network multi-source data repair technology:
[0117] (1) If the abnormal data does not have other data sequences with a high correlation, it is impossible to confirm whether the abnormal data is caused by sensor failure or equipment abnormality. In this case, data cleaning is not performed, and an "abnormal cause to be determined" warning signal is issued. The operation and maintenance personnel and inspection personnel manually determine the cause of the abnormal data.
[0118] (2) If abnormal data appears in multiple data sequences with a high degree of correlation at a similar time, it is determined that the abnormal data is caused by equipment abnormality. At this time, data cleaning is not performed and an "equipment abnormality" warning signal is issued, and maintenance personnel can perform maintenance work on the equipment.
[0119] (3) If only a single data sequence among multiple data sequences with a high degree of correlation produces an anomaly, it is determined that the abnormal data is caused by a sensor failure. At this time, these data sequences are input into the multi-source data repair technology that combines fuzzy rules and neural networks for repair.
[0120] The process of data cleaning of integrated energy system is shown in Figure 5 .
[0121] 2. The multi-source data repair technology combining fuzzy rules and neural networks includes the following steps:
[0122] Step 1: Obtain fuzzy business rules through business research and rewrite them in first-order logic. The operating rule language description should be fuzzy rules.
[0123] For example, "When indicator A is high and indicator B is high, indicator C is high" can be rewritten as
[0124] "When indicator A is at a high level, indicator C is at a medium level" can be rewritten as
[0125] "When indicator A is high or indicator B is high, indicator C is high" can be rewritten as
[0126] "When indicator A is at a high level, indicator C is not at a high level" can be rewritten as
[0127] Step 2: Design and train the indicator rule satisfaction classifier offline.
[0128] Design an indicator rule satisfaction classifier. Its structure is a neural network, which is divided into three parts: input part, indicator part, and output part (each part contains a hidden layer). Its structure is as follows:
[0129] (1) The input part is all the data of the integrated energy system at the same time, such as "current", "temperature", "power", "voltage", "flow rate", and "pressure";
[0130] (2) The indicator part is the category possibility of all indicators involved in the fuzzy business rules in step 1 (taking the above data as an example, this layer will have "current is high level", "current is medium level", "current is low level", "temperature is high level", "temperature is medium level", "temperature is low level", "power is high level", "power is medium level", "power is low level", "voltage is high level", "voltage is medium level", "voltage is low level", "flow rate is high level", "flow rate is medium level", "flow rate is low level", "pressure is high level", "pressure is medium level", "pressure is low level"). The higher the value, the greater the possibility. For example, if "current is high level", "current is medium level", and "current is low level" are 10, 6, and -1 respectively, it means that the probability of current being high level is greater than the probability of current being medium or low level.
[0131] In order to make the indicator part able to express its meaning more objectively in the initial situation, the connection weight between the input layer "data A" and the hidden layer "A is a high level" in the neural network is set to 1, the connection weight between the input layer "data A" and the hidden layer "A is a medium level" is set to 0, and the connection weight between the input layer "data A" and the hidden layer "A is a low level" is set to -1. The purpose of this is to enable the neural network to objectively meet the objective law that "the larger the observation value, the greater the probability of it being a high level, and the smaller the observation value, the greater the probability of it being a low level" when it is not trained.
[0132] (3) The output part is the degree of violation of each fuzzy rule. The neural network design from the indicator part to the output part is designed according to the first-order logic representation of each fuzzy rule. The number of neurons in the output part is the number of business rules. The details are as follows:
[0133] The design of the intersection operation is to connect all neurons involved in the intersection operation to a new neuron, with a connection coefficient of 1 and a bias of -n, where n is the number of neurons involved in the intersection operation. The intersection operation diagram is shown in the figure below: Figure 6 , the new neuron output satisfies the following formula:
[0134]
[0135] This formula can make the calculated value as low as possible because the probability of multiple first-order logic intersections is lower than the probability of a single first-order logic. a0 and b0 are the parameters of the Sigmoid function.
[0136] The design of the union operation is to connect all neurons involved in the intersection operation to a new neuron, with a connection coefficient of 1 and a bias of 0. The diagram of the union operation is shown in Figure 7 , the new neuron output satisfies the following formula:
[0137]
[0138] This formula can make the calculated value as high as possible because the probability of multiple first-order logic values is higher than the probability of a single first-order logic value. a1 and b1 are the parameters of the Sigmoid function.
[0139] The design of the non-operation is to connect the neurons involved in the intersection operation to a new neuron, with a connection coefficient of -1 and a bias of 0. The schematic diagram of the non-operation is shown in Figure 8 , the new neuron output satisfies the following formula:
[0140]
[0141] This formula allows the calculated value to be the difference between 1 and the original value, because the sum of the probabilities of a single first-order logic event and its non-event is 1.
[0142] in The probability value of each first-order logic is generally calculated using the sigmoid function to keep it between [0, 1].
[0143] The design of basic operations follows the conventional method of neural network.
[0144] The smaller the output, the less the classification result violates the rules.
[0145] (4) The loss function of the trainer is the sum of the squares of the prediction results, and training is performed using the gradient descent method. Since the above loss function and neuron calculations are composed of sigmoid functions and basic operation symbols, they can be differentiated, and the gradient descent method is feasible.
[0146] The training data is the historical non-abnormal data of the integrated energy system. The output result mark value of the data at each moment is set to 0, which means that the violation of the fuzzy rules by the non-abnormal data is 0.
[0147] Step 3: Use the wavelet neural network with momentum online to predict the data to be repaired using non-abnormal data that has a high correlation with the data to be repaired.
[0148] The predicted result (repaired abnormal data) replaces the abnormal data and is input into the offline trained indicator rule satisfaction classifier. If the value of the loss function is lower than the set threshold (generally set to 0.05), the predicted result is used as the repair value; if the value of the loss function is higher than the set threshold, it is determined that the abnormal data is caused by equipment abnormality. At this time, no data cleaning is performed, and an "equipment abnormality" warning signal is issued, and maintenance personnel can perform maintenance work on the equipment.
[0149] The specific process of multi-source data repair technology combining fuzzy rules and neural networks is shown in Figure 9 .
[0150] The integrated energy system data cleaning process provided by this technology takes into account various possibilities of data anomalies in the integrated energy system, and sets up four processing solutions, including "no operation required", "abnormal cause to be determined", "equipment abnormality", and "multi-source data abnormality repair". Before cleaning the data, the cause of the abnormality is analyzed to determine the processing solution, avoiding multi-source data repair of all abnormal data and improving system efficiency.
[0151] In summary, this application has the following beneficial effects:
[0152] (1) It realizes the organic integration of steady-state, transient and dynamic data of source, grid, load and storage, as well as panoramic data of equipment status, images and operating conditions, which facilitates the sharing of various system resources and reduces system monitoring costs.
[0153] (2) Convert all collected data into data objects of the integrated energy system information model.
[0154] (3) The seamless connection between the secondary equipment model following IEC61850 SCL and the primary equipment model following IEC61970 CIM is achieved.
[0155] (4) A remote service interface based on the integrated energy system information model is provided, and third-party analysis programs can easily obtain and analyze the collected data, providing convenience for related application work, flexible deployment, and low maintenance costs.
[0156] The basic principles, main features and advantages of the present invention are shown and described above. It should be understood by those skilled in the art that the present invention is not limited to the above embodiments. The above embodiments and descriptions are merely illustrative of the principles of the present invention. Various changes and modifications may be made to the present invention without departing from the spirit and scope of the present invention. Such changes and modifications are intended to fall within the scope of the present invention, which is defined by the appended claims and their equivalents.
Claims
1. A full-cycle data center and 5G network-based data panorama collection architecture, characterized by: The architecture is a distributed architecture, which includes a device acquisition layer, a communication layer, a master station layer, a storage layer, and a business layer. The equipment collection layer is the collection terminal, responsible for collecting and reporting data; The communication layer is responsible for the communication from the equipment acquisition layer to the master station layer; The master station layer is responsible for communicating with the collection device, aggregating and collecting data, and processing the data and storing it in the database; The storage layer is responsible for storing the measurement data collected by the system and archival structured data; The business layer implements data cleaning and preprocessing, and situation prediction functions, and uses edge computing control technology to upload measurement data and download control instructions from the equipment collection end.
2. The data panorama collection architecture based on a full-cycle data center and 5G network according to claim 1 is characterized by: The device collection layer collects data in real time; the collected data is ensured to correspond one-to-one with the actual status data, and the daily loss rate of collected data is no higher than 0.1%. The data security is ensured through the link encryption protocol; it supports data storage above PB level and uses a non-relational database for real-time device message storage. This type of database supports mixed row and column storage.
3. The data panorama collection architecture based on a full-cycle data center and 5G network according to claim 1 is characterized by: The communication layer uses network components to manage various network services and is responsible for receiving / sending messages; Use the device gateway to connect devices and push device messages to the event bus; use the event bus to forward data within the process, forward data to the database or push messages.
4. The data panorama acquisition architecture based on a full-cycle data center and 5G network according to claim 1 is characterized by: The equipment acquisition layer and communication layer involve industrial Ethernet and equipment Internet of Things communication protocols; The equipment acquisition layer and communication layer also include the integration of data models: the communication object server can access subsystems with different data models, and the communication object server integrates these different data models to unify the comprehensive energy system information model; The device acquisition layer and communication layer also include the extension of the data model: the communication object server extends the data model in the online monitoring field according to the extension principles of the IEC61850 standard; The device acquisition layer and communication layer also include data model conversion: the communication object server converts the data model as needed, including conversion between models based on the IEC61850 standard and models based on traditional linear point tables; Conversion between models based on IEC61850 and models based on IEC61970; The equipment acquisition layer and communication layer also include the conversion of communication protocols: the communication object server converts the general communication protocol outside the station according to different data models: models and interfaces based on the IEC61850 standard and models and interfaces based on the IEC61970 standard; The equipment acquisition layer and communication layer also need to collect the wind turbine measurement data together with the video data after image processing.
5. The data panorama collection architecture based on a full-cycle data center and 5G network according to claim 1 is characterized by: The main station layer mainly includes five modules: front-end service, buffer, stream computing, data storage, and business management. The details are as follows: (1) The front-end service implements the parsing and upward forwarding of uplink messages, as well as the encapsulation and issuance of downlink terminal control commands, and has the characteristics of high concurrency, strong timing, and large-scale data processing; (2) The buffer is responsible for buffering the large batches of message data forwarded by the front-end service, shielding the instantaneous pressure of data surges on subsequent programs, and distributing the data to the stream computing component for processing according to the strategy; the buffer adopts a message queue and is implemented based on the message queue component of the big data platform; (3) The stream computing module is responsible for performing real-time computing on the collected data distributed by the buffer, realizing data preprocessing, real-time alarming, and event processing; the stream computing function is implemented based on the stream computing component of the big data platform; (4) The data storage module realizes the storage of measurement data and archival data into the analysis domain and the temporary storage of short-term measurement data; (5) The business management module is responsible for managing the main station system.
6. The data panorama collection architecture based on a full-cycle data center and 5G network according to claim 1 is characterized by: The storage layer includes two modes: row storage scheme to store device data and column storage scheme to store device data.
7. The data panorama collection architecture based on a full-cycle data center and 5G network according to claim 6 is characterized in that: The row-based storage solution stores device data. Each device attribute value is saved as an index record. Its typical application scenario is that the device only reports a portion of the attributes at a time and supports reading some attribute data. The columnar storage solution stores device data, with one attribute as a column and one attribute message as an index record. This solution is suitable for scenarios where the device reports all attribute values every time.
8. The data panorama collection architecture based on a full-cycle data center and 5G network according to claim 1 is characterized by: The business layer specifically includes: integrated energy system data cleaning process, multi-source data repair combining fuzzy rules and neural networks; (I) Integrated energy system data cleaning process: To identify and repair abnormal data caused by sensor failure, the following steps are included: Step 1: Time series data construction; After preprocessing the data, correlation analysis and calculation are performed on the preprocessed data. Based on the data list during preprocessing, the program is used to import it into the memory space to form a calculation matrix or list data format. Based on the preprocessing results, the horizontal and vertical marks of each data are used to extract it from the total table and perform direct calculations. Not only can it index each data element, but it can also index each other type of element. To facilitate subsequent indicator calculation and feature analysis, based on these preprocessed data, the program is first used to construct a time series. The so-called time series data, that is, data with a time stamp, is necessary data for studying the trend of each variable changing over time, and is also necessary data for calculating parameters such as volatility and load rate. Therefore, it is necessary to construct the time series data of each variable within the research time scale based on historical data; the voltage of each topological node, the power at the I end of each line, the active output of each generator and other information are all regarded as a variable. A file of power grid source data records the information at this section at a time section, so a file cannot construct time series data. It is only a storage file for the variable data of the entire network at a time section. Therefore, to construct a time series, it is necessary to traverse all QS files; each file only provides a value for the construction of a sequence of a specific variable, which needs to be extracted in sequence from the data files continuously recorded by the power grid. Assuming that the data at the first time point is taken out from the first file, the data at the second time point needs to be extracted from the second file, and so on; Through the above steps, the massive historical data of the integrated energy system can be converted into time series data that can be included in the program calculation and analysis, so as to conduct further mining; Step 2: Association rule mining: Use the generalized association rules (GAR) algorithm to analyze the support and confidence between the data sequences of the integrated energy system and mine data sequences with high correlation; Step 3: Abnormal data detection: Use the abnormal data detection technology based on the sliding time window to detect each data sequence of the integrated energy system. If abnormal data exists, perform abnormal data repair in step 4; Step 4: Abnormal data repair: When there is an associated sequence of abnormal data, and only one data in the associated sequence can be abnormal within a time period, the abnormal data is repaired by combining fuzzy rules and neural network multi-source data repair technology: (1) If the abnormal data does not have other data sequences with a high correlation, it is impossible to confirm whether the abnormal data is caused by sensor failure or equipment abnormality. In this case, no data cleaning is performed and an "abnormal cause to be determined" warning signal is issued. The operation and maintenance personnel and inspection personnel manually determine the cause of the abnormal data. (2) If abnormal data appears in multiple data series with a strong correlation at a similar time, it is determined that the abnormal data is caused by equipment abnormality. In this case, no data cleaning is performed and an "equipment abnormality" warning signal is issued, so that maintenance personnel can carry out equipment maintenance work; (3) If only a single data sequence among multiple data sequences with strong correlation produces an anomaly, it is determined that the abnormal data is caused by sensor failure. In this case, these data sequences are input into the multi-source data repair technology that combines fuzzy rules and neural networks for repair; (2) Multi-source data repair combining fuzzy rules and neural networks includes the following steps: Step 1: Obtain fuzzy business rules through business research and rewrite them in the form of first-order logic; the language description of the operating rules should be fuzzy rules; Step 2: Design and offline train the indicator rule satisfaction classifier; Design an indicator rule satisfaction classifier, whose structure is a neural network, divided into three parts: input part, indicator part, and output part. Its structure is as follows: (1) The input part is all the data of the integrated energy system at the same time. (2) The indicator part is the category possibility of all indicators involved in the fuzzy business rules in step 1; (3) The output part is the degree of violation of each fuzzy rule. The neural network design from the indicator part to the output part is designed according to the first-order logic representation of each fuzzy rule. The number of neurons in the output part is the number of business rules; the details are as follows: The intersection operation is designed to connect all neurons involved in the intersection operation to a new neuron with a connection coefficient of 1 and a bias of -n, where n is the number of neurons involved in the intersection operation. (4) The loss function of the trainer is the sum of the squares of the prediction results, and training is performed using the gradient descent method. Since the above loss function and neuron calculations are composed of sigmoid functions and basic operation symbols, they can be derived, and the gradient descent method is feasible; The training data is the historical non-abnormal data of the integrated energy system. The output result mark value of the data at each moment is set to 0, which means that the violation of the fuzzy rule by the non-abnormal data is 0; Step 3: Use the wavelet neural network with momentum online to predict the data to be repaired using non-abnormal data that has a high correlation with the data to be repaired; The predicted result replaces the abnormal data and is input into the offline-trained indicator rule satisfaction classifier. If the loss function value is lower than the set threshold, the predicted result is used as the repair value. If the loss function value is higher than the set threshold, the abnormal data is determined to be caused by equipment abnormality. In this case, no data cleaning is performed, and an "equipment abnormality" warning signal is issued, allowing maintenance personnel to perform equipment maintenance work.