Power system multi-source heterogeneous data processing method and related device

Through technologies such as data dictionary mapping, incremental data synchronization, Kafka message queue and zero-trust dynamic access control, the problems of multi-source heterogeneous data access, shared exchange and security management in the new power system are solved, and efficient and secure data transmission and dispatching instructions are achieved.

CN120162376APending Publication Date: 2025-06-17CHINA ELECTRIC POWER RESEARCH INSTITUTE CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510242977.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

There are huge challenges in the access, sharing and security management of multi-source heterogeneous data in new power systems, including diverse data formats, inconsistent communication protocols, and data security risks, resulting in low data access efficiency, complex and inefficient data sharing, and risks of security leakage and tampering.

Method used

The data dictionary mapping method is used to convert multi-source heterogeneous data into a unified format, and the data change part is obtained using incremental data synchronization technology and encrypted to generate data digests. The Kafka message queue technology is used to realize high concurrent data transmission and real-time sharing and exchange, and the zero-trust dynamic access control method is used to ensure data security.

Benefits of technology

It realizes unified access and management of multi-source heterogeneous data, improves data transmission efficiency, ensures the confidentiality and integrity of data during transmission, solves the problems of inconsistent data transmission delay, redundancy and synchronization, and ensures the rapid and secure issuance of scheduling instructions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120162376A_ABST
    Figure CN120162376A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of electric power data management, and discloses an electric power system multi-source heterogeneous data processing method and related devices.The method comprises the steps that collected electric power system data are obtained and converted into a unified format through a data dictionary mapping method; the method comprises the following steps: acquiring a change part of power system data by adopting an incremental data synchronization technology, encrypting the change part and generating a data abstract to obtain to-be-transmitted data; sharing the data to be transmitted to a transmission target node by using a Kafka message queue technology; and receiving and checking the issued scheduling instruction, and issuing the checked scheduling instruction to an execution node of the scheduling instruction according to an instruction priority rule and a first-in first-out rule. Unified access and management of multi-source heterogeneous power system data are achieved, high-concurrency data transmission and real-time sharing exchange are achieved, data transmission efficiency is improved, confidentiality and integrity of data in the transmission process are guaranteed, safety and legality of dispatching instructions are guaranteed, and quick response and execution of high-priority instructions are guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of power data management, and relates to a method for processing multi-source heterogeneous data in a power system and related devices. Background Art

[0002] With the development of the new power system, especially the large-scale access of new energy and the wide application of distributed energy systems, the data environment of the power system has become increasingly complex. The new power system covers various energy forms such as wind power, photovoltaic, energy storage, and flexible loads. The data generated by these energy systems are diverse in type, different in format, and inconsistent in communication protocols, posing huge challenges to data access, management, and application.

[0003] In the new power system, data sources are extensive, including sensor data, device status data, environmental data, market transaction data, etc. These data are not only diverse in format, such as structured data, semi-structured data, and unstructured data, but also different in communication protocols. This multi-source heterogeneity makes data access extremely difficult. Traditional data access methods often cannot adapt to this complexity, resulting in low data access efficiency or even inability to access some key data. At the same time, due to the existence of multi-source heterogeneous data, data sharing and exchange become complex and inefficient. The differences in data formats and protocols between different energy systems require a large amount of conversion and adaptation work during data transmission, which not only increases the data transmission delay but also reduces the data transmission efficiency. At the same time, with the continuous increase in the amount of data, the pressure of data sharing and exchange is also increasing day by day, and traditional data sharing and exchange methods are difficult to meet the needs of the new power system.

[0004] In addition, the wide access and sharing of multi-source heterogeneous data also bring serious data security management risks. On the one hand, data may be intercepted, tampered with, or lost during transmission, threatening the confidentiality, integrity, and availability of data; on the other hand, the data security policies and management mechanisms of different energy systems are different, making it difficult to achieve unified data security management. This data security management risk may not only affect the normal operation of the power system but also have a serious impact on national security and social stability. Therefore, there is an urgent need for a method for processing multi-source heterogeneous data in a power system to solve the problems of multi-source heterogeneous data access, sharing, and security management in the new power system. Summary of the Invention

[0005] The purpose of the present invention is to overcome the above-mentioned shortcomings of the prior art and provide a method for processing multi-source heterogeneous data in a power system and related devices.

[0006] To achieve the above object, the present invention adopts the following technical solutions:

[0007] In the first aspect of the present invention, a method for processing multi-source heterogeneous data in a power system is provided, including: acquiring the collected power system data and converting it into a unified format through a data dictionary mapping method; using an incremental data synchronization technology to acquire the changed part of the power system data, encrypting the changed part and generating a data digest to obtain the data to be transmitted; using the Kafka message queue technology to share the data to be transmitted to the transmission target node; receiving the dispatched scheduling instruction and performing verification, and dispatching the verified scheduling instruction to the execution node of the scheduling instruction according to the instruction priority rule and the first-in, first-out rule.

[0008] Optionally, before using the incremental data synchronization technology to acquire the changed part of the power system data, it further includes: preprocessing the power system data and storing it in a distributed storage manner; according to the edge computing requirements, processing the preprocessed power system data stored in the distributed storage based on a real-time data stream computing framework to obtain the real-time processed power system data.

[0009] Optionally, the encrypting the changed part and generating a data digest to obtain the data to be transmitted includes: using the AES-256 symmetric encryption method and the RSA asymmetric encryption method to encrypt the changed part to obtain encrypted data, using the SHA-256 hash algorithm to generate a data digest for the changed part, and combining the encrypted data and the data digest to obtain the data to be transmitted.

[0010] Optionally, the using the Kafka message queue technology to share the data to be transmitted to the transmission target node includes: dividing the data to be transmitted into blocks and applying Reed-Solomon error correction coding to each data block to be transmitted to obtain the coded data of each data block to be transmitted; sharing the coded data of each data block to be transmitted to the transmission target node using the Kafka message queue technology, and the transmission target node uses the Reed-Solomon decoding algorithm to decode and combine the coded data of each data block to be transmitted to obtain the data to be transmitted.

[0011] Optionally, the receiving the dispatched scheduling instruction and performing verification includes: receiving the dispatched scheduling instruction and performing integrity verification and timestamp verification to obtain the verification result of the scheduling instruction.

[0012] Optionally, before sharing the data to be transmitted to the transmission target node using the Kafka message queue technology, the following steps are further included: verifying the transmission target node using the zero-trust dynamic access control method, and after the transmission target node passes the verification, sharing the data to be transmitted to the transmission target node using the Kafka message queue technology; before receiving the dispatched scheduling instruction, the following steps are further included: verifying the scheduling instruction issuing node using the zero-trust dynamic access control method, and after the scheduling instruction issuing node passes the verification, receiving the dispatched scheduling instruction; before dispatching the verified scheduling instruction to the execution node of the scheduling instruction according to the instruction priority rule and the first-in-first-out rule, the following steps are further included: verifying the execution node of the scheduling instruction using the zero-trust dynamic access control method, and after the execution node passes the verification, dispatching the verified scheduling instruction to the execution node of the scheduling instruction according to the instruction priority rule and the first-in-first-out rule.

[0013] Optionally, the following steps are further included: obtaining data change information and converting the data change information into log entries to be added to the log, and generating a log digest according to the log entries; sending the log entries and the log digest to each data synchronization node, and each data synchronization node generates a log digest according to the received log entries and compares it with the received log digest, and after the comparison is consistent, copies the received log entries to its own log.

[0014] In the second aspect of the present invention, a power system multi-source heterogeneous data processing system is provided, including: a data access module, configured to obtain the collected power system data and convert it into a unified format through a data dictionary mapping method; a data processing module, configured to use the incremental data synchronization technology to obtain the changed part of the power system data, encrypt the changed part and generate a data digest to obtain the data to be transmitted; a data sharing module, configured to share the data to be transmitted to the transmission target node using the Kafka message queue technology; an instruction management module, configured to receive the dispatched scheduling instruction and verify it, and dispatch the verified scheduling instruction to the execution node of the scheduling instruction according to the instruction priority rule and the first-in-first-out rule.

[0015] In the third aspect of the present invention, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above-mentioned power system multi-source heterogeneous data processing method are implemented.

[0016] In the fourth aspect of the present invention, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned power system multi-source heterogeneous data processing method are implemented.

[0017] Compared with the prior art, the present invention has the following beneficial effects:

[0018] The multi-source heterogeneous data processing method for the power system of the present invention converts the collected multi-source heterogeneous power system data into a unified format through a data dictionary mapping method, realizes the unified access and management of multi-source heterogeneous power system data, and provides basic support for subsequent data processing and power dispatching optimization. By adopting the incremental data synchronization technology, the changed part of the power system data is obtained, and the changed part is encrypted and a data digest is generated to obtain the data to be transmitted. Then, by using the Kafka message queue technology, the data to be transmitted is shared to the transmission target node, realizing high-concurrency data transmission and real-time sharing and exchange, improving the data transmission efficiency, and ensuring the confidentiality and integrity of the data during the transmission process, effectively solving the problems that the wide-area distributed energy system has data transmission delay, redundancy and inconsistent synchronization, and thus it is difficult to meet the real-time dispatching requirements, and the problems that there are risks of security leakage and tampering during the data transmission process, which affect the data integrity and system stability. By receiving the dispatched scheduling instructions and verifying them, and sending the verified scheduling instructions to the execution nodes of the scheduling instructions according to the instruction priority rules and the first-in-first-out rules, the security and legality of the scheduling instructions are ensured, the rapid response and execution of high-priority instructions are guaranteed, the rapid and safe dispatch of the scheduling instructions is realized, the real-time response and execution accuracy of the power grid dispatching system are guaranteed, and the problems that it is difficult to ensure the rapid transmission and execution of the scheduling instructions caused by the lack of a caching mechanism, poor execution timeliness and insufficient security in the current scheduling instruction management are effectively solved. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 It is a flowchart of the multi-source heterogeneous data processing method for the power system according to an embodiment of the present invention.

[0020] Figure 2 It is a block diagram of the structure of the multi-source heterogeneous data processing system for the power system according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0021] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0022] It should be noted that the terms "first", "second", etc. in the specification, claims and the above-mentioned drawings of the present invention are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0023] The present invention will be further described in detail below with reference to the accompanying drawings:

[0024] See Figure 1 , in an embodiment of the present invention, a method for processing multi-source heterogeneous data of a power system is provided, specifically a method for processing multi-source heterogeneous data of a new power system suitable for accessing multi-energy systems, which can meet the data processing requirements of real-time, high-efficiency and security of the new power system, and effectively support the stable and efficient operation and energy coordination optimization of the new power system.

[0025] Specifically, the method for processing multi-source heterogeneous data of the power system of the present invention includes the following steps:

[0026] S1: Obtain the collected power system data and convert it into a unified format through a data dictionary mapping method.

[0027] S2: Use incremental data synchronization technology to obtain the changed part of the power system data, encrypt the changed part and generate a data digest to obtain the data to be transmitted.

[0028] S3: Use Kafka message queue technology to share the data to be transmitted to the transmission target node.

[0029] S4: Receive the dispatched scheduling instruction and check it, and issue the scheduling instruction that passes the check to the execution node of the scheduling instruction according to the instruction priority rule and the first-in-first-out rule.

[0030] The multi-source heterogeneous data processing method for the power system of the present invention converts the collected multi-source heterogeneous power system data into a unified format through a data dictionary mapping method, realizes the unified access and management of multi-source heterogeneous power system data, and provides basic support for subsequent data processing and power dispatching optimization. By adopting incremental data synchronization technology, the changed part of the power system data is obtained, and the changed part is encrypted and a data digest is generated to obtain the data to be transmitted. Then, by using the Kafka message queue technology, the data to be transmitted is shared to the transmission target node, realizing high-concurrency data transmission and real-time sharing and exchange, improving data transmission efficiency, and ensuring the confidentiality and integrity of data during the transmission process. It effectively solves the problems that the wide-area distributed energy system has data transmission delay, redundancy and inconsistent synchronization, and thus it is difficult to meet the real-time dispatching requirements, and the problems that there are risks of security leakage and tampering during the data transmission process, which affect the data integrity and system stability. By receiving and verifying the dispatched scheduling instructions, and distributing the verified scheduling instructions to the execution nodes of the scheduling instructions according to the instruction priority rules and the first-in first-out rules, the security and legality of the scheduling instructions are ensured, the quick response and execution of high-priority instructions are guaranteed, the quick and safe distribution of the scheduling instructions is realized, the real-time response and execution accuracy of the power grid dispatching system are guaranteed, and it effectively solves the problems that it is difficult to ensure the quick transmission and execution of the scheduling instructions caused by the lack of a caching mechanism, poor execution timeliness and insufficient security in the current scheduling instruction management.

[0031] Exemplarily, when acquiring the collected power system data, the collected power system data is acquired through a multi-protocol adapter. Among them, the multi-protocol adapter adapts to multiple communication protocols and is used to realize the standardized access of data, such as IEC 61850, Modbus, and MQTT, etc. Among them, IEC 61850 is an international standard for power system communication and is mainly applied to substation automation systems. Modbus is an industrial automation communication protocol and is widely used for data transmission of power equipment. MQTT is a lightweight Internet of Things communication protocol and is suitable for data collection of distributed devices such as wind power and photovoltaic. The multi-protocol adapter detects the communication protocol to which the input data belongs, and then extracts data fields according to the protocol rules, completes the conversion from the protocol to the standardized format, and realizes the unified access of data. Through the multi-protocol adapter, cross-protocol data sharing and exchange are realized, different device communication specifications are adapted, and the compatibility and universality of the system are improved.

[0032] Exemplarily, the conversion to the unified format through the data dictionary mapping method can be expressed as:

[0033] X standard =Map(X raw ,D dict )

[0034] Wherein, X rawIs the original dataset, representing the data directly collected from multi-source systems such as wind power, photovoltaic, energy storage, and load; it has the characteristics of diverse data formats (such as floating-point numbers and strings), different units (such as kW and MW, etc.), and various communication protocols. D dict Is the data dictionary, containing the rules and mapping relationships for data standardization. Generally includes data field name mapping: the correspondence between the original field name (such as "PV_Power") and the standard field name (such as "PhotovoltaicPower"). Data unit conversion rules: uniformly convert the unit of the original data (such as "kW") to the standard unit (such as "MW"). Data format standardization rules: unify the numerical types of the original data (such as floating-point numbers or integers, etc.). The data dictionary, as a unified reference standard, formats the data from different sources, mainly including mapping rules for specific protocols and data sources. X standard Is the standardized dataset, representing the data that conforms to the unified standard after mapping and conversion, with the characteristics of unified data fields, consistent data units, and unified data formats. Map() is the mapping function, representing the conversion process of data from the original format to the standardized unified format.

[0035] Explanatory. The incremental data synchronization technology is used to obtain the changed part of the power system data to achieve data transmission, which can effectively reduce data redundancy and improve transmission efficiency. The incremental data synchronization technology is an efficient data synchronization method that only transmits the data that has changed since the last synchronization. This technology reduces the amount of data transmission and improves the synchronization efficiency by identifying the changed part in the dataset. The implementation methods include using timestamp method, log analysis method, or trigger method, etc., to capture data changes. The incremental data synchronization technology can effectively reduce the network burden and improve the real-time performance and accuracy of data synchronization. In the face of the current situation of high latency and synchronization difficulties in the data interaction between the widely distributed energy system and the power grid dispatching system, by using the incremental data synchronization technology to obtain the changed part of the power system data and then only transmit the changed part, the transmission of redundant data can be reduced, the data transmission latency can be lowered, the incremental data can be efficiently synchronized, and the waste of transmission resources can be reduced.

[0036] In a possible implementation manner, before obtaining the changed part of the power system data by using the incremental data synchronization technology, it further includes: preprocessing the power system data and storing it in a distributed storage manner; according to the edge computing requirements, processing the preprocessed power system data stored in the distributed storage based on the real-time data stream computing framework to obtain the real-time processed power system data.

[0037] Explanatory, preprocessing power system data is a crucial step. Since power system data often has the characteristics of being massive, complex, and multi-source, it may contain errors, duplicates, or invalid data. Therefore, before performing incremental data synchronization, it is necessary to clean this data, remove noise and outliers, and ensure the accuracy and reliability of the data. At the same time, to reduce the overhead of data transmission and storage, the data can also be compressed to improve data compactness and transmission efficiency.

[0038] At the same time, adopting a distributed storage method to store the preprocessed power system data is the key to achieving highly available storage and fast query of data. Distributed storage disperses the data across multiple nodes, improving data fault tolerance and availability. Even if a certain node fails, other nodes can still provide data services to ensure data continuity and reliability. In addition, distributed storage also supports parallel processing and fast query of data, meeting the requirements of power systems for real-time and efficiency.

[0039] Explanatory, edge computing requirements can be part of the overall task decomposition. Specifically, first, according to the task decomposition principle of edge computing, the overall data processing task is refined into multiple small tasks. Then, these small tasks are assigned to different edge nodes for execution, and each edge node is responsible for processing its assigned data subset. On the edge node, a real-time data stream computing framework is deployed and run. It pulls the preprocessed power system data from the distributed storage system in real time and performs streaming processing on the data according to the predefined processing logic. During the processing, the real-time data stream computing framework can respond to data changes in real time, perform necessary operations such as calculation, filtering, and aggregation, and finally generate the results of real-time processed power system data.

[0040] First, through task decomposition and edge node processing, the computing pressure on the central system is significantly reduced. Second, the application of the real-time data stream computing framework makes data processing more efficient and flexible. Third, the application of the distributed storage system provides high reliability and scalability for data processing. The real-time data stream computing framework based on edge computing processes the preprocessed power system data stored in a distributed manner, not only reducing the computing pressure on the central system, improving the efficiency and flexibility of data processing, but also ensuring the reliability and scalability of data processing. This processing mode makes full use of the computing power of edge nodes, realizes the local processing of data, reduces data transmission latency and bandwidth consumption, and provides strong support for the real-time monitoring, fault prediction, and intelligent decision-making of power systems. With the continuous development of edge computing and real-time data stream computing technologies, this processing mode will play an increasingly important role in power system data management and applications.

[0041] In a possible implementation, the encryption of the changed part and the generation of a data digest to obtain the data to be transmitted include: encrypting the changed part using the AES-256 symmetric encryption method and the RSA asymmetric encryption method to obtain encrypted data, generating a data digest for the changed part using the SHA-256 hashing algorithm, and combining the encrypted data and the data digest to obtain the data to be transmitted.

[0042] Exemplarily, the AES-256 symmetric encryption method is a symmetric encryption algorithm and belongs to the Advanced Encryption Standard (AES) series. AES is a widely used block cipher algorithm that supports three different key lengths: 128 bits, 192 bits, and 256 bits. AES-256 uses a 256-bit key length, providing extremely high security. The RSA asymmetric encryption method is an asymmetric encryption algorithm. The RSA algorithm is based on the mathematical problem of large number factorization, that is, it is computationally infeasible to factorize the product of two large prime numbers. The SHA-256 hashing algorithm is a cryptographic hash function and belongs to the SHA-2 series of algorithms. SHA-256 can map input data of any length to a hash value of a fixed length (256 bits).

[0043] Specifically, first, the incremental data synchronization technology is used to obtain the changed part of the power system data, and the data of these changed parts is extracted to prepare for subsequent encryption and digest generation operations. Using the pre-generated 256-bit AES symmetric key, the extracted changed part data is used as input and encrypted through the AES-256 encryption algorithm. During the encryption process, the data is divided into appropriate block sizes, and each block of data undergoes the encryption process of the AES-256 algorithm to generate encrypted data blocks. All the encrypted data blocks are combined together to form the complete encrypted data. The SHA-256 hashing algorithm is used to generate a digest for the original changed part data. The changed part data is used as input and undergoes a hashing operation through the SHA-256 algorithm to generate a data digest (hash value) of a fixed length. This data digest is used for subsequent data integrity verification. The encrypted data and the generated data digest are combined together to form the data packet to be transmitted. Optionally, the data packet may also include some additional metadata, such as the algorithm information used for encryption, time stamps, etc., so that the receiving party can correctly parse and verify the data.

[0044] By encrypting the changed parts, the confidentiality of the data is ensured, unauthorized access and theft are prevented, and the security of sensitive information is effectively protected. The use of RSA asymmetric encryption technology further enhances the encryption strength and reliability, making it much more difficult to crack the encrypted data. At the same time, the SHA-256 hash algorithm is used to generate a data digest for the changed parts, providing a means for data integrity verification. Any minor tampering with the data can be quickly detected. Combining the encrypted data and the data digest to obtain the data to be transmitted ensures both the confidentiality and integrity of the data, providing comprehensive security protection for data transmission and processing, and greatly enhancing the security and credibility of data communication.

[0045] Explanatory, based on Kafka message queue technology, it can achieve high-concurrency transmission and distribution of multi-source heterogeneous data in a multi-energy power system. Kafka message queue technology is a distributed, high-throughput, and highly scalable message queue system. It is based on the publish / subscribe model, allowing producers to publish messages in the Kafka cluster, and consumers to subscribe to and consume messages from the cluster. Kafka achieves high availability and load balancing of data through partition and replication mechanisms, and is suitable for scenarios such as big data processing, stream computing, and log collection. It can process a large amount of data in real time and meet the requirements of high-concurrency access and persistent storage.

[0046] In a possible implementation manner, the sharing of the data to be transmitted to the transmission target node by using Kafka message queue technology includes: dividing the data to be transmitted into blocks and applying Reed-Solomon error correction coding to each block of the data to be transmitted to obtain the coded data of each block of the data to be transmitted; sharing the coded data of each block of the data to be transmitted to the transmission target node by using Kafka message queue technology, and the transmission target node uses the Reed-Solomon decoding algorithm to decode and combine the coded data of each block of the data to be transmitted to obtain the data to be transmitted.

[0047] Explanatory, data redundancy protection and fault tolerance recovery are achieved using Reed-Solomon error correction coding to enhance system reliability. Reed-Solomon error correction code is based on algebraic geometry and finite field theory. It represents data as polynomials over a finite field and generates parity-check codes to achieve error correction. Specifically, in the process of sharing the data to be transmitted to the transmission target node using Kafka message queue technology, the data to be transmitted is first block-processed, and then the Reed-Solomon error correction coding algorithm is applied to each data block to generate coded data containing redundant information. Then, these coded data are sent to the transmission target node through Kafka message queue technology. After receiving the coded data, the transmission target node decodes each data block using the Reed-Solomon decoding algorithm, and finally combines all the decoded data blocks to restore the original data to be transmitted.

[0048] Based on Reed-Solomon error correction coding, the reliability of data transmission can be effectively improved. Reed-Solomon error correction coding can resist certain data loss or damage during data transmission. Even if some of the coded data is lost or in error during transmission, the transmission target node can still restore the complete data to be transmitted through the decoding algorithm, thus ensuring the accuracy and integrity of data transmission.

[0049] In a possible implementation manner, the receiving and verifying the dispatched instruction includes: receiving the dispatched instruction and performing integrity verification and timestamp verification to obtain the verification result of the dispatched instruction.

[0050] Specifically, after receiving the dispatched instruction, the system first performs integrity verification on the instruction to ensure that the instruction has not been tampered with or damaged during transmission. Then, the system performs timestamp verification to verify whether the timestamp of the instruction is legal to ensure the timeliness and sequentiality of the instruction. After completing these two verifications, the system obtains the verification result of the dispatched instruction for subsequent processing decisions. By performing integrity verification and timestamp verification on the dispatched instruction, the accuracy and reliability of the instruction can be effectively ensured. Integrity verification can prevent the instruction from being maliciously tampered with or damaged during transmission, ensuring the authenticity and integrity of the instruction. Timestamp verification can ensure the timeliness and sequentiality of the instruction, preventing outdated or out-of-order instructions from interfering with or damaging the system. In this way, the system can execute the dispatched instruction more safely and reliably, improving the stability and availability of the system. In addition, an instruction monitoring method can be designed to automatically identify and intercept abnormal or illegal instructions.

[0051] Exemplarily, when the passed inspection scheduling instructions are sent to the execution nodes of the scheduling instructions according to the instruction priority rule and the first-in-first-out rule, the AES-256 encryption method can be used to achieve end-to-end encrypted transmission, and at the same time, the SHA-256 hash algorithm is used to ensure that the transmitted scheduling instructions are not tampered with. In addition, the execution results of the scheduling instructions can be recorded, the execution status can be fed back, and at the same time, log records can be generated to form a closed-loop management, which generally includes: monitoring the execution status of the scheduling instructions and the node response situation, if the scheduling instruction execution fails, automatically triggering a remedial mechanism or redistributing the instruction, and recording the execution log to provide historical scheduling instruction tracking and query.

[0052] Explanatorily, when the passed inspection scheduling instructions are sent to the execution nodes of the scheduling instructions according to the instruction priority rule and the first-in-first-out rule, that is, the received scheduling instructions are classified according to the priority (such as high priority, medium priority, and low priority), and then the scheduling instruction management is carried out in combination with the first-in-first-out rule. Under different priorities, it is ensured that the scheduling instructions with high priority are processed first, and under the same priority, it is ensured that the scheduling instructions received first are processed first. By sending the scheduling instructions in the order of the priority queue, the high-level scheduling instructions are preferentially executed to meet the real-time requirements of power grid scheduling. Optionally, monitor the execution time of the high-priority scheduling instructions, and automatically trigger a compensation mechanism to redistribute the scheduling instructions after timeout. Optionally, after receiving the sent scheduling instructions, delete the duplicate scheduling instructions and update the cache data in real time to maintain cache consistency.

[0053] In a possible implementation manner, before sharing the data to be transmitted to the transmission target node by using the Kafka message queue technology, it further includes: using the zero-trust dynamic access control method to verify the transmission target node, and after the transmission target node passes the verification, sharing the data to be transmitted to the transmission target node by using the Kafka message queue technology; before receiving the sent scheduling instructions, it further includes: using the zero-trust dynamic access control method to verify the scheduling instruction sending node, and after the scheduling instruction sending node passes the verification, receiving the sent scheduling instructions; before sending the passed inspection scheduling instructions to the execution nodes of the scheduling instructions according to the instruction priority rule and the first-in-first-out rule, it further includes: using the zero-trust dynamic access control method to verify the execution nodes of the scheduling instructions, and after the execution nodes pass the verification, sending the passed inspection scheduling instructions to the execution nodes of the scheduling instructions according to the instruction priority rule and the first-in-first-out rule.

[0054] Explanatory, the zero-trust dynamic access control method is a security access control strategy based on the zero-trust concept. It breaks the concept of building a trust domain based on network boundaries in the traditional security architecture and believes that both the internal and external networks are insecure. It is necessary to continuously and dynamically evaluate and authorize all access requests. The zero-trust dynamic access control method collects context information such as the user's identity, device status, network environment, and application security policies, and evaluates the user's access requests in real time. The system will automatically adjust the access permissions according to these dynamic information to ensure that only legitimate and secure users can access resources and applications.

[0055] Specifically, before sharing the data to be transmitted using the Kafka message queue technology, the transmission target node is verified to ensure that the data is only sent to trusted nodes; before receiving the scheduling instruction, the instruction issuing node is verified to prevent the intrusion of malicious instructions; before issuing the scheduling instruction to the execution node, the execution node is verified again to ensure that the instruction is correctly executed. This multi-level verification mechanism effectively prevents unauthorized access and potential security threats, and enhances the overall security and reliability of the system.

[0056] In a possible implementation manner, the multi-source heterogeneous data processing method for the power system further includes: obtaining data change information and converting the data change information into log entries and adding them to the log, and generating a log summary according to the log entries; sending the log entries and the log summary to each data synchronization node, and each data synchronization node generates a log summary according to the received log entries and compares it with the received log summary, and copies the received log entries to its own log after the comparison is consistent.

[0057] Explanatory, data synchronization and consistency are maintained through the Raft consensus algorithm between different nodes. The Raft consensus algorithm is a consensus algorithm used to manage replicated logs in a distributed system. It elects a leader to be responsible for managing and coordinating log replication to ensure data consistency among all nodes. The algorithm decomposes the consistency problem into sub-problems such as leader election, log replication, and security. Through clear leader election and log replication mechanisms, it simplifies the handling of consistency problems in a distributed system.

[0058] In a possible implementation manner, the multi-source heterogeneous data processing method for the power system further includes: achieving the balance between power grid supply and demand and optimizing intelligent scheduling through the spatio-temporal data fusion algorithm and the reinforcement learning model.

[0059] Explanatory, the data characteristics of multi - energy systems are complex. Traditional scheduling methods are difficult to efficiently optimize the supply - demand balance and cannot cope with the uncertainty of new energy. Through the Transformer spatio - temporal data fusion model, spatio - temporal feature extraction and fusion analysis of multi - source energy data are realized, improving data utilization. Reinforcement learning and genetic algorithms are used for power grid supply - demand balance and intelligent scheduling optimization, dynamically adjusting scheduling strategies. Based on the robust optimization model and Monte Carlo simulation, the uncertainty of new energy output is processed. Data interconnection and collaborative scheduling among multi - energy systems are realized, energy resource allocation is optimized, the new energy consumption and supply - demand balance capabilities are improved, effectively coping with power grid scheduling management under extreme working conditions, and improving the safety and stability of power grid operation.

[0060] The method for processing multi - source heterogeneous data in the power system of the present invention is applicable under the background of new power systems, and has the best application effects especially in the following scenarios: 1. The wide - area access grid scheduling scenario of multi - energy entities; 2. Real - time data interaction and security management of high - proportion new - energy power systems; 3. Data sharing and fusion analysis of large - scale distributed energy systems; 4. Cross - regional collaborative scheduling and supply - demand balance optimization of power systems; 5. Coping with power grid scheduling order management and execution under complex extreme working conditions.

[0061] For the wide - area access grid scheduling scenario of multi - energy entities, with the rapid development of distributed energy (such as wind power, photovoltaic, energy storage, flexible loads), wide - area distributed multi - energy systems need to be connected to the power grid scheduling platform, forming the real - time interaction and scheduling management requirements of multi - source heterogeneous data. There are challenges such as diverse data sources, inconsistent formats, identifiers, and communication protocols; multi - energy data need to be shared and synchronized in real - time, with large data volumes and wide distribution; and the security and integrity of data access need to be guaranteed. By standardizing the acquisition of data from multi - energy systems such as wind power, photovoltaic, and energy storage, adapting to multiple communication protocols, and completing data standardization and mapping; using Kafka message queues and incremental data synchronization technology to achieve efficient sharing and real - time synchronization of wide - area data; adopting end - to - end encryption (AES - 256 encryption method) and fault - tolerance protection to ensure the security and reliability of data transmission. Finally, the problem of accessing multi - source heterogeneous data is efficiently solved, ensuring the security and integrity of data transmission; and supporting the real - time monitoring and management of wide - area distributed energy data by the power grid scheduling platform.

[0062] For the real-time data interaction and security management of a new energy power system with a high proportion, in the new energy power system, the proportion of wind power and photovoltaic power generation is gradually increasing. The power grid operation faces problems such as large fluctuations and strong randomness in new energy output, and real-time data interaction and efficient management are required. There are challenges such as large fluctuations in new energy output, high demand for real-time data interaction; rapid response to dispatching instructions to ensure supply-demand balance and system stability; and security risks faced during data sharing and transmission. By deploying edge computing nodes at the new energy access end, data preprocessing, compression, and real-time analysis are realized; through distributed data synchronization and the Raft consensus algorithm, real-time synchronization and reliable transmission of new energy data are ensured; and instruction caching, priority scheduling, and secure issuance are realized to ensure the rapid response and execution of power grid dispatching. Finally, real-time monitoring and management of new energy power generation data are carried out to optimize the power grid supply-demand balance; and the execution efficiency and security of dispatching instructions are improved to ensure the stable operation of the power grid.

[0063] For the data sharing and fusion analysis of a large-scale distributed energy system, in the distributed energy system, data is scattered and comes from various sources, and data sharing and fusion analysis have become the key to realizing energy collaborative dispatching. There are challenges such as wide data distribution and great difficulty in fusion analysis; high real-time requirements and large data volume for distributed energy data; and the need for an efficient data sharing and processing mechanism. By uniformly collecting distributed energy data, data cleaning and standardization are carried out; at the same time, the Transformer model is combined to fuse spatio-temporal data features, and the reinforcement learning algorithm is combined to optimize intelligent dispatching, and efficient data sharing and synchronization are realized through Kafka and incremental data synchronization technology. Finally, efficient data sharing, real-time fusion, and collaborative dispatching of the distributed energy system are realized, and the visualization management and decision-making ability of distributed energy data are improved.

[0064] For the cross-regional collaborative dispatching and supply-demand balance optimization of the power system, in the cross-regional operation of the power grid, it is necessary to coordinate energy resources in each region to achieve cross-regional supply-demand balance and dispatching optimization. There are challenges such as great difficulty in cross-regional energy data interaction and dispatching coordination; high reliability and low latency required for real-time data transmission; and the need for cross-regional data fusion and dispatching optimization. By realizing the synchronization and sharing of large-scale cross-regional data to ensure data consistency, and by deploying computing nodes in a distributed manner to realize local data processing and reduce the cross-regional transmission load; at the same time, the reinforcement learning algorithm is combined to optimize the cross-regional supply-demand balance and realize efficient energy utilization. Finally, the cross-regional energy dispatching efficiency is improved, the power grid supply-demand balance is ensured, and data fusion and cross-regional collaborative dispatching optimization are realized to enhance system stability.

[0065] For the management and execution of power grid dispatching instructions under complex and extreme working conditions, in extreme weather conditions such as typhoons and blizzards, the power grid operation faces the impacts of extreme loads and new energy fluctuations. The dispatching system needs to have the capabilities of rapid response and fault tolerance. There are challenges such as rapid caching and issuing of dispatching instructions to ensure real-time execution; data transmission needs to have security and high fault tolerance to avoid data loss; and it is necessary to optimize the dispatching plan under extreme working conditions to ensure the stability of the power grid. Through priority caching management and timestamp verification, quickly process and safely issue dispatching instructions; by using AES encryption and Reed-Solomon error correction coding, enhance data security and transmission fault tolerance; at the same time, combine with a robust optimization model to optimize the dispatching plan and quickly respond to extreme load changes. Finally, ensure the rapid and safe execution of dispatching instructions, improve the stability and response ability of power grid operation, and reduce the power grid operation risk under extreme working conditions through intelligent optimization of the dispatching plan.

[0066] The method for processing multi-source heterogeneous data in the power system of the present invention achieves the following remarkable effects: 1. Standardized data access: Solve the problem of accessing multi-source heterogeneous data and realize unified management and standardized processing of data. 2. Efficient real-time data sharing: Through high-concurrency message queues and data synchronization technologies, improve the real-time performance and consistency of data transmission. 3. Data security and fault tolerance: Ensure data security and transmission reliability through encrypted transmission and fault tolerance protection technologies. 4. Edge computing to accelerate response: Realize preprocessing and distributed computing of large-volume data on the edge side to improve the real-time response ability of the system. 5. Safe management of dispatching instructions: Realize caching, priority management and safe issuing of dispatching instructions to improve the accuracy and timeliness of dispatching. At the same time, intelligent dispatching and supply-demand balance can also be achieved: Optimize the dispatching strategy through intelligent algorithms to improve the efficiency and stability of power grid dispatching.

[0067] The following is an apparatus embodiment of the present invention, which can be used to execute the method embodiment of the present invention. For the details not disclosed in the apparatus embodiment, please refer to the method embodiment of the present invention.

[0068] See Figure 2 , in another embodiment of the present invention, a power system multi-source heterogeneous data processing system is provided, which can be used to implement the above-mentioned power system multi-source heterogeneous data processing method. Specifically, the power system multi-source heterogeneous data processing system includes a data access module, a data processing module, a data sharing module, and an instruction management module.

[0069] Among them, the data access module is used to obtain the collected power system data and convert it into a unified format through the data dictionary mapping method; the data processing module is used to adopt the incremental data synchronization technology to obtain the changed part of the power system data, encrypt the changed part and generate a data digest to obtain the data to be transmitted; the data sharing module is used to share the data to be transmitted to the transmission target node by using the Kafka message queue technology; the instruction management module is used to receive the issued scheduling instructions for inspection, and issue the scheduling instructions that pass the inspection to the execution nodes of the scheduling instructions according to the instruction priority rule and the first-in-first-out rule.

[0070] In a possible implementation manner, it further includes a consistency module, which is used to obtain data change information, convert the data change information into log entries and add them to the log, and generate a log digest according to the log entries; send the log entries and the log digest to each data synchronization node, and each data synchronization node generates a log digest according to the received log entries and compares it with the received log digest, and after the comparison is consistent, copy the received log entries to its own log.

[0071] All relevant contents of each step involved in the embodiments of the foregoing power system multi-source heterogeneous data processing method can be cited in the function description of the corresponding functional modules of the power system multi-source heterogeneous data processing system in the embodiments of the present invention, and will not be elaborated here.

[0072] The division of modules in the embodiments of the present invention is illustrative, only a logical function division. In actual implementation, there may be other division methods. In addition, in each embodiment of the present invention, the functional modules can be integrated in one processor, or can exist separately physically, or two or more modules can be integrated in one module. The above integrated modules can be implemented in the form of hardware or in the form of software functional modules.

[0073] In another embodiment of the present invention, a computer device is provided. The computer device includes a processor and a memory. The memory is used to store a computer program, and the computer program includes program instructions. The processor is used to execute the program instructions stored in the computer storage medium. The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, and is suitable for implementing one or more instructions. Specifically, it is suitable for loading and executing one or more instructions in the computer storage medium to implement the corresponding method flow or corresponding function. The processor described in the embodiment of the present invention can be used for the operation of the multi-source heterogeneous data processing method in the power system.

[0074] In another embodiment of the present invention, a storage medium is also provided, specifically a computer-readable storage medium (Memory). The computer-readable storage medium is a memory device in the computer device and is used to store programs and data. It can be understood that the computer-readable storage medium here can include both the built-in storage medium in the computer device and, of course, the extended storage medium supported by the computer device. The computer-readable storage medium provides a storage space, and the operating system of the terminal is stored in this storage space. And, one or more instructions suitable for being loaded and executed by the processor are also stored in this storage space. These instructions can be one or more computer programs (including program codes). It should be noted that the computer-readable storage medium here can be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory. The one or more instructions stored in the computer-readable storage medium can be loaded and executed by the processor to implement the corresponding steps of the multi-source heterogeneous data processing method in the power system in the above embodiments.

[0075] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0076] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0077] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0078] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0079] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that it is still possible to modify the specific embodiments of the present invention or make equivalent substitutions. Any modification or equivalent substitution that does not depart from the spirit and scope of the present invention shall be covered by the protection scope of the claims of the present invention.

Claims

1. A method for processing multi-source heterogeneous data in a power system, characterized in that: include: Acquire the collected power system data and convert it into a unified format through data dictionary mapping method; The incremental data synchronization technology is used to obtain the changed part of the power system data, and the changed part is encrypted and a data summary is generated to obtain the data to be transmitted; Use Kafka message queue technology to share the data to be transmitted to the transmission target node; Receive and inspect the dispatching instructions issued, and send the dispatching instructions that pass the inspection to the execution node of the dispatching instructions according to the instruction priority rule and the first-in-first-out rule.

2. The method for processing multi-source heterogeneous data in a power system according to claim 1, characterized in that: Before the incremental data synchronization technology is used to obtain the changed part of the power system data, the method further includes: Preprocess the power system data and store it in a distributed storage manner; According to the edge computing requirements, the pre-processed power system data in distributed storage is processed based on the real-time data stream computing framework to obtain real-time processed power system data.

3. The method for processing multi-source heterogeneous data in a power system according to claim 1, characterized in that: The encrypting of the changed part and generating a data summary to obtain the data to be transmitted includes: The AES-256 symmetric encryption method and the RSA asymmetric encryption method are used to encrypt the changing part to obtain encrypted data, and the SHA-256 hash algorithm is used to generate a data summary for the changing part, and the encrypted data and the data summary are combined to obtain the data to be transmitted.

4. The method for processing multi-source heterogeneous data in a power system according to claim 1, characterized in that: The use of Kafka message queue technology to share the data to be transmitted to the transmission target node includes: The data to be transmitted is divided into blocks and Reed-Solomon error correction coding is applied to each data block to obtain the encoded data of each data block to be transmitted; the encoded data of each data block to be transmitted is shared to the transmission target node using Kafka message queue technology, and the transmission target node uses the Reed-Solomon decoding algorithm to decode and combine the encoded data of each data block to be transmitted to obtain the data to be transmitted.

5. The method for processing multi-source heterogeneous data in a power system according to claim 1, characterized in that: The receiving and checking of the dispatched scheduling instruction includes: receiving the dispatched scheduling instruction and performing integrity check and timestamp check to obtain the check result of the scheduling instruction.

6. The method for processing multi-source heterogeneous data in a power system according to claim 1, characterized in that: Before the Kafka message queue technology is used to share the data to be transmitted to the transmission target node, it also includes: using a zero-trust dynamic access control method to verify the transmission target node, and after the transmission target node passes the verification, using the Kafka message queue technology to share the data to be transmitted to the transmission target node; Before receiving the dispatching instruction, the method further includes: verifying the dispatching instruction issuing node using a zero-trust dynamic access control method, and receiving the dispatching instruction after the dispatching instruction issuing node passes the verification; Before sending the verified scheduling instructions to the execution node of the scheduling instructions according to the instruction priority rule and the first-in-first-out rule, it also includes: using a zero-trust dynamic access control method to verify the execution node of the scheduling instructions, and after the execution node is verified, sending the verified scheduling instructions to the execution node of the scheduling instructions according to the instruction priority rule and the first-in-first-out rule.

7. The method for processing multi-source heterogeneous data in a power system according to claim 1, characterized in that: Also includes: Obtain data change information and convert the data change information into log entries and add them to the log, and generate log summaries based on the log entries; The log entries and log summaries are sent to each data synchronization node. Each data synchronization node generates a log summary based on the received log entries and compares it with the received log summary. After the comparison is consistent, the received log entry is copied to its own log.

8. A multi-source heterogeneous data processing system for a power system, characterized in that: include: Data access module, used to obtain the collected power system data and convert it into a unified format through a data dictionary mapping method; A data processing module is used to obtain the changed part of the power system data by using the incremental data synchronization technology, and to encrypt the changed part and generate a data summary to obtain the data to be transmitted; The data sharing module is used to share the data to be transmitted to the transmission target node using Kafka message queue technology; The instruction management module is used to receive and check the dispatching instructions issued, and send the dispatching instructions that pass the inspection to the execution node of the dispatching instructions according to the instruction priority rule and the first-in-first-out rule.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method for processing multi-source heterogeneous data in a power system as described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method for processing multi-source heterogeneous data in a power system as claimed in any one of claims 1 to 7 are implemented.