A centralized archiving and processing system and method for power engineering data based on cloud storage
Through the hash ring construction and load balancing mechanism based on cloud storage and distributed storage, the problem of low storage and query efficiency in power engineering data management is solved, efficient and secure data management and dynamic storage are achieved, and the reliability and flexibility of the system are improved.
Patent Information
- Application Number
- CN202510067970.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2045-01-16
AI Technical Summary
The existing centralized archiving and processing system for power engineering data has deficiencies in data management and storage efficiency, security, and reliability. In particular, it is difficult to achieve efficient, secure, and flexible storage and query in the management of massive data.
By building a hash ring based on cloud storage and distributed storage technology, and combining it with a load balancing mechanism, digital and intelligent management of power engineering data can be achieved, including digital formatting, correlation matrix analysis, distributed hash storage, and dynamic archiving management.
It improves the storage and retrieval efficiency of data, optimizes the data organization structure, enhances the system's reliability and disaster recovery capabilities, and ensures the rational allocation of storage resources and the efficient operation of the system.
Smart Images

Figure CN119903021B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to data processing load balancing, and in particular to a centralized archiving processing system and method for power engineering data based on cloud storage. Background Art
[0002] A cloud-based centralized archiving and processing system and method for power engineering data uploads digitized power engineering data to a cloud storage platform. Using correlation analysis techniques, the system obtains the union of random data packets and generates a correlation matrix based on the proportion of the union to the entire data set. Subsequently, a distributed storage hash space is constructed based on a hash ring, and a distributed hash algorithm is used to disperse the data based on correlation. By introducing a load balancing mechanism, dynamic and intelligent archiving and management of power engineering data is achieved. This method not only improves the flexibility and reliability of data storage but also effectively optimizes storage resource allocation and access performance, providing technical support for the efficient management of power engineering data.
[0003] The centralized archiving and processing systems and methods for power engineering data currently available on the market primarily rely on cloud computing, distributed storage, and big data technologies to achieve efficient management of massive amounts of engineering data. Such systems typically centrally aggregate and categorize the design drawings, construction records, equipment information, and maintenance data generated in power engineering projects. The system supports rapid entry, encrypted storage, and intelligent retrieval of files in multiple formats, ensuring data security and traceability. At the same time, the integrated archiving method also supports permission management and collaboration functions, providing hierarchical access rights based on different user needs to improve work efficiency. Furthermore, by introducing artificial intelligence technology, the system can perform in-depth analysis of historical data, providing support for engineering optimization and decision-making. Compared with traditional decentralized management methods, this method significantly improves data integrity, query efficiency, and archiving quality, and is an important component of the digital transformation of power engineering. Summary of the Invention
[0004] In order to improve the existing centralized archiving and processing system and method for electric power engineering data, a centralized archiving and processing system and method for electric power engineering data based on cloud storage is provided. The method realizes efficient management of data through cloud computing and distributed storage technology, integrates data uploading, correlation analysis, intelligent storage and dynamic optimization functions, and improves archiving efficiency and security.
[0005] In order to achieve the above objects, the technical solution adopted by the present invention is:
[0006] A centralized archiving and processing method for power engineering data based on cloud storage, comprising:
[0007] Digitally format all power engineering data through scanners;
[0008] Upload digitized power engineering data to the cloud storage platform based on distributed storage technology;
[0009] Based on the acquired power engineering data, analyze the correlation between the data and obtain the correlation matrix;
[0010] Based on the data management requirements of power engineering materials, a hash space with distributed storage based on hash ring is constructed;
[0011] Based on the obtained correlation matrix, the data is dispersed and stored in the hash ring according to the correlation through the distributed hash algorithm;
[0012] Dynamic archiving and management of power engineering data based on load balancing storage.
[0013] Preferably, the digital formatting of all power engineering data by a scanner specifically includes:
[0014] Obtain electronic image data of all power engineering materials through scanners;
[0015] Based on the image data, digital image processing is performed to obtain the content data of the power engineering data;
[0016] Based on the content data of power engineering materials, basic information data of each material is obtained.
[0017] Preferably, the analyzing the correlation between the data based on the acquired power engineering data to obtain the correlation matrix specifically includes:
[0018] Based on the data in the power engineering data of each project, obtain all data packages of the data;
[0019] Based on the data contained in the data packages of all projects, the percentage of the combined data of each data package in the total project data is analyzed to obtain the correlation matrix of the data packages.
[0020] Preferably, the step of analyzing the percentage of the combined data of each data package in the total project data package based on the data package contained in all projects to obtain the data package relevance matrix specifically includes:
[0021] Based on the data contained in the data packets of all the items obtained, obtain the data packet set P={ , }, each data packet Contains a subset of data , the global data set is ;
[0022] The entire data set Each data item in Mapped to index k, and each packet is represented by a binary vector , the formula is:
[0023]
[0024] Based on two random data packets from all projects and , calculate the union of the two data packets , the formula is:
[0025]
[0026] Based on the obtained union, calculate its size using the formula:
[0027]
[0028] Based on the size of the obtained union, calculate the percentage of the obtained union in the entire set. The formula is:
[0029]
[0030] Based on the calculation of the union ratio data between all data packets, the correlation matrix R is constructed. The matrix is:
[0031]
[0032] Preferably, the construction of a hash space for distributed storage based on a hash ring based on the power engineering data management requirements specifically includes:
[0033] Construct a hash ring of length N and determine the hash value range to be [0, -1];
[0034] Add a unique identifier to each node;
[0035] Based on the hash function H(x), the data packet Mapped onto the ring.
[0036] Preferably, the step of distributing and storing the data packets in a hash ring according to the association by using a distributed hash algorithm based on the obtained association matrix specifically includes:
[0037] Based on the correlation matrix, data packets with high correlation are placed at both ends of a diameter of the hash ring, and data packets with low correlation are randomly and evenly distributed on the hash ring.
[0038] Preferably, placing data packets with high correlation at both ends of a diameter of a hash ring based on the correlation matrix and randomly and evenly distributing data packets with low correlation on the hash ring specifically includes:
[0039] like and is a group of highly correlated data packets placed at both ends of the diameter, and the formula is:
[0040]
[0041] Where N is the total length of the hash ring, the opposite side;
[0042] The data packets with low correlation are randomly placed in the free positions on the ring. The formula is:
[0043]
[0044] If the location If it is occupied, create a virtual node and map the data packet to the virtual node. The formula is:
[0045]
[0046] Where k is the virtual machine node number, is the offset of the virtual node.
[0047] Preferably, the dynamic archiving management of power engineering data based on load balancing storage specifically includes:
[0048] Calculate the number of packets at each location on the hash ring and determine whether the distribution is uniform by the ratio of the maximum load to the minimum load. The formula is:
[0049]
[0050] Based on the polling method, part of the data in the high-load area is migrated to other nodes, and the storage location of the data is determined according to the real-time load situation of each node.
[0051] Furthermore, a centralized archiving and processing system for power engineering data based on cloud storage is proposed, including:
[0052] Scanning module: The scanning module is mainly used to scan the power engineering data to obtain clear image data of the power engineering data, which is convenient for subsequent analysis and classification operations;
[0053] A cloud storage module is mainly used to store power engineering data on a remote server through cloud computing technology, so that the power engineering data can be accessed, managed, shared and backed up anytime and anywhere through the network;
[0054] A correlation module, wherein the correlation module is mainly used to calculate the correlation between the data packets by comparing the data in the data packets;
[0055] Construction module: The construction module is mainly used to construct a hash space of distributed storage based on a hash ring;
[0056] A storage processing module, wherein the storage processing module is mainly used to disperse and store associated data packets on a hash ring;
[0057] A load processing module, which is mainly used to optimize the use of storage resources to avoid performance degradation or service unavailability caused by overloading of a single storage node;
[0058] A cloud storage and a processor, wherein the storage and the processor are connected to each other via a network communication, and the processor executes the above method by executing computer instructions.
[0059] Compared with the prior art, the advantages of the present invention are:
[0060] By uploading the digitized data to the cloud storage platform, efficient centralized management of data is achieved, improving storage and retrieval efficiency. Secondly, using the correlation matrix to analyze the relationship between data helps to discover the inherent correlation of data and optimize the data organization structure. Based on the distributed hash algorithm, highly correlated data is stored in a distributed hash ring, which not only improves storage reliability, but also effectively avoids single points of failure and improves the system's disaster recovery capabilities. In addition, combined with the load balancing mechanism, dynamic archiving and balanced distribution of data are achieved, avoiding storage node overload and further ensuring the efficient operation of the system. This method fully utilizes the technical advantages of cloud computing and distributed storage, providing strong support for the intelligent management and long-term storage of power engineering data. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] Figure 1 This is a schematic diagram of the data archiving processing method proposed by the present invention;
[0062] Figure 2 This is a schematic diagram of the digital formatting process proposed by the present invention;
[0063] Figure 3 This is a schematic diagram of the association acquisition proposed by the present invention;
[0064] Figure 4 This is a schematic diagram of the correlation matrix calculation proposed by the present invention;
[0065] Figure 5 This is a schematic diagram of the hash ring construction proposed by the present invention;
[0066] Figure 6 This is a schematic diagram of data storage proposed by the present invention;
[0067] Figure 7 This is a schematic diagram of the high-low correlation data storage proposed by the present invention;
[0068] Figure 8 This is a schematic diagram of the load balancing storage proposed by the present invention. DETAILED DESCRIPTION
[0069] The following description is intended to disclose the present invention so that those skilled in the art can implement the present invention. The preferred embodiments described below are merely examples, and those skilled in the art may conceive of other obvious variations.
[0070] A centralized archiving and processing system for power engineering data based on cloud storage, comprising:
[0071] Scanning module: The scanning module is mainly used to scan the power engineering data to obtain clear image data of the power engineering data, which is convenient for subsequent analysis and classification operations;
[0072] A cloud storage module is mainly used to store power engineering data on a remote server through cloud computing technology, so that the power engineering data can be accessed, managed, shared and backed up anytime and anywhere through the network;
[0073] A correlation module, wherein the correlation module is mainly used to calculate the correlation between the data packets by comparing the data in the data packets;
[0074] Construction module: The construction module is mainly used to construct a hash space of distributed storage based on a hash ring;
[0075] A storage processing module, wherein the storage processing module is mainly used to disperse and store associated data packets on a hash ring;
[0076] A load processing module, which is mainly used to optimize the use of storage resources to avoid performance degradation or service unavailability caused by overloading of a single storage node;
[0077] A cloud storage and a processor, wherein the storage and the processor are connected to each other via a network communication, and the processor executes a cloud storage-based centralized archiving and processing method for power engineering data by executing computer instructions.
[0078] See Figure 1 As shown, the centralized archiving and processing method for power engineering data based on cloud storage specifically includes:
[0079] Step 1: Use a scanner to digitally format all electrical engineering data;
[0080] Step 2: Upload the digitized power engineering data to the cloud storage platform based on distributed storage technology;
[0081] Step 3: Based on the acquired power engineering data, analyze the correlation between the data and obtain the correlation matrix;
[0082] Step 4: Based on the data management requirements of power engineering information, a hash space with distributed storage based on a hash ring is constructed;
[0083] Step 5: Based on the obtained correlation matrix, the data is dispersed and stored in the hash ring according to the correlation through the distributed hash algorithm;
[0084] Step 6: Dynamically archive and manage power engineering data based on load balancing storage.
[0085] See Figure 2 As shown, digital formatting of all power engineering data through a scanner specifically includes:
[0086] Obtain electronic image data of all power engineering materials through scanners;
[0087] Based on the image data, digital image processing is performed to obtain the content data of the power engineering data;
[0088] Based on the content data of power engineering materials, basic information data of each material is obtained.
[0089] See Figure 3 As shown in the figure, based on the acquired power engineering data, the correlation between the data is analyzed to obtain the correlation matrix, which specifically includes:
[0090] Based on the data in the power engineering data of each project, obtain all data packages of the data;
[0091] Based on the data contained in the data packages of all projects, the percentage of the combined data of each data package in the total project data is analyzed to obtain the correlation matrix of the data packages.
[0092] See Figure 4 As shown, based on the data contained in the data packages of all projects, the percentage of the combined data of each data package in the total project data is analyzed to obtain the data package relevance matrix, which specifically includes:
[0093] Based on the data contained in the data packets of all the items obtained, obtain the data packet set P={ , }, each data packet Contains a subset of data , the global data set is ;
[0094] The entire data set Each data item in Mapped to index k, and each packet is represented by a binary vector , the formula is:
[0095]
[0096] Based on two random data packets from all projects and , calculate the union of the two data packets , the formula is:
[0097]
[0098] Based on the obtained union, calculate its size using the formula:
[0099]
[0100] Based on the size of the obtained union, calculate the percentage of the obtained union in the entire set. The formula is:
[0101]
[0102] Based on the calculation of the union ratio data between all data packets, the correlation matrix R is constructed. The matrix is:
[0103]
[0104] See Figure 5 As shown in the figure, based on the data management requirements of power engineering materials, the construction of a distributed storage hash space based on a hash ring specifically includes:
[0105] Construct a hash ring of length N and determine the hash value range to be [0, -1];
[0106] Add a unique identifier to each node;
[0107] Based on the hash function H(x), the data packet Mapped onto the ring.
[0108] Specifically, the hash ring distributes data in a ring-shaped node space through a consistent hashing algorithm, allowing data to be stored in a distributed manner based on association, avoiding the node load that may be caused by traditional centralized storage. Secondly, the design of the hash ring allows only a small amount of data to be reallocated when adding or deleting storage nodes, without the need for full migration. This dynamic expansion capability is very suitable for meeting the growing storage needs of power engineering data, and has little impact on the system, ensuring service continuity.
[0109] See Figure 6 As shown, based on the obtained correlation matrix, the distributed hash algorithm is used to disperse and store the data packets in the hash ring according to the correlation, specifically including:
[0110] Based on the correlation matrix, data packets with high correlation are placed at both ends of a diameter of the hash ring, and data packets with low correlation are randomly and evenly distributed on the hash ring.
[0111] It is understandable that placing highly correlated data at both ends of the diameter, while reflecting a logical association, will cause the data to physically span a greater distance in the hash ring. By introducing logical partitioning through virtualization technology for hash ring nodes, the access path to highly correlated data can be optimized to the shortest path. At the same time, combined with local caching mechanisms, such as hotspot data caching, the cross-node access cost of highly correlated data can be reduced. At the same time, randomly distributed low-correlation data may be completely irregularly scattered on the ring, resulting in a lack of correlation in the data stored on some nodes, increasing the coordination overhead during overall data processing. By establishing weak association rules for low-correlation data during the distribution process, we can ensure that even random distribution can optimize the storage layout to a certain extent. By globally analyzing low-correlation data, we can introduce a weighted hash distribution algorithm to enhance the intelligence of data layout.
[0112] See Figure 7 As shown, based on the correlation matrix, data packets with high correlation are placed at both ends of a diameter of the hash ring, and data packets with low correlation are randomly and evenly distributed on the hash ring. Specifically, the following steps are included:
[0113] like and is a group of highly correlated data packets placed at both ends of the diameter, and the formula is:
[0114]
[0115] Where N is the total length of the hash ring, the opposite side;
[0116] The data packets with low correlation are randomly placed in the free positions on the ring. The formula is:
[0117]
[0118] If the location If it is occupied, create a virtual node and map the data packet to the virtual node. The formula is:
[0119]
[0120] Where k is the virtual machine node number, is the offset of the virtual node.
[0121] See Figure 8 As shown in the figure, dynamic archiving management of power engineering data based on load balancing storage specifically includes:
[0122] Calculate the number of packets at each location on the hash ring and determine whether the distribution is uniform by the ratio of the maximum load to the minimum load. The formula is:
[0123]
[0124] Based on the polling method, part of the data in the high-load area is migrated to other nodes, and the storage location of the data is determined according to the real-time load situation of each node.
[0125] Specifically, for data with high access frequency, a multi-copy mechanism can be used to retain a copy on the target node while retaining the master copy on the original node, reducing changes in user access paths. A temporary cache is established for hot data to reduce the impact of migration on user requests. The migration process is optimized through hot and cold data classification, with cold data prioritized to reduce the impact of migrated data on node performance. Dynamic hot and cold classification monitors data access frequency in real time and dynamically adjusts the hot and cold data division. During the migration process, if the target node fails, the migration task is reassigned to the next target node or a list of available nodes is pre-maintained to quickly switch to the target node.
[0126] In summary, the advantages of the present invention are: by uploading digital data to the cloud storage platform, efficient centralized management of data is achieved, which improves the speed and convenience of data access; by using union analysis and correlation matrix, the correlation between different data is accurately mined, which helps to optimize the data organization structure and query efficiency; constructing a distributed storage hash space of a hash ring and combining the correlation to disperse the storage data, which not only improves the reliability and scalability of storage, but also effectively reduces the risk of single point failure; dynamic archiving management is achieved through load balancing storage, making storage resource allocation more reasonable and ensuring the efficient operation and stability of the system. This method combines efficient storage with intelligent management, is suitable for the large-scale and complex storage needs of power engineering data, and helps the industry's digital transformation and upgrading.
[0127] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The above embodiments and descriptions merely illustrate the principles of the present invention. Various changes and modifications may be made to the present invention without departing from the spirit and scope of the present invention. Such changes and modifications are intended to fall within the scope of the present invention. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.
Claims
1. A centralized archiving and processing method for power engineering data based on cloud storage, characterized in that: include: Digitally format all power engineering data through scanners; Upload digitized power engineering data to the cloud storage platform based on distributed storage technology; Analyzing the correlation between each data based on the acquired power engineering data to obtain a correlation matrix; analyzing the correlation between each data based on the acquired power engineering data to obtain a correlation matrix specifically includes: Based on the data in the power engineering data of each project, obtain all data packages of the data; Based on the data contained in the data packages of all projects, the percentage of the combined data of each data package in the total project data is analyzed to obtain the correlation matrix of the data packages; The data contained in the data packages of all projects, analyzing the percentage of the combined data of each data package in the total project data, and obtaining the data package relevance matrix specifically includes: Based on the data contained in the data packets of all the items obtained, obtain the data packet set P={ , }, each data packet Contains a subset of data , the global data set is ; The entire data set Each data item in Mapped to index s, and each packet is represented by a binary vector , the formula is: ; Based on two random data packets from all projects and , calculate the union of the two data packets , the formula is: ; Based on the obtained union, calculate its size using the formula: ; Based on the size of the obtained union, calculate the percentage of the obtained union in the entire set. The formula is: ; Based on the calculation of the union ratio data between all data packets, the correlation matrix R is constructed. The matrix is: ; Based on the data management requirements of power engineering materials, a hash space with distributed storage based on hash ring is constructed; Based on the obtained correlation matrix, the data is dispersed and stored in the hash ring according to the correlation through the distributed hash algorithm; The method of distributing and storing the data packets in the hash ring according to the association based on the obtained association matrix by using a distributed hash algorithm specifically includes: Based on the correlation matrix, data packets with high correlation are placed at both ends of a diameter of the hash ring, and data packets with low correlation are randomly and evenly distributed on the hash ring; The method of placing data packets with high correlation at both ends of a diameter of a hash ring based on the correlation matrix and randomly and evenly distributing data packets with low correlation on the hash ring specifically includes: like and is a group of highly correlated data packets placed at both ends of the diameter, and the formula is: ; Where N is the total length of the hash ring, the opposite side; The data packets with low correlation are randomly placed in the free positions on the ring. The formula is: ; If the location If it is occupied, create a virtual node and map the data packet to the virtual node. The formula is: ; Where k is the virtual machine node number, is the offset of the virtual node; Dynamic archiving and management of power engineering data based on load balancing storage.
2. The method for centralized archiving and processing of power engineering data based on cloud storage according to claim 1, characterized in that: The digital formatting of all power engineering data by means of a scanner specifically includes: Obtain electronic image data of all power engineering materials through scanners; Based on the image data, digital image processing is performed to obtain the content data of the power engineering data; Based on the content data of power engineering materials, basic information data of each material is obtained.
3. The method for centralized archiving and processing of power engineering data based on cloud storage according to claim 1, characterized in that: The construction of a hash space for distributed storage based on a hash ring based on the power engineering data management requirements specifically includes: Construct a hash ring of length N and determine the hash value range to be [0, -1]; Add a unique identifier to each node; Based on the hash function H(x), the data packet Mapped onto the ring.
4. The method for centralized archiving and processing of power engineering data based on cloud storage according to claim 1, characterized in that: The dynamic archiving management of power engineering data based on load balancing storage specifically includes: Calculate the number of packets at each location on the hash ring and determine whether the distribution is uniform by the ratio of the maximum load to the minimum load. The load calculation formula is: ; Based on the polling method, part of the data in the high-load area is migrated to other nodes, and the storage location of the data is determined according to the real-time load situation of each node.
5. A cloud storage-based centralized archiving and processing system for electric power engineering data, used to implement the cloud storage-based centralized archiving and processing method for electric power engineering data according to any one of claims 1 to 4, characterized in that: include: Scanning module: The scanning module is mainly used to scan the power engineering data to obtain clear image data of the power engineering data, which is convenient for subsequent analysis and classification operations; A cloud storage module is mainly used to store power engineering data on a remote server through cloud computing technology, so that the power engineering data can be accessed, managed, shared and backed up anytime and anywhere through the network; A correlation module, wherein the correlation module is mainly used to calculate the correlation between the data packets by comparing the data in the data packets; Construction module: The construction module is mainly used to construct a hash space of distributed storage based on a hash ring; A storage processing module, wherein the storage processing module is mainly used to disperse and store associated data packets on a hash ring; A load processing module, which is mainly used to optimize the use of storage resources to avoid performance degradation or service unavailability caused by overloading of a single storage node; A cloud storage and a processor, wherein the storage and the processor are connected to each other via a network communication, and the processor executes the method according to any one of claims 1 to 4 by executing computer instructions.
Citation Information
Patent Citations
Neighbor storage method based on data mapping algorithm
CN106020724A
Emergency data information management system and method
CN117726279A