Container isolation and double-node management fused data integration treatment method and system

By adopting a data integration governance method that integrates container isolation and two-node management in the hydropower station system, and using neural network models to analyze communication status, the problems of data synchronization difficulties and communication instability in traditional methods are solved, and efficient and secure data governance is achieved.

CN120216596APending Publication Date: 2025-06-27HUANENG LANCANG RIVER HYDROPOWER CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510277996.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

Traditional data integration governance methods have problems such as data synchronization difficulties, unstable communication status and data security risks in hydropower station systems, which are difficult to meet the needs of modern hydropower station systems for efficient, accurate and safe data governance.

Method used

A data integration governance method that integrates container isolation and two-node management is adopted. By monitoring the data and communication timing data of hydropower station nodes in real time, a neural network model CNN is used to establish a communication state model, deeply analyze the data to obtain key indicators, and trigger corresponding governance strategies based on these indicators to optimize data transmission and isolation strategies.

Benefits of technology

It improves the accuracy and data consistency of data synchronization among nodes, enhances data security and system anti-interference capabilities, optimizes resource allocation and system operation efficiency, and adapts to dynamic changes in the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216596A_ABST
    Figure CN120216596A_ABST
Patent Text Reader

Abstract

The invention discloses a container isolation and double-node management fused data integration governance method and system, and relates to the technical field of data integration governance. Communication time sequence data between a plurality of groups of main nodes and branch nodes of a hydropower station is monitored in real time, and deep analysis is performed by using a neural network model CNN; indexes such as data drift rate and data coverage difference can be accurately identified, and the data synchronization condition can be evaluated in real time. The accuracy of data synchronization between the nodes is promoted to be improved, and the data consistency and timeliness are ensured. By constructing the communication state model and deeply analyzing the data, the communication state can be accurately evaluated, and key communication indexes such as the uplink and downlink transmission ratio, the burst load index and the like can be identified. Once it is found that the communication state is unstable, the first governance strategy, the second governance strategy and the third governance strategy can be triggered, so that the data transmission path is optimized, the network load is adjusted, resource allocation is effectively optimized, and the overall operation efficiency of the system is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data integration and governance, and specifically to a data integration and governance method and system that integrates container isolation and dual-node management. Background Art

[0002] With the continuous development of information technology, the importance of data in all walks of life has become increasingly prominent. Especially in infrastructure fields such as hydropower stations, the real-time monitoring and analysis of data are of great significance for improving system operation efficiency, ensuring equipment safety, and dealing with emergencies. A hydropower station usually consists of multiple main nodes and branch nodes, which are distributed in different geographical locations and require a highly efficient and stable communication system for data transmission and interaction.

[0003] Traditional data integration and governance methods usually have multiple problems, such as difficult data synchronization between nodes, unstable communication status, and data security risks, making it difficult to meet the requirements of modern hydropower station systems for efficient, accurate, and secure data governance. Due to the scattered geographical locations of nodes and differences in network environments, data synchronization often has problems such as inconsistent time sequences and data delays. Summary of the Invention

[0004] In view of the deficiencies of the prior art, the present invention provides a data integration and governance method and system that integrates container isolation and dual-node management to solve the problems mentioned in the background art.

[0005] To achieve the above objectives, the present invention is realized through the following technical solutions: A data integration and governance method that integrates container isolation and dual-node management,

[0006] including the following steps:

[0007] Step 1: Establish a database by real-time monitoring various data in several groups of main nodes and branch nodes of the hydropower station;

[0008] Step 2: Real-time monitor the first communication time sequence data between several groups of main nodes and branch nodes of the hydropower station;

[0009] Use the neural network model CNN to establish and train a communication status model, and conduct in-depth analysis and calculation on the first communication time sequence data to obtain the data drift rate D fl 、data coverage difference index D cd 、up and down transmission ratio U pr 、burst load index F bi and sensitive data exposure rate Sen Ex ; and the data drift rate D fl 、data coverage difference index D cd 、up and down transmission ratio U pr, the sudden load index F bi and the sensitive data exposure rate Sen Ex are correlated to obtain the first communication state index C sc ; when the first communication state index C sc is lower than 0.8, trigger and implement the first governance strategy;

[0010] Step 3. After implementing the first governance strategy, re-collect and obtain several groups of second communication timing data between the master node and the branch nodes, and deeply analyze and calculate through the communication state model to obtain the abnormal behavior response delay Atr l , the node working switching rate S wr , the two-node mutual trust index T xi and the isolation efficiency index G; and calculate the comprehensive governance index Z through correlation;

[0011] Step 4. Preset the efficiency threshold Zy, and compare the comprehensive governance index Z with the preset efficiency threshold Zy, and determine the data governance level status according to the result:

[0012] If the governance status is the first high-efficiency level, maintain the current two-node management architecture;

[0013] If the governance status is the second critical level, optimize the task allocation and load balancing of the branch nodes;

[0014] If the governance status is the third low-efficiency level, adjust the data collaboration strategy between the master node and the branch nodes and optimize the container isolation parameters.

[0015] Preferably, a resource database is established based on the collected data for storing various types of data of each node of the hydropower station;

[0016] The various types of data include but are not limited to: master node data: reservoir water level, flow rate, power generation, equipment operation status, and energy output power; branch node data: equipment health status, real-time grid load, flow rate fluctuation, sensor data, and security monitoring data;

[0017] Data cleaning and standardization: Clean the various types of data collected from the master node and the branch nodes, including: removing abnormal data points, filling in missing data, and standardizing the cleaned data based on dimensionless processing techniques such as Z-Score standardization.

[0018] Preferably, Step 2 includes:

[0019] S21. Collect the first communication timing data of several groups of master nodes and branch nodes, and establish a communication state model through the neural network model CNN. The communication state model contains several convolutional layers for extracting local features of the first communication timing data. The convolutional kernels of each layer scan the input first communication timing data and apply convolutional operations in the local area, including weighted summation and bias terms;

[0020] The several convolutional layers include the feature maps extracted by the convolutional kernels of each layer, which represent the local or global features of the input data at that layer;

[0021] S22. Extract the output feature values of each layer of the master node and the output feature values of each layer of the branch node from the feature maps extracted by the convolutional kernels of each layer, and calculate the feature difference value δ of the master node in the j-th layer of the convolutional neural network through the following formula m,j and the feature difference value δ of the branch node in the j-th layer of the convolutional neural network s,j :

[0022]

[0023] In the formula, is the output feature value of the master node at the i-th moment of the j-th convolutional layer extracted by the communication state model, is the output feature value of the master node at the previous moment of the i-th moment of the j-th convolutional layer extracted by the communication state model. The feature difference value δ of the master node in the j-th layer of the convolutional neural network m,j , indicating the change of the master node in the feature map of this layer;

[0024] In the formula, is the output feature value of the branch node at the i-th moment of the j-th convolutional layer extracted by the communication state model, is the output feature value of the branch node at the previous moment of the i-th moment of the j-th convolutional layer extracted by the communication state model. The feature difference value δ of the branch node in the j-th layer of the convolutional neural network s,j , indicating the change of the branch node in the feature map of this layer;

[0025] S23. Based on the feature difference value δ of the master node in the j-th layer of the convolutional neural network m,j and the feature difference value δ of the branch node in the j-th layer of the convolutional neural network s,j , calculate the data drift rate D fl . The data drift rate D fl represents the change trend of the synchronized data between the master node and the branch node, and is calculated through the following formula:

[0026]

[0027] In the formula, f m,iRepresents the data value of the main node at the $i$-th moment, $f$ s,i Represents the data value of the branch node at the $i$-th moment: $N$ represents the size of the time window, specifically the number of data points in the time series, $\delta$ m,j Represents the feature difference value of the main node in the $j$-th layer of the convolutional neural network; $\delta$ s,j Represents the feature difference value of the branch node in the $j$-th layer of the convolutional neural network, $m$ represents the $m$-th main node, $s$ represents the $s$-th branch node, $L$ represents the total number of convolutional layers, $\alpha$ j Represents the weight coefficient of the $j$-th convolutional layer; ($f$ m,i -$f$ s,i ) represents the out-of-sync difference term of the data between two nodes, Integrates the differences of the convolutional feature maps of each layer in the communication state model network and assigns different weights according to the different layers. The convolutional feature difference $\delta$ m,j -$\delta$ s,j will affect the final drift rate.

[0028] Preferably, step two further includes:

[0029] S24. Extract the first communication timing data and establish a communication state model through the neural network model CNN, and deeply calculate and analyze to obtain the data coverage difference index $D$ cd , the uplink-downlink transmission ratio $U$ pr , the burst load index $F$ bi and the sensitive data exposure rate $Sen$ Ex ;

[0030] S241. The data coverage difference index $D$ cd is calculated and obtained through the following formula:

[0031]

[0032] In the formula, $C$ m,i represents the coverage rate of the main node at the $i$-th moment, $C$ s,i represents the coverage rate of the branch node at the $i$-th moment: $N$ represents the size of the time window; $Z$ m,j represents the coverage difference value of the main node in the $j$-th layer of the convolutional neural network; $Z$ s,j represents the coverage difference value of the branch node in the $j$-th layer of the convolutional neural network, $\alpha$ j represents the weight coefficient of the $j$-th convolutional layer;

[0033] S242. The uplink-downlink transmission ratio $U$ pr is calculated and obtained through the following formula:

[0034]

[0035] In the formula, $D$ up,irepresents the uplink data volume at the i-th moment, specifically the transmission volume from the master node to the branch node, D down,i represents the downlink data volume at the i-th moment, specifically the transmission volume from the branch node to the master node, and N represents the time window size;

[0036] S243. The burst load index F bi is calculated and obtained through the following formula:

[0037]

[0038] In the formula, the burst load index F bi is used to measure the change rate of the branch node load in the case of instantaneous high concurrency, L s,i represents the CPU usage rate of the branch node at the i-th moment, and Δt represents the time interval, represents taking the maximum value of the maximum load change rate in the entire time window;

[0039] S244. The sensitive data exposure rate Sen Ex is calculated and obtained through the following formula:

[0040]

[0041] In the formula, D ex,i represents the exposed sensitive data volume at the i-th moment, which is the transmission volume from the master node to the branch node, P i represents the data type sensitivity weight coefficient, including:

[0042] The first sensitivity data weight coefficient, including:

[0043] Master node data: reservoir water level 1.0, flow rate 0.9, power generation 0.8, equipment operation status 0.95, and energy output power 0.85;

[0044] The second sensitivity data weight coefficient, including:

[0045] Branch node data: equipment health status 0.9, real-time grid load 0.8, flow rate fluctuation 0.7, sensor data 0.75, and security monitoring data 0.9;

[0046] Among them, D dwon,i represents the downlink data volume at the i-th moment, which is the transmission volume from the branch node to the master node, and N represents the time window size.

[0047] Preferably, step two further includes: S25. Extract the data drift rate D fl , data coverage difference index D cd , uplink and downlink transmission ratio U pr, Sudden load index F bi and sensitive data exposure rate Sen Ex , after dimensionless processing, the first communication status index C is calculated through the following associated formula sc :

[0048]

[0049] In the formula, a1, a2, a3, a4 and a5 respectively represent the data drift rate D fl , data coverage difference index D cd , sudden load index F bi , sensitive data exposure rate Sen Ex and uplink-downlink transmission ratio U pr of the weight coefficients, and the sum of the weight coefficients is 1.

[0050] Preferably, step two further includes: S26. Preset the communication threshold to 0.8, and compare the first communication status index C sc with the communication threshold. When the first communication status index C sc < 0.8, it means that the communication status between the current master node and the branch node is unqualified, and the first governance strategy is triggered, including:

[0051] S2611. Use Docker's volumes to allocate independent storage volumes for each master node and branch node, and set access permissions for the storage volumes of each node through container volume isolation;

[0052] S2612. Increase the current synchronization frequency of the current master node and branch node by 30%, increase the data transmission compression rate of the current master node and branch node by 30%, and increase the current 10%-30% upload bandwidth;

[0053] S2613. And encrypt the data type sensitivity weight coefficient Pi of the data transmitted between the master node and the branch node that exceeds 0.9 with the AES encryption algorithm with a 128-bit key length, generate the first encryption instruction and implement it;

[0054] When the first communication status index C sc ≥ 0.8, it means that the communication status between the current master node and the branch node is qualified, and continuous monitoring is carried out.

[0055] Preferably, step three includes:

[0056] S31. After the implementation of the first governance strategy, re-collect and obtain several groups of second communication timing data between the master node and the branch node;

[0057] S32. Input the second communication timing data into the communication status model for in-depth analysis and calculation to obtain the abnormal behavior response delay Atr l , the node working switching rate S wr , the dual-node mutual trust index T xi and the isolation efficiency index G;

[0058] S321. The abnormal behavior response delay Atr l is obtained through the following formula:

[0059]

[0060] In the formula, T resp,i represents the response time of the master node to the abnormal behavior of the branch node at the i-th moment. The abnormal behavior includes data packet replay or forgery behavior; T detect,i represents the time when the abnormal behavior of the branch node is detected at the i-th moment, and R represents the total number of abnormal behavior monitoring times;

[0061] S322. The node working switching rate S wr is obtained through the following formula:

[0062]

[0063] In the formula, ΔR ole,i represents the change in the role switching between the master node and the branch node at the i-th moment. If the master node becomes the branch node role, the value is 1; if the branch node becomes the master node role, the value is -1; if there is no switching, the value is 0; V represents the total number of role switches;

[0064] S323. The dual-node mutual trust index T xi is obtained through the following formula;

[0065]

[0066] In the formula, T success,i represents the number of successfully transmitted tasks between the master node and the branch node at the i-th moment. Whether a specific task is successfully completed or synchronized successfully, E i represents the weight coefficient of task success at the i-th moment, which is related to factors such as the complexity or data volume of the task, and N represents the time window size;

[0067] S324. The isolation efficiency index G is obtained through the following formula:

[0068]

[0069] In the formula, D isolated represents the amount of data successfully isolated, D total represents the total amount of data; Eencrypt Represents the amount of encrypted data, T transmit Represents the data transmission time after encryption; S protected Represents the successfully protected sensitive data, S total Represents the amount of sensitive data; F switch Represents the error rate caused by node switching;

[0070] S33. And the abnormal behavior response delay Atr l The node working switching rate S wr The two-node mutual trust index T xi And the isolation efficiency index G. After dimensionless processing, the comprehensive governance index Z is calculated through the following associated formula:

[0071]

[0072] In the formula, b1, b2, b3, and b4 respectively represent the abnormal behavior response delay Atr l The node working switching rate S wr The two-node mutual trust index T xi And the weight coefficients of the isolation efficiency index G, and the sum of the weight coefficients is 1.

[0073] Preferably, step four includes:

[0074] Preset an efficiency threshold Zy, and preset the efficiency threshold Zy, and compare the comprehensive governance index Z with the preset efficiency threshold Zy to obtain the following judgment results, including:

[0075] When the comprehensive governance index Z > the efficiency threshold Zy, it means that the current two-node operation meets the standard, and the governance status is the first high-efficiency level;

[0076] When the efficiency threshold Zy * 80% ≤ the comprehensive governance index Z ≤ the efficiency threshold Zy, it means that the current two-node operation does not meet the standard, and the governance status is the second critical level;

[0077] When the comprehensive governance index Z < the efficiency threshold Zy * 80%, it means that the current two-node operation does not meet the standard, and the governance status is the third low-efficiency level;

[0078] If the governance status is the first high-efficiency level, maintain the current two-node management architecture;

[0079] If the governance status is the second critical level, optimize the task allocation and load balance of the branch nodes to generate the second governance strategy;

[0080] If the governance status is the third low-efficiency level, adjust the data coordination strategy between the main node and the branch nodes, optimize the container isolation parameters, and generate the third governance strategy.

[0081] Preferably, the second governance strategy includes: based on the real-time monitoring results, reducing the communication task volume between the current master node and the busiest branch node by 15%-20% and transferring it to the low-load node; increasing the synchronization frequency by 30% compared with the first governance strategy, and here maintaining the synchronization frequency increase of 15%-20%; using AES-256 encryption instead of AES-128, and encrypting the data types with a sensitivity weight coefficient Pi exceeding 0.8 in the data transmitted between the master node and the branch node using the AES encryption algorithm with a 256-bit key length, generating a second encryption instruction and implementing it;

[0082] The third governance strategy includes: the second governance strategy includes: based on the real-time monitoring results, reducing the communication task volume between the current master node and the busiest branch node by 21%-30% and transferring it to the low-load node; increasing the synchronization frequency by 40% compared with the first governance strategy and maintaining the synchronization frequency increase of 21%-30%;

[0083] Using Docker volumes configuration, each node is separately allocated a storage volume; reconfiguring the independent volume capacity of the inefficient nodes and expanding the storage volume capacity by 30%-50%;

[0084] Using AES-256 encryption instead of AES-128, and increasing the data block transfer before encryption, and encrypting the data types with a sensitivity weight coefficient Pi exceeding 0.8 in the data transmitted between the master node and the branch node using the AES encryption algorithm with a 256-bit key length, generating a third encryption instruction and implementing it.

[0085] A data integration governance system integrating container isolation and dual-node management includes:

[0086] A database establishment unit for establishing a database by real-time monitoring various types of data in several groups of master nodes and branch nodes of the hydropower station;

[0087] A node communication data acquisition unit for real-time monitoring of the first communication timing data between several groups of master nodes and branch nodes of the hydropower station;

[0088] A communication status model establishment unit for using the neural network CNN to establish a communication status model, deeply analyzing the collected data to obtain the data drift rate D fl 、data coverage difference index D cd 、up and down transmission ratio U pr 、burst load index F bi and sensitive data exposure rate Sen Ex ; and the data drift rate D fl 、data coverage difference index D cd 、up and down transmission ratio U pr 、burst load index Fbi and the sensitive data exposure rate Sen Ex are associated to obtain the first communication status index C sc ; when the first communication status index C sc is lower than 0.8, trigger and implement the first governance strategy;

[0089] A second acquisition and analysis unit, which is used to, after the implementation of the first governance strategy, re-acquire and obtain several groups of second communication timing data between the master node and the branch node, and deeply analyze and calculate through the communication status model to obtain the abnormal behavior response delay Atr l 、the node working switching rate S wr 、the dual-node mutual trust index T xi and the isolation efficiency index G; and calculate the comprehensive governance index Z through correlation;

[0090] A governance level evaluation unit, which is used to preset an efficiency threshold Zy, compare the comprehensive governance index Z with the preset efficiency threshold Zy, determine the data governance level status according to the result, and generate corresponding governance strategies.

[0091] The present invention provides a data integration governance method and system integrating container isolation and dual-node management. It has the following

[0092] beneficial effects:

[0093] (1) Through the container isolation technology, the present invention can effectively isolate the data storage and transmission processes of different nodes, ensuring the data security and reliability of the master node and the branch node. Each node realizes strict protection of sensitive data through an independent storage volume and a customized encryption policy, preventing malicious attacks and data leakage, and enhancing the anti-interference ability and stability of the entire system.

[0094] (2) The present invention can accurately identify indicators such as the data drift rate and data coverage difference by real-time monitoring of the communication timing data between several groups of master nodes and branch nodes in the hydropower station and using the neural network model CNN for in-depth analysis, and can evaluate the data synchronization situation in real time. This can greatly improve the accuracy of data synchronization between nodes, ensuring data consistency and timeliness. By constructing a communication status model and deeply analyzing the data, the present invention can accurately evaluate the communication status and identify key communication indicators such as the up and down transmission ratio and the burst load index. Once it is found that the communication status is unstable (for example, the first communication status index C sc is lower than 0.8), the system can immediately trigger the first governance strategy, thereby optimizing the data transmission path and adjusting the network load to avoid system failures caused by communication problems.

[0095] (3) In infrastructure systems such as hydropower stations, data security is of particular importance. Traditional methods often carry the risk of sensitive data exposure. By monitoring the sensitive data exposure rate and combining indicators such as data drift rate and data coverage difference, the present invention can identify data leakage risks in advance and take targeted measures for governance. By optimizing data transmission and isolation strategies, the present invention significantly improves data security and reduces the possibility of sensitive data exposure. Due to the wide distribution of nodes in hydropower stations, traditional data governance methods often cannot perform load balancing in a timely manner, resulting in some nodes being overloaded while other nodes are idle. By implementing a dual-node management architecture and optimizing the task allocation and load balancing of branch nodes, the system can automatically adjust the working status of nodes based on real-time data analysis, thereby effectively optimizing resource allocation and improving the overall operating efficiency of the system.

[0096] (4) In traditional technologies, data governance often relies on preset static rules and is difficult to cope with the dynamic changes of the system. By introducing a comprehensive governance index Z and comparing it with a preset efficiency threshold Zy, the present invention can automatically adjust the system's strategy according to different data governance level states (such as the first high-efficiency level, the second critical level, and the third low-efficiency level). If the governance state is the low-efficiency level, the system can optimize the data collaboration strategy between the main node and the branch node, adjust the container isolation parameters, and flexibly adapt to the different needs of the system. This adaptive ability enables the hydropower station system to maintain efficient and stable operation in a changing environment. BRIEF DESCRIPTION OF THE DRAWINGS

[0097] Figure 1 It is a schematic diagram of the steps of a data integration governance method that combines container isolation and dual-node management according to the present invention;

[0098] Figure 2 It is a schematic diagram of the trigger conditions and specific measures of the governance strategy according to the present invention;

[0099] Figure 3 It is a schematic diagram of the block diagram process of a data integration governance system that combines container isolation and dual-node management according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0100] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0101] Embodiment 1

[0102] Please refer to Figure 1, the present invention provides a data integration governance method that combines container isolation and dual-node management, including the following steps:

[0103] Step 1: Establish a database by real-time monitoring various types of data in several groups of master nodes and branch nodes of a hydropower station;

[0104] Step 2: Real-time monitor the first communication timing data between several groups of master nodes and branch nodes of a hydropower station;

[0105] Use the neural network model CNN to establish and train a communication status model, and conduct in-depth analysis and calculation on the first communication timing data to obtain the data drift rate D fl 、data coverage difference index D cd 、up and down transmission ratio U pr 、burst load index F bi and sensitive data exposure rate Sen Ex ; and correlate the data drift rate D fl 、data coverage difference index D cd 、up and down transmission ratio U pr 、burst load index F bi and sensitive data exposure rate Sen Ex to obtain the first communication status index C sc ; when the first communication status index C sc is lower than 0.8, trigger and implement the first governance strategy;

[0106] Step 3: After implementing the first governance strategy, re-collect and obtain the second communication timing data between several groups of master nodes and branch nodes, and conduct in-depth analysis and calculation through the communication status model to obtain the abnormal behavior response delay Atr l 、node working switching rate S wr 、dual-node mutual trust index T xi and isolation efficiency index G; and calculate the comprehensive governance index Z through correlation;

[0107] Step 4: Preset an efficiency threshold Zy, compare the comprehensive governance index Z with the preset efficiency threshold Zy, and determine the data governance level status based on the result:

[0108] If the governance status is the first high-efficiency level, maintain the current dual-node management architecture;

[0109] If the governance status is the second critical level, optimize the task allocation and load balancing of branch nodes;

[0110] If the governance status is the third low-efficiency level, adjust the data collaboration strategy between master nodes and branch nodes, and optimize the container isolation parameters.

[0111] In this embodiment, the present invention can accurately identify metrics such as data drift rate and data coverage difference by monitoring the communication timing data between several groups of master nodes and branch nodes in a hydropower station in real time and performing in-depth analysis using the neural network model CNN, and can evaluate the data synchronization situation in real time. This can greatly improve the accuracy of data synchronization between nodes and ensure data consistency and timeliness. By constructing a communication state model and performing in-depth analysis on the data, the present invention can accurately evaluate the communication state and identify key communication metrics such as the uplink-downlink transmission ratio and burst load index. Once it is found that the communication state is unstable (for example, the first communication state index C sc is lower than 0.8), the system can immediately trigger the first governance strategy to optimize the data transmission path and adjust the network load, avoiding system failures caused by communication problems.

[0112] In infrastructure systems such as hydropower stations, data security is particularly important. Traditional methods often carry the risk of exposing sensitive data. The present invention can identify the risk of data leakage in advance by monitoring the sensitive data exposure rate and combining metrics such as data drift rate and data coverage difference, and take targeted measures for governance. By optimizing data transmission and isolation strategies, the present invention significantly improves data security and reduces the possibility of exposing sensitive data. Due to the wide distribution of nodes in hydropower stations, traditional data governance methods often cannot perform load balancing in a timely manner, resulting in some nodes being overloaded while other nodes are idle. By implementing a dual-node management architecture and optimizing the task allocation and load balancing of branch nodes, the system can automatically adjust the working status of nodes based on real-time data analysis, thereby effectively optimizing resource allocation and improving the overall operating efficiency of the system. In traditional technologies, data governance often relies on preset static rules and is difficult to cope with the dynamic changes of the system. By introducing a comprehensive governance index Z and comparing it with a preset efficiency threshold Zy, the present invention can automatically adjust the system's strategy according to different data governance level states (such as the first high-efficiency level, the second critical level, and the third low-efficiency level). If the governance state is the low-efficiency level, the system can optimize the data collaboration strategy between the master node and the branch node and adjust the container isolation parameters to flexibly adapt to different requirements of the system. This adaptive ability enables the hydropower station system to maintain efficient and stable operation in a changing environment.

[0113] Embodiment 2

[0114] This embodiment is an explanatory description based on Embodiment 1. Specifically, a resource database is established based on the collected data to store various types of data of each node in the hydropower station;

[0115] The various types of data include but are not limited to: master node data: reservoir water level, flow rate, power generation, equipment operation status, and energy output power; branch node data: equipment health status, real-time grid load, flow rate fluctuations, sensor data, and security monitoring data;

[0116] Data cleaning and standardization: Clean various types of data collected from the main node and branch nodes, including: removing abnormal data points, filling in missing data, and standardizing the cleaned data based on dimensionless processing techniques such as Z-Score standardization.

[0117] In this embodiment, by establishing a resource database and performing cleaning and standardization processing on various types of data, the quality and consistency of hydropower station data can be effectively improved, ensuring the accuracy and reliability of subsequent analysis, supporting precise decision-making and optimization strategies, and enhancing the system operation efficiency and security.

[0118] Embodiment 3

[0119] This embodiment is an explanatory description based on Embodiment 1. Specifically, Step 2 includes:

[0120] S21. Collect several groups of first communication timing data of the main node and branch nodes, and establish a communication state model through the neural network model CNN. The communication state model contains several convolutional layers for extracting local features of the first communication timing data. Each layer of convolutional kernel scans the input first communication timing data and applies a convolution operation in the local area, including weighted summation and bias terms;

[0121] The several convolutional layers include the feature maps extracted by each layer of convolutional kernel, which represent the local or global features of the input data at that layer;

[0122] S22. Extract the output feature values of each layer of the main node and the output feature values of each layer of the branch node in the feature maps extracted by each layer of convolutional kernel, and calculate the feature difference value δ of the main node in the j-th layer of the convolutional neural network through the following formula m,j and the feature difference value δ of the branch node in the j-th layer of the convolutional neural network s,j :

[0123]

[0124]

[0125] In the formula, is the output feature value of the main node at the i-th moment convolutional layer in the j-th layer extracted by the communication state model, is the output feature value of the main node at the convolutional layer at the previous moment of the i-th moment in the j-th layer extracted by the communication state model. The feature difference value δ of the main node in the j-th layer of the convolutional neural network m,j , indicating the change of the main node in the feature map of this layer;

[0126] In the formula, The output eigenvalue of the convolutional layer of the branch node extracted for the communication state model at the j-th layer at the i-th moment The output eigenvalue of the convolutional layer of the branch node extracted for the communication state model at the j-th layer at the (i-1)-th moment, and the feature difference value δ of the branch node in the convolutional neural network at the j-th layer s,j , indicating the change of the feature map of the branch node in this layer;

[0127] S23. According to the feature difference value δ of the main node in the convolutional neural network at the j-th layer m,j and the feature difference value δ of the branch node in the convolutional neural network at the j-th layer s,j , calculate the data drift rate D fl , the data drift rate D fl represents the change trend of the synchronized data between the main node and the branch node, and is obtained by calculation through the following formula:

[0128]

[0129] In the formula, f m,i represents the data value of the main node at the i-th moment, and f s,i represents the data value of the branch node at the i-th moment: N represents the size of the time window, specifically the number of data points in the time series. It is set to monitor 1 hour of data, sampled once per second, and N is 3600; δ m,j represents the feature difference value of the main node in the convolutional neural network at the j-th layer; δ s,j represents the feature difference value of the branch node in the convolutional neural network at the j-th layer, m represents the m-th main node, s represents the s-th branch node, L represents the total number of convolutional layers, and α j represents the weight coefficient of the j-th convolutional layer; (f m,i -f s,i ) represents the desynchronization difference term of the data between the two nodes, synthesizes the differences of the convolutional feature maps of each layer in the communication state model network and gives different weights according to the different layers. The convolutional feature difference δ m,j -δ s,j of each layer will affect the final drift rate and is weighted by the weight coefficient α j so that the influence of some layers is more prominent.

[0130] In this embodiment, by extracting the feature difference values of the main node and the branch node in different convolutional layers, the feature changes of the two during the communication process can be quantified, the dynamic differences and potential problems between the nodes can be accurately identified, providing strong support for subsequent anomaly detection and optimization strategies, and improving the accuracy of data synchronization and processing. Calculate the data drift rate D flIt can effectively measure the data synchronization between the master node and the branch node. By comprehensively considering the differences in the feature maps of the convolutional layers and assigning different weights, this method can accurately capture the data drift trend, thereby helping to quickly identify the unstable factors in the communication state and improving the data processing efficiency and system stability.

[0131] Embodiment 4

[0132] This embodiment is an explanatory description based on Embodiment 3. Specifically, Step 2 further includes:

[0133] S24. Extract the first communication timing data, and establish a communication state model through the neural network model CNN, and deeply calculate and analyze to obtain the data coverage difference index D cd , the uplink and downlink transmission ratio U pr , the burst load index F bi and the sensitive data exposure rate Sen Ex ;

[0134] S241. The data coverage difference index D cd is calculated and obtained through the following formula:

[0135]

[0136] In the formula, C m,i represents the coverage rate of the master node at the i-th moment, and C s,i represents the coverage rate of the branch node at the i-th moment; N represents the time window size; Z m,j represents the coverage difference value of the master node in the j-th convolutional neural network; X s,j represents the coverage difference value of the branch node in the j-th convolutional neural network, and α j represents the weight coefficient of the j-th convolutional layer; calculating the data coverage difference index D cd can measure the difference in data coverage between the master node and the branch node, thereby identifying the coverage blind spots or inconsistent areas in the data transmission process and providing guarantee for data integrity and accuracy.

[0137] S242. The uplink and downlink transmission ratio U pr is calculated and obtained through the following formula:

[0138]

[0139] In the formula, D up,i represents the uplink data volume at the i-th moment, specifically the transmission volume from the master node to the branch node, and D down,i represents the downlink data volume at the i-th moment, specifically the transmission volume from the branch node to the master node, and N represents the time window size; by calculating the uplink and downlink transmission ratio U pr, it can monitor the imbalance of data transmission direction in real time, help identify network bottlenecks or transmission efficiency problems, thereby optimizing data flow and communication paths, and improving the overall performance of the system.

[0140] S243. The burst load index F bi is calculated and obtained through the following formula:

[0141]

[0142] In the formula, the burst load index F bi is used to measure the change rate of the load of the branch node under instantaneous high concurrency. L s,i represents the CPU usage rate of the branch node at time i, and Δt represents the time interval. represents taking the maximum value of the maximum load change rate in the entire time window; the calculation of the burst load index F bi can quickly identify the load changes of nodes under high concurrency, especially the pressure situation of branch nodes, and timely adopt load balancing or resource adjustment strategies to effectively prevent system overload or collapse and ensure the stable operation of the system.

[0143] S244. The sensitive data exposure rate Sen Ex is calculated and obtained through the following formula:

[0144]

[0145] In the formula, D ex,i represents the amount of sensitive data exposed at time i, represents the transmission volume from the master node to the branch node, and P i represents the sensitivity weight coefficient of the data type, including:

[0146] The first sensitivity data weight coefficient, including:

[0147] Master node data: reservoir water level 1.0, flow rate 0.9, power generation 0.8, equipment operation status 0.95, and energy output power 0.85;

[0148] The second sensitivity data weight coefficient, including:

[0149] Branch node data: equipment health status 0.9, real-time grid load 0.8, flow rate fluctuation 0.7, sensor data 0.75, and security monitoring data 0.9;

[0150] Among them, D down,i represents the amount of downstream data at time i, represents the transmission volume from the branch node to the master node, and N represents the time window size. The sensitive data exposure rate Sen ExThe calculation helps to evaluate the security exposure risk of sensitive data during communication. Based on the sensitivity weights of different data, it can accurately identify high-risk sensitive data transmissions, thereby providing real-time warnings and protection measures for data protection and privacy security.

[0151] S25. Extract the data drift rate D between each group of master nodes and branch nodes fl and the data coverage difference index D cd and the uplink-downlink transmission ratio U pr and the burst load index F bi and the sensitive data exposure rate Sen Ex , after dimensionless processing, the first communication status index C is calculated through the following associated formula sc :

[0152]

[0153] In the formula, a1, a2, a3, a4, and a5 respectively represent the data drift rate D fl , the data coverage difference index D cd , the burst load index F bi , the sensitive data exposure rate Sen Ex and the uplink-downlink transmission ratio U pr 's weight coefficients, and the sum of the weight coefficients is 1.

[0154] In this embodiment, by performing dimensionless processing and weighted synthesis on the data drift rate D fl , the data coverage difference index D cd , the burst load index F bi , the sensitive data exposure rate Sen Ex and the uplink-downlink transmission ratio U pr , the first communication status index C sc is obtained, which can comprehensively evaluate the overall performance of the communication status, help to timely identify potential communication problems, optimize network performance and improve the accuracy and reliability of data governance.

[0155] Embodiment 5

[0156] Please refer to Figure 2 , this embodiment is an explanatory description based on Embodiment 4. Specifically, step two further includes: S26. Preset the communication threshold to 0.8, and compare the first communication status index C sc with the communication threshold. When the first communication status index C sc < 0.8, it means that the communication status between the current master node and branch node is unqualified, and the first governance strategy is triggered, including:

[0157] S2611. Use Docker volumes to allocate independent storage volumes to each master node and branch node, which are used to isolate the storage volumes of each node through container volumes and set access permissions;

[0158] S2612, increase the current synchronization frequency of the current master node and branch node by 30%, increase the data transmission compression rate of the current master node and branch node by 30%, and increase the current upload bandwidth by 10%-30%;

[0159] S2613, encrypt the data type transmitted by the main node and the branch node with a sensitivity weight coefficient Pi exceeding 0.9 using the encryption algorithm AES with a 128-bit key length, generate a first encryption instruction and implement it;

[0160] When the first communication state index C sc When ≥0.8, it means that the communication status between the current master node and the branch node is qualified and continues to be monitored.

[0161] In this embodiment, by setting the communication threshold and setting the first communication state index C sc In contrast, it is possible to effectively determine whether the communication status is qualified, trigger the first governance strategy in time, and ensure the stability and efficiency of the hydropower station communication system. This method can help automatically monitor and optimize the communication quality, avoid manual intervention, and improve the system's adaptability. By allocating independent storage volumes to each master node and branch node and setting access rights, effective isolation and management of data storage can be achieved, reducing the risk of data conflict or leakage, and improving the data security and stability of the system, especially in a multi-node environment. By increasing the synchronization frequency, improving the data transmission compression rate and upload bandwidth, the data transmission efficiency can be significantly improved, the transmission delay can be reduced, especially in the case of unsatisfactory network conditions, ensuring the real-time and accuracy of the data, and optimizing the overall communication performance. Encrypting sensitive data, especially data with a sensitivity weight coefficient exceeding 0.9, can effectively prevent data leakage and unauthorized access, improve the security and privacy protection level of the system, and ensure that the data is strictly protected during transmission and meets compliance requirements. The first governance strategy can quickly take a series of effective optimization measures when the communication status is unqualified, improve the data communication efficiency, security and stability of the hydropower station system, ensure that the best performance can be maintained in various environments, and enhance the system's self-regulation and emergency response capabilities.

[0162] Example 6

[0163] This embodiment is an explanation of the embodiment 1. Specifically, step 3 includes:

[0164] S31. After the implementation of the first governance strategy, re-collect and obtain several groups of second communication timing data between the master node and the branch nodes. Re-collecting the second communication timing data includes the abnormal behavior logs of the branch nodes, which can provide more detailed historical data for subsequent anomaly detection and analysis, ensure comprehensive monitoring of node behavior after the implementation of the governance strategy, and promptly capture potential security hazards and performance issues.

[0165] S32. Then input the second communication timing data into the communication status model for in-depth analysis and calculation to obtain the abnormal behavior response delay Atr l , the node working switching rate S wr , the dual-node mutual trust index T xi and the isolation efficiency index G;

[0166] S321. The abnormal behavior response delay Atr l is calculated and obtained through the following formula:

[0167]

[0168] In the formula, T resp,i represents the response time of the master node to the abnormal behavior of the branch node at the i-th moment, and the abnormal behavior includes data packet replay or forgery behavior; T detect,i represents the time when the abnormal behavior of the branch node at the i-th moment is detected, and R represents the total number of abnormal behavior monitoring times; calculating the abnormal behavior response delay Atr l can quantify the response speed of the system to abnormal events, provide specific data support for the evaluation of the effectiveness of the governance strategy, help optimize the processing process and timely response ability of the system to abnormal events, and improve the robustness of the system.

[0169] S322. The node working switching rate S wr is calculated and obtained through the following formula:

[0170]

[0171] In the formula, ΔR ole,i represents the change in the role switch between the master node and the branch node at the i-th moment. If the master node becomes the branch node role, the value is 1; if the branch node becomes the master node role, the value is -1; if there is no switch, the value is 0; V represents the total number of role switches; by calculating the node working switching rate S wr , the frequency and stability of the role switch between the master node and the branch node can be evaluated, ensuring the smooth operation of the system under high load or abnormal conditions, and avoiding performance fluctuations and increased management complexity caused by frequent switching.

[0172] S323. The dual-node mutual trust index T xi is calculated and obtained through the following formula;

[0173]

[0174] In the formula, T success,i represents the number of successfully transmitted tasks of the master node and the branch node at the i-th moment. Whether a specific task is successfully completed or synchronized successfully, E i represents the weight coefficient of task success at the i-th moment, which is related to factors such as the complexity of the task or the amount of data. N represents the time window size; through the calculation of the mutual trust index T xi between the two nodes, the degree of cooperation between the master node and the branch node can be quantified, the stability and reliability of the system task transmission can be evaluated, a quantifiable basis can be provided for system optimization, and data consistency and transmission success rate can be improved.

[0175] S324. The isolation efficiency index G is obtained by calculating through the following formula:

[0176]

[0177] In the formula, D isolated represents the amount of data successfully isolated, D total represents the total amount of data; E encrypt represents the amount of encrypted data, T transmit represents the data transmission time after encryption; S protected represents the sensitive data successfully protected, S total represents the amount of sensitive data; E switch represents the error rate caused by node switching; calculating the isolation efficiency index G helps to evaluate the data isolation and protection effects in abnormal situations, ensure the secure transmission and isolation of sensitive data, avoid the risk of data leakage, and improve the overall security and data protection capabilities of the system.

[0178] S33. And the abnormal behavior response delay Atr l and the node working switching rate S wr , the mutual trust index T xi between the two nodes, and the isolation efficiency index G, after dimensionless processing, the comprehensive governance index Z is obtained by calculating through the following associated formula:

[0179]

[0180] In the formula, b1, b2, b3, and b4 respectively represent the abnormal behavior response delay Atr l , the node working switching rate S wr , the mutual trust index T xi between the two nodes, and the weight coefficients of the isolation efficiency index G, and the sum of the weight coefficients is 1.

[0181] In this embodiment, by comprehensively considering the abnormal behavior response delay Atrl 1. Node working switching rate S wr 2. Mutual trust index T of two nodes xi and isolation efficiency index G, calculate the comprehensive governance index Z to make the governance effect more comprehensive and accurate. After dimensionless processing, it can reduce the interference of the unit differences of various indicators, make the weights of each indicator more balanced, and improve the stability and reliability of the overall system.

[0182] Example 7

[0183] Please refer to Figure 2 , this example is an explanatory note based on Example 1. Specifically, Step 4 includes:

[0184] Preset the efficiency threshold Zy, and compare the comprehensive governance index Z with the preset efficiency threshold Zy to obtain the following judgment results, including:

[0185] When the comprehensive governance index Z > the efficiency threshold Zy, it means that the current operation of the two nodes meets the standard, and the governance status is the first high-efficiency level;

[0186] When the efficiency threshold Zy * 80% ≤ the comprehensive governance index Z ≤ the efficiency threshold Zy, it means that the current operation of the two nodes does not meet the standard, and the governance status is the second critical level;

[0187] When the comprehensive governance index Z < the efficiency threshold Zy * 80%, it means that the current operation of the two nodes does not meet the standard, and the governance status is the third low-efficiency level;

[0188] If the governance status is the first high-efficiency level, maintain the current two-node management architecture; when the system reaches the first high-efficiency level, it can continuously maintain the current management architecture, reduce unnecessary resource waste, and ensure the efficient operation of the system.

[0189] When the system is in a critical or low-efficiency state, corresponding governance strategies will be triggered to avoid the continuous low-efficiency operation of the system.

[0190] If the governance status is the second critical level, optimize the task allocation and load balancing of branch nodes to generate the second governance strategy;

[0191] If the governance status is the third low-efficiency level, adjust the data collaboration strategy between the main node and branch nodes, optimize the container isolation parameters, and generate the third governance strategy.

[0192] The second governance strategy includes: based on the real-time monitoring results, reducing the communication task volume between the current master node and the busiest branch node by 15%-20% and transferring it to the low-load node; increasing the synchronization frequency by 30% compared with the first governance strategy, and here maintaining the synchronization frequency increase of 15%-20%; using AES-256 encryption instead of AES-128, and encrypting the data types with a sensitivity weight coefficient Pi exceeding 0.8 in the data transmitted between the master node and the branch node using the AES encryption algorithm with a 256-bit key length, generating a second encryption instruction and implementing it;

[0193] The third governance strategy includes: The second governance strategy includes: based on the real-time monitoring results, reducing the communication task volume between the current master node and the busiest branch node by 21%-30% and transferring it to the low-load node; increasing the synchronization frequency by 40% compared with the first governance strategy and maintaining the synchronization frequency increase of 21%-30%;

[0194] Using Docker volumes configuration, each node is separately allocated a storage volume; reconfiguring the independent volume capacity of the inefficient nodes and expanding the storage volume capacity by 30%-50%;

[0195] Using AES-256 encryption instead of AES-128, and increasing the block transfer of data before encryption, and encrypting the data types with a sensitivity weight coefficient Pi exceeding 0.8 in the data transmitted between the master node and the branch node using the AES encryption algorithm with a 256-bit key length, generating a third encryption instruction and implementing it.

[0196] In this embodiment, the second governance strategy and the third governance strategy adjust task allocation based on real-time monitoring results, reduce the communication task volume between the master node and the busiest branch node, and transfer tasks to low-load nodes. This strategy effectively avoids overload of some nodes, improves the load balance of the overall system, and thus enhances the stability and response ability of the system. In the second governance strategy, compared with the first governance strategy, the increase range of the synchronization frequency is moderately reduced (15%-20%). By maintaining a reasonable synchronization frequency, the network pressure and delay caused by frequent synchronization are reduced. In the third governance strategy, the synchronization frequency is further enhanced (increased by 40%) to cope with more complex situations and ensure the timeliness and accuracy of data transmission. By upgrading the AES-128 encryption algorithm to AES-256 and enhancing the encryption processing of sensitive data (especially data with a weight coefficient Pi exceeding 0.8), the security of data transmission is improved, and data leakage or tampering during transmission is avoided. Especially in the third governance strategy, combined with data block transmission and enhanced encryption, the protection ability of the system for sensitive data is effectively improved. In the third governance strategy, Docker volumes are used to allocate independent storage volumes to each node, and the capacity of the independent volumes of inefficient nodes is expanded (30%-50%). This makes storage management more flexible and avoids data transmission delays or losses caused by storage bottlenecks. At the same time, through capacity adjustment, the system can adapt to the storage requirements under different load conditions and enhance the overall elasticity of the system. The second governance strategy and the third governance strategy adopt different governance measures according to the actual operation data and load conditions, such as adjusting the synchronization frequency, task allocation, storage configuration, and encryption intensity, to ensure that the system can achieve the best performance under different load and security requirements. This adaptive governance method enables the system to have better scalability and processing ability when dealing with high-load or data-intensive tasks. Through reasonable resource allocation, optimized storage management, and enhanced data encryption and other measures, the second and third governance strategies enhance the stability and scalability of the system while improving the system performance, enabling the system to adapt to the demand fluctuations under different operating states and maintain efficient operation in the long term.

[0197] Please refer to Figure 3 , a data integration governance system integrating container isolation and dual-node management, including:

[0198] A database establishment unit, used to establish a database by real-time monitoring various types of data in several groups of master nodes and branch nodes of a hydropower station;

[0199] A node communication data acquisition unit, which real-time monitors the first communication timing data between several groups of master nodes and branch nodes of a hydropower station;

[0200] A communication status model establishment unit is used to establish a communication status model using the neural network CNN, deeply analyze the collected data, and obtain the data drift rate D between each group of master nodes and branch nodes fl , the data coverage difference index D cd , the uplink-downlink transmission ratio U pr , the burst load index F bi and the sensitive data exposure rate Sen Ex ; and correlate the data drift rate D fl , the data coverage difference index D cd , the uplink-downlink transmission ratio U pr , the burst load index F bi and the sensitive data exposure rate Sen Ex to obtain the first communication status index C sc ; when the first communication status index C sc is lower than 0.8, trigger and implement the first governance strategy;

[0201] A second collection and analysis unit is used to, after the implementation of the first governance strategy, re-collect and obtain several groups of second communication timing data between the master node and the branch node, and deeply analyze and calculate through the communication status model to obtain the abnormal behavior response delay Atr l , the node working switching rate S wr , the two-node mutual trust index T xi and the isolation efficiency index G; and calculate the comprehensive governance index Z through correlation;

[0202] A governance level evaluation unit is used to preset an efficiency threshold Zy, compare the comprehensive governance index Z with the preset efficiency threshold Zy, determine the data governance level status according to the result, and generate corresponding governance strategies.

[0203] In this embodiment, the data integration governance system that integrates container isolation and dual-node management of the present invention provides an efficient, flexible and secure data governance solution by integrating real-time monitoring, data analysis and intelligent decision-making mechanisms. The system first uses the database establishment unit to collect and store various types of data of the main node and branch nodes of the hydropower station in real time, forming a complete historical data record, providing a basis for subsequent analysis. The node communication data collection unit monitors the communication timing data between the main node and the branch node in real time to ensure the timeliness and accuracy of data collection.

[0204] On this basis, the communication status model establishment unit deeply analyzes the data using the neural network CNN technology, calculates key indicators including data drift rate, data coverage difference index, uplink-downlink transmission ratio, burst load index, and sensitive data exposure rate, etc., and forms the first communication status index. This index can effectively reflect the qualification of the communication status between the main node and the branch node. When the index is lower than the set threshold of 0.8, the first governance strategy is automatically triggered to improve the system reliability and data security.

[0205] Through the second acquisition and analysis unit, after the implementation of the first governance strategy, the system re-acquires and deeply analyzes the communication timing data, calculates the abnormal behavior response delay, node working switching rate, mutual trust index between double nodes, and isolation efficiency index G, etc., to further improve the data governance strategy. Finally, the governance level evaluation unit compares the comprehensive governance index with the preset efficiency threshold, dynamically evaluates the system operation status, and generates the corresponding second governance strategy and second governance strategy according to the governance level.

[0206] Generally speaking, the system of the present invention effectively improves the security, stability and operation efficiency of the hydropower station communication system through intelligent data monitoring and analysis, container isolation technology and dynamic governance strategies, and can flexibly adjust the strategies according to different loads and security requirements to ensure the efficient and reliable operation of the system and adapt to the changing operation environment.

[0207] The setting of the size of the threshold is for the convenience of comparison. Regarding the size of the threshold, it depends on the amount of sample data and the number of base numbers set by those skilled in the art for each group of sample data; as long as it does not affect the proportional relationship between the parameters and the quantified values.

[0208] The above formulas are all obtained by collecting a large amount of data for software simulation and selecting a formula close to the true value. The coefficients in the formulas are set by those skilled in the art according to the actual situation. The above is only the preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and should be covered by the protection scope of the present invention.

Claims

1. A data integration management method integrating container isolation and dual-node management, characterized in that: The following steps are involved: Step 1: Establish a database by real-time monitoring of various data in several groups of main nodes and branch nodes of the hydropower station; Step 2: Real-time monitoring of the first communication time series data between several groups of main nodes and branch nodes of the hydropower station; using the neural network model CNN to establish and train the communication state model, and perform in-depth analysis and calculation on the first communication time series data to obtain the data drift rate D between each group of main nodes and branch nodes. fl , Data coverage difference index D cd , uplink and downlink transmission ratio U pr , burst load index F bi And sensitive data exposure rate Sen Ex ; and the data drift rate D fl , Data coverage difference index D cd , uplink and downlink transmission ratio U pr , burst load index F bi and sensitive data exposure rate Sen Ex Perform association and obtain the first communication state index C sc ; When the first communication state index C sc When it is lower than 0.8, the first governance strategy is triggered and implemented; Step 3: After the first governance strategy is implemented, re-collect and obtain the second communication timing data between several groups of master nodes and branch nodes, and deeply analyze and calculate through the communication state model to obtain the abnormal behavior response delay Atr l , node work switching rate S wr , Dual-node mutual trust index T xi and isolation efficiency index G; And the comprehensive governance index Z is calculated in association; Step 4: Preset the efficiency threshold Zy, and compare the comprehensive governance index Z with the preset efficiency threshold Zy, and determine the data governance level status based on the results: If the governance status is the first efficient level, maintain the current dual-node management architecture; If the governance status is at the second critical level, optimize branch node task allocation and load balancing; If the governance status is the third inefficient level, adjust the data collaboration strategy between the master node and the branch node, and optimize the container isolation parameters.

2. According to claim 1, a data integration management method integrating container isolation and dual-node management is characterized in that: A resource database is established based on the collected data to store various data of each node of the hydropower station; Various types of data include but are not limited to: Main node data: reservoir water level, flow, power generation, equipment operating status and energy output power; Branch node data: equipment health status, real-time grid load, flow fluctuations, sensor data and security monitoring data; Data cleaning and standardization: Clean all types of data collected by the main node and branch nodes, including: removing abnormal data points, filling missing data, and standardizing the cleaned data based on dimensionless processing techniques such as Z-Score standardization.

3. According to claim 1, a data integration management method integrating container isolation and dual-node management is characterized in that: Step 2 includes: S21, collecting first communication time series data of several groups of main nodes and branch nodes, and establishing a communication state model through a neural network model CNN, wherein the communication state model includes several convolution layers for extracting local features of the first communication time series data, and the convolution kernel of each layer scans the input first communication time series data and applies a convolution operation in a local area, including a weighted sum bias term; Several convolutional layers include feature maps extracted by convolution kernels at each layer, which represent the local or global features of the input data at that layer; S22, extract the output feature value of each layer of the main node and the output feature value of each layer of the branch node in the feature graph extracted by each layer of the convolution kernel, and calculate the feature difference value δ of the main node in the jth layer of the convolutional neural network by the following formula m,j and the feature difference value δ of the branch node in the jth layer of the convolutional neural network s,j : In the formula, is the output feature value of the convolutional layer of the master node at the jth layer at the i-th time extracted by the communication status model, is the output feature value of the convolutional layer of the master node at the jth layer before the ith moment extracted by the communication state model, and the feature difference value δ of the convolutional neural network of the master node at the jth layer m,j , indicating the change of the feature graph of the master node in this layer; In the formula, is the output feature value of the convolutional layer at the jth layer and the i-th time of the branch node extracted by the communication state model. is the output feature value of the convolutional layer of the branch node in the jth layer at the i-th moment extracted by the communication state model, and the feature difference value δ of the convolutional neural network of the branch node in the jth layer s,j , represents the change of the feature map of the branch node at this layer; S23, based on the feature difference value δ of the main node in the jth layer of the convolutional neural network m,j and the feature difference value δ of the branch node in the jth layer of the convolutional neural network s,j , calculate the data drift rate D fl , data drift rate D fl Indicates the change trend of the synchronized data between the master node and the branch node, which is calculated using the following formula: In the formula, f m,i represents the data value of the master node at the i-th moment, f s,i Represents the data value of the branch node at the i-th moment: N represents the time window size, specifically the number of data points in the time series, δ m,j Represents the feature difference value of the main node in the jth layer of the convolutional neural network; δ s,j represents the feature difference value of the branch node in the jth layer of the convolutional neural network, m represents the mth main node, s represents the sth branch node, L represents the total number of convolutional layers, and α j represents the weight coefficient of the j-th convolutional layer; (f m,i -f s,i ) represents the asynchronous difference item of data between two nodes, The difference of convolution feature maps of each layer in the communication state model network is synthesized, and different weights are given according to the number of layers. The difference of convolution feature maps of each layer δ m,j -δ s,j Will affect the final drift rate.

4. According to claim 3, a data integration management method integrating container isolation and dual-node management is characterized in that: Step 2 also includes: S24, extract the first communication time series data, and establish a communication state model through a neural network model CNN, and obtain a data coverage difference index D through deep calculation and analysis. cd , uplink and downlink transmission ratio U pr , burst load index F bi and sensitive data exposure rate Sen Ex ; S241, the data coverage difference index D cd Calculated by the following formula: In the formula, C m,i represents the coverage rate of the master node at the i-th moment, C s,i represents the coverage rate of the branch node at the i-th moment: N represents the time window size; Z m,j Z represents the coverage difference value of the master node in the jth layer of the convolutional neural network; s,j represents the coverage difference value of the branch node in the jth layer of the convolutional neural network, α j Represents the weight coefficient of the j-th convolutional layer; S242, the uplink and downlink transmission ratio U pr Calculated by the following formula: Where D up,i represents the amount of uplink data at the i-th moment, specifically the amount of transmission from the master node to the branch node, D down,i represents the amount of downlink data at the i-th moment, specifically the amount of data transmitted from the branch node to the main node, and N represents the time window size; S243, the burst load index F bi Calculated by the following formula: Where, the burst load index F bi It is used to measure the rate of change of branch node load under instantaneous high concurrency conditions, L s,i represents the CPU usage of the branch node at time i, Δt represents the time interval, It means to find the maximum value of the maximum load change rate in the entire time window; S244, the sensitive data exposure rate Sen Ex Calculated by the following formula: Where D ex,i represents the amount of exposed sensitive data at the i-th moment, represents the amount of transmission from the main node to the branch node, and P i Indicates the data type sensitivity weight coefficient, including: The first sensitivity data weight coefficient includes: Main node data: reservoir water level 1.0, flow 0.9, power generation 0.8, equipment operation status 0.95 and energy output power 0.85; The second sensitivity data weight coefficient includes: Branch node data: equipment health status 0.9, real-time grid load 0.8, flow fluctuation 0.7, sensor data 0.75, and security monitoring data 0.9; Among them, D down,i It represents the amount of downlink data at the i-th moment, represents the amount of transmission from the branch node to the main node, and N represents the time window size.

5. According to claim 4, a data integration management method integrating container isolation and dual-node management is characterized in that: Step 2 also includes: S25, extracting the data drift rate D between each group of main nodes and branch nodes fl , Data coverage difference index D cd , uplink and downlink transmission ratio U pr , burst load index F bi and sensitive data exposure rate Sen Ex After dimensionless processing, the first communication state index C is calculated by the following associated formula sc : Where a1, a2, a3, a4 and a5 represent the data drift rate D fl , Data coverage difference index D cd , burst load index F bi , Sensitive data exposure rateSen Ex And the uplink and downlink transmission ratio U pr The weight coefficients of , and the sum of the weight coefficients is 1.

6. According to claim 5, a data integration management method integrating container isolation and dual-node management is characterized in that: Step 2 also includes: S26, presetting the communication threshold to 0.8, and setting the first communication state index C sc Compared with the communication threshold, when the first communication state index C sc When <0.8, it means that the communication status between the current master node and the branch node is unqualified, triggering the first governance strategy, including: S2611. Use Docker volumes to allocate independent storage volumes to each master node and branch node, which are used to isolate the storage volumes of each node through container volumes and set access permissions; S2612, increase the current synchronization frequency of the current master node and branch node by 30%, increase the data transmission compression rate of the current master node and branch node by 30%, and increase the current upload bandwidth by 10%-30%; S2613, encrypt the data type transmitted by the main node and the branch node with a sensitivity weight coefficient Pi exceeding 0.9 using the encryption algorithm AES with a 128-bit key length, generate a first encryption instruction and implement it; When the first communication state index C sc When ≥0.8, it means that the communication status between the current master node and the branch node is qualified and continues to be monitored.

7. According to claim 1, a data integration management method integrating container isolation and dual-node management is characterized in that: Step three includes: S31. After the first governance strategy is implemented, re-collect and obtain second communication timing data between several groups of master nodes and branch nodes; S32, inputting the second communication timing data into the communication state model, performing in-depth analysis and calculation to obtain the abnormal behavior response delay Atr l , node work switching rate S wr , Dual-node mutual trust index T xi and isolation efficiency index G; S321, the abnormal behavior response delay Atr l Calculated by the following formula: Where, T resp,i T represents the response time of the master node to the abnormal behavior of the branch node at time i, which includes data packet replay or forgery; deetect,i It indicates the time when the abnormal behavior of the branch node at the i-th moment is detected, and R indicates the total number of abnormal behavior monitoring times; S322, the node operation switching rate S wr Calculated by the following formula: In the formula, ΔR ole,i Indicates the change in the role switch between the master node and the branch node at the i-th moment. If the master node becomes a branch node, the value is 1; if the branch node becomes a master node, the value is -1; if there is no switch, the value is 0; V represents the total number of role switches; S323, the dual-node mutual trust index T xi Calculated by the following formula; Where, T success,i Indicates the number of tasks successfully transmitted between the master node and the branch node at the i-th moment, whether the specific task is successfully completed or synchronized, E i The weight coefficient of the success of the task at the i-th moment is related to the complexity of the task or the amount of data. N represents the time window size. S324. The isolation efficiency index G is calculated by the following formula: Where D isolated Indicates the amount of successfully isolated data, D total Indicates the total amount of data; E encrypt Indicates the amount of encrypted data, T transmit Indicates the encrypted data transmission time; S protected Indicates the sensitive data that is successfully protected, S total The amount of data representing sensitive data; E switch Indicates the error rate caused by node switching; S33, and abnormal behavior response delay Atr l , node work switching rate S wr , Dual-node mutual trust index T xi And the isolation efficiency index G, after dimensionless processing, the comprehensive governance index Z is calculated by the following related formula: Where b1, b2, b3 and b4 represent the abnormal behavior response delay Atr l , node work switching rate S wr , Dual-node mutual trust index T xi And the weight coefficient of the isolation efficiency index G, and the sum of the weight coefficients is 1.

8. According to claim 1, a data integration management method integrating container isolation and dual-node management is characterized in that: Step 4 includes: The efficiency threshold Zy is preset, and the efficiency threshold Zy is preset, and the comprehensive governance index Z is compared with the preset efficiency threshold Zy to obtain the following judgment results, including: When the comprehensive governance index Z>efficiency threshold Zy, it means that the current dual-node operation meets the standard and the governance status is the first high-efficiency level; When the efficiency threshold Zy*80% ≤ comprehensive governance index Z ≤ efficiency threshold Zy, it means that the current dual-node operation does not meet the standard and the governance status is the second critical level; When the comprehensive governance index Z is less than the efficiency threshold Zy*80%, it means that the current dual-node operation is not up to standard and the governance status is the third inefficient level; If the governance status is the first efficient level, maintain the current dual-node management architecture; If the governance status is the second critical level, optimize the branch node task allocation and load balancing, and generate the second governance strategy; If the governance status is the third inefficient level, adjust the data coordination strategy between the master node and the branch node, optimize the container isolation parameters, and generate the third governance strategy.

9. A data integration management method integrating container isolation and dual-node management according to claim 8, characterized in that: The second governance strategy includes: based on real-time monitoring results, reducing the communication task volume between the current master node and the busiest branch node by 15%-20%, and transferring it to the low-load node; increasing the synchronization frequency by 30% compared to the first governance strategy, and maintaining the synchronization frequency increased by 15%-20% here; using AES-256 encryption instead of AES-128, and encrypting the data type sensitivity weight coefficient Pi transmitted by the master node and the branch node with a 256-bit key length through the encryption algorithm AES, generating and implementing the second encryption instruction; The third governance strategy includes: The second governance strategy includes: based on real-time monitoring results, reducing the communication task volume between the current master node and the busiest branch node by 21%-30%, and transferring it to the low-load node; compared with the first governance strategy, increasing the synchronization frequency by 40%, and maintaining the synchronization frequency increased by 21%-30%; Use Docker volumes configuration to allocate storage volumes to each node separately; reconfigure independent volume capacity for inefficient nodes and expand storage volume capacity by 30%-50%; Use AES-256 encryption instead of AES-128, increase the data block transmission before encryption, and encrypt the data type transmitted by the main node and branch node with a sensitivity weight coefficient Pi exceeding 0.8 through the encryption algorithm AES with a 256-bit key length, generate the third encryption instruction and implement it.

10. A data integration management system integrating container isolation and dual-node management, characterized in that: include: A database establishment unit, used to establish a database by real-time monitoring of various data in several groups of main nodes and branch nodes of the hydropower station; A node communication data acquisition unit monitors first communication time series data between several groups of main nodes and branch nodes of a hydropower station in real time; The communication status model building unit is used to build a communication status model using a neural network CNN, and conduct in-depth analysis on the collected data to obtain the data drift rate D between each group of main nodes and branch nodes. fl , Data coverage difference index D cd , uplink and downlink transmission ratio U pr , burst load index F bi and sensitive data exposure rate Sen Ex ; and the data drift rate D fl , Data coverage difference index D cd , uplink and downlink transmission ratio U pr , burst load index F bi and sensitive data exposure rate Sen Ex Perform association and obtain the first communication state index C sc ; When the first communication state index C sc When it is lower than 0.8, the first governance strategy is triggered and implemented; The second collection and analysis unit is used to re-collect and obtain the second communication timing data between several groups of master nodes and branch nodes after the first governance strategy is implemented, and deeply analyze and calculate the communication state model to obtain the abnormal behavior response delay Atr l , node work switching rate S wr , Dual-node mutual trust index T xi and isolation efficiency index G; And the comprehensive governance index Z is calculated in association; The governance level evaluation unit is used to preset the efficiency threshold Zy, and compare the comprehensive governance index Z with the preset efficiency threshold Zy, determine the data governance level status based on the results, and generate a corresponding governance strategy.

Citation Information

Cited By

  • Intelligent photographing lamp box automatic photographing system based on order number driving

    CN121283975A

  • Intelligent photographing light box automatic shooting system based on order number driving

    CN121283975B