Hydropower data analysis system and method based on data blood relationship

By using a data lineage-based hydropower data analysis method, the correlation strength between equipment and the region is automatically calculated, solving the data link problem caused by the dynamic changes of equipment in temporary building complexes. This achieves efficient and accurate data management and reduces operation and maintenance costs.

CN121807922BActive Publication Date: 2026-04-28GUIZHOU WUJIANG HYDROPOWER DEV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
GUIZHOU WUJIANG HYDROPOWER DEV
Filing Date
2026-03-11
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

In the management of water and electricity data in temporary building complexes, existing technologies are unable to cope with the characteristics of non-fixed equipment installation, which leads to data link breakage or mismatch. Furthermore, the lack of dynamic optimization and traceability mechanisms results in high labor costs and low efficiency.

Method used

The hydropower data analysis method based on data lineage acquires data from metering nodes and analysis units, defines lineage feature dimensions, calculates lineage association strength, optimizes mapping schemes, and generates lineage traceability files to achieve automated equipment-region association.

Benefits of technology

It reduces the labor costs of operating and maintaining temporary building complexes, improves the accuracy and adaptability of water and electricity data mapping, provides a reliable data foundation, and supports subsequent metering statistics and energy consumption analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121807922B_ABST
    Figure CN121807922B_ABST
Patent Text Reader

Abstract

The application discloses a water and electricity data analysis system and method based on data blood relationship, relates to the technical field of data analysis, and comprises the following steps: obtaining a metering node, an analysis unit and deployment event data; defining a blood relationship characteristic dimension and obtaining corresponding original support data; obtaining a blood relationship benchmark system based on the analysis unit and metering node data; defining a blood relationship characteristic dimension calculation rule and configuring a weight; calculating blood relationship correlation strength based on the blood relationship characteristic dimension calculation rule; determining an initial mapping scheme based on the blood relationship correlation strength; defining a constraint rule to optimize the initial mapping scheme to obtain an optimal mapping scheme; solidifying the optimal mapping scheme into a final mapping relationship, generating a blood relationship tracing archive and an optimization process report and feeding back. The application can effectively improve the situation that the data link is unstable due to the non-fixed installation of temporary building water and electricity equipment and the short use cycle in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data analysis technology, specifically to a hydropower data analysis system and method based on data lineage. Background Technology

[0002] The water and electricity systems of temporary building complexes (such as exhibition halls and construction sites) are characterized by their temporary and flexible nature. Their water and electricity equipment is mostly modular and temporarily connected, encompassing key components such as metering instruments, distribution switches, and sensor terminals. The usage cycle in these scenarios typically ranges from several days to several months, requiring dynamic adjustments to equipment installation locations based on actual needs, often accompanied by frequent additions, removals, and relocations of equipment. Water and electricity data, as the core support for operation and management, needs to achieve real-time collection and analysis of information such as energy consumption metering, equipment operating status, and load distribution to ensure the stability of power and water supply, control operating costs, and meet safety regulatory requirements. Because the physical layout of temporary building complexes lacks fixed planning, the network access methods and power distribution topologies of water and electricity equipment are also dynamically changing. This places special demands on the continuity and reliability of data links; data must be accurately linked to the corresponding functional areas to provide effective support for subsequent statistical analysis and operation and maintenance decisions.

[0003] Currently, the management of water and electricity data for temporary building complexes mostly adopts traditional, fixed-scenario adaptation solutions. These solutions rely heavily on manual configuration of the association between equipment and the analysis area, or on establishing data links through simple network IP binding and physical location labeling. Existing technologies generally suffer from the following shortcomings: First, they struggle to handle the non-fixed nature of equipment installation. Manual configuration cannot keep up with equipment migration and additions in real time, easily leading to data link breaks or mismatches, resulting in problems such as "equipment has been moved, but data is still associated with the original area." Second, they rely on long-term accumulated benchmark data to build association standards, which cannot adapt to the short-term use of temporary building complexes. In new scenarios, there is a lack of effective association criteria, leading to delays in data link establishment. Furthermore, existing technologies lack dynamic optimization and traceability mechanisms, resulting in low efficiency in troubleshooting data link anomalies and requiring significant manual maintenance costs, further reducing application adaptability in temporary scenarios. Summary of the Invention

[0004] The purpose of this invention is to provide a hydropower data analysis system and method based on data lineage, so as to solve the problems raised in the prior art.

[0005] To achieve the above objectives, the present invention provides the following technical solution: a hydropower data analysis method based on data lineage, the method comprising the following steps:

[0006] Step 1: Obtain metering node, analysis unit, and deployment event data; define lineage characteristic dimensions and obtain corresponding raw supporting data;

[0007] The metering node refers to the hydropower metering equipment deployed in a single cycle, such as voltage transformers, current transformers, and power meters; the analysis unit refers to the analysis object in a fixed area, such as by dividing it into defined zones.

[0008] Step 2: Obtain the kinship benchmark system based on the analysis unit and measurement node data; define the calculation rules for kinship feature dimensions and configure the weights;

[0009] Step 3: Calculate the blood relation strength based on the blood relation feature dimension calculation rules; determine the initial mapping scheme based on the blood relation strength; define constraint rules to optimize the initial mapping scheme and obtain the optimal mapping scheme;

[0010] Step 4: Solidify the optimal mapping scheme into the final mapping relationship, generate a bloodline tracing file and an optimization process report, and provide feedback.

[0011] In step 1, the metering node is represented as M, and each node is bound to a unique identifier ID; the metering node's full-cycle runtime sequence data D(m) is obtained through the acquisition terminal, including continuous acquisition data of parameters such as voltage, current, power, flow rate, and load;

[0012] The analysis unit is represented by Z, and each unit is bound to a unique identifier ID; the analysis unit attribute data, including physical boundaries, area, and power distribution topology, is obtained through system archives;

[0013] Deployment event data is collected via mobile scanning and form submission, recording node deployment / removal information, including node ID, timestamp, metadata, and network access information;

[0014] Define the lineage feature dimension F, including metadata matching dimension f1 (matching key fields such as device affiliation area and measurement parameter type to initially judge the rationality of association), network topology dimension f2 (reflecting communication correlation based on network connection stability), time sequence pattern dimension f3 (verifying data consistency by matching operational patterns), power distribution topology dimension f4 (strengthening association logic based on physical topology relationship), and business time sequence dimension f5 (verifying adaptability by combining business scenario time sequence). The intensity value of each dimension is normalized to [0,1].

[0015] Collect the original supporting data corresponding to each lineage characteristic dimension, including regional registration information (including the list of equipment access for analysis units and information of responsible persons), gateway topology diagram (showing the coverage area of ​​the gateway and connected equipment), equipment operation curves (such as a 24-hour power change trend chart), power distribution topology relationship (clearly defining the affiliation of switches, lines and analysis units), and business schedule (recording the time of events such as inspection, scheduling, and maintenance).

[0016] In step 2, the process of obtaining the kinship benchmark system based on analysis unit and measurement node data (providing a reference standard at the analysis unit level for calculating association strength, avoiding the influence of randomness in single node data) specifically includes: selecting kinship mapping node data that has been manually verified and confirmed in historical periods as a sample set; and extracting features from the sample data at the analysis unit granularity.

[0017] Statistical characteristics: Calculate the periodic statistical indicators of the measurement node data under the same analysis unit, including the mean μ. z σ z ² and the percentage of peak interval ρ z ;

[0018] Pattern-based features: Analyze the changing patterns of time-series data at metering nodes and extract consistent time-series patterns P across periods. z (t);

[0019] The bloodline benchmark system B is integrated to form analysis unit Z. z ={μ z ,σ z ²,ρ z ,P z (t)}, and iterates and updates periodically;

[0020] The specific rules for calculating the bloodline feature dimension include:

[0021] Metadata matching dimension: f1(m,z)=N match (m,z) / N total (m,z); where N match (m,z) represents the number of key fields that match between the metadata of metering node m and the metadata of analysis unit z, where N is the number of fields that match between them. total (m,z) represents the preset total number of metadata key fields, N total (m,z)>0; For example: the preset key fields can be set to 5, including the region, equipment type, measurement parameters, installation location, and management department;

[0022] Network topology dimension: f2(m,z)=α×T stable (m,z) / T total (m,z); where α is the gateway home consistency weight, 0≤α≤1, T stable (m,z) represents the stable connection duration between the gateway of metering node m and the gateway of analysis unit z, T. total (m,z) represents the total duration within the statistical period, T total (m,z)>0;

[0023] Temporal pattern dimension: f3(m,z)=sim(D(m),P z(t)); where sim() is the temporal similarity calculation function, which can be implemented using the Dynamic Time Warping (DTW) algorithm;

[0024] Distribution topology dimension: f4(m,z)=β(m,z); where β(m,z) is the preset matching weight of the distribution topology level between metering node m and analysis unit z, for example: β=0.9 for the same branch, β=0.6 for adjacent branches, and β=0.1 for non-associated branches;

[0025] Business time sequence dimension: f5(m,z)=N aligned (m,z) / N events (z); where N aligned (m,z) represents the number of events in the runtime sequence of metering node m and the business events of analysis unit z that satisfy the preset alignment standard, N. events (z) represents the total number of business events in analysis unit z, N events (z)>0;

[0026] Assign weights W={w1,w2,w3,w4,w5} to each bloodline characteristic dimension. These weights can be adjusted according to the scenario. For example, if physical affinity is emphasized, increase the weight of w4.

[0027] In step 3, the calculation of bloodline association strength based on the bloodline feature dimension calculation rule (comprehensive multi-dimensional features to quantify the degree of association between nodes and analysis units) specifically includes: for each measurement node m∈M and candidate analysis unit z∈Z, the bloodline association strength is calculated through a weighted fusion algorithm: S(m→z)=σ(∑w i ·f i (m,z));

[0028] Where S(m→z) represents the kinship association strength from the measurement node m to the analysis unit z; σ(·) is the normalization function, which can be either the Sigmoid function or the Min-Max normalization; i∈{1,2,…,5};

[0029] The initial mapping scheme based on the strength of kinship specifically includes: for each measurement node m, arranging all analysis units z in descending order of S(m→z), and selecting the unit z with the highest comprehensive correlation strength. top1 This forms the initial mapping scheme M0;

[0030] Defining global lineage constraint rules (ensuring the mapping scheme is reasonable and feasible, and avoiding logical conflicts) specifically includes:

[0031] Node uniqueness constraint: A single metering node can only be mapped to one analysis unit;

[0032] Element capacity constraint: The total load L(z) of all nodes associated with element z must satisfy |L(z) - L base(z)|≤θ·L base (z); where L base (z) represents the preset baseline load; θ represents the constraint coefficient; node load data is obtained based on D(m);

[0033] Topology consistency constraint: Nodes under the same distribution topology are mapped to analysis units that are the same or adjacent in the same distribution topology level;

[0034] Establish the optimization objective: max∑S(m→z)·x mz -λ·∑C; The first term maximizes the global correlation strength, and the second term minimizes the constraint violation;

[0035] Where, x mz For 0-1 decision variables, x mz =1 indicates that m is mapped to z, x mz =0 indicates no mapping; C is the constraint violation penalty term, C=Σ k=1 3 w k ·C k k=1,2,3: correspond to node uniqueness constraint, cell capacity constraint, and topology consistency constraint, respectively; w k This represents the penalty weight for the k-th type of constraint, configured according to the importance of the constraint, with a total weight sum of 1; C k This represents the penalty value for a single violation of the k-th type of constraint;

[0036] The optimal mapping scheme M* that satisfies all global constraints is solved by using an integer programming algorithm or a heuristic greedy algorithm.

[0037] Meanwhile, the algorithm includes a relaxation mechanism (to avoid having no solution under strict constraints): when there is no solution under strict constraints, the constraint with the lowest penalty coefficient is relaxed first (such as appropriately expanding the fluctuation range of the capacity constraint) in order to seek a feasible and as optimal a compromise solution M*.

[0038] In step 4, the optimal mapping scheme M* obtained from the optimization solution is solidified into the final mapping relationship;

[0039] The generation of the bloodline tracing file specifically includes: when ∑w i ·f i When (m,z)≠0, record the S(m→z) and the contribution percentage of each dimension for each mapping (m→z): contrib(f i )=w i ·f i (m,z) / ∑w i ·f i (m,z) form a bloodline link;

[0040] The generated optimization process report specifically includes: the global total strength of bloodline association and constraint satisfaction of the final solution M*;

[0041] When the relaxation mechanism is enabled, the relaxed constraints and the degree of relaxation should be clearly reported.

[0042] When the optimization algorithm cannot find a feasible solution, it outputs a no-solution state and outputs the core conflict points that led to the no-solution (e.g., which specific constraints cannot be satisfied).

[0043] A hydropower data analysis system based on data lineage includes a data acquisition module, a benchmark construction module, a mapping optimization module, and a result generation module.

[0044] The data acquisition module is used to acquire data on metering nodes, analysis units, and deployment events; define kinship feature dimensions and acquire corresponding original supporting data; the benchmark construction module is used to obtain a kinship benchmark system based on the analysis unit and metering node data; define kinship feature dimension calculation rules and configure weights; the mapping optimization module is used to calculate kinship association strength based on the kinship feature dimension calculation rules; determine the initial mapping scheme based on the kinship association strength; define constraint rules to optimize the initial mapping scheme to obtain the optimal mapping scheme; the result generation module is used to solidify the optimal mapping scheme into the final mapping relationship, generate kinship traceability files and optimization process reports, and provide feedback.

[0045] The data acquisition module includes a metering node acquisition unit, an analysis unit acquisition unit, a deployment event collection unit, a supporting data collection unit, and a feature dimension definition unit.

[0046] The metering node acquisition unit is used to acquire the unique identifier and full-cycle runtime sequence data of the metering node; the analysis unit acquisition unit is used to acquire the unique identifier and attribute data of the analysis unit; the deployment event acquisition unit is used to collect deployment event data; the support data acquisition unit is used to collect the original support data corresponding to each lineage feature dimension; the feature dimension definition unit is used to define the lineage feature dimensions, including metadata matching dimension, network topology dimension, time sequence pattern dimension, power distribution topology dimension, and service time sequence dimension.

[0047] The benchmark construction module includes a sample selection unit, a feature extraction unit, a benchmark update unit, a rule definition unit, and a weight configuration unit;

[0048] The sample selection unit is used to select bloodline mapping node data that has been manually verified and confirmed in the historical period as a sample set; the feature extraction unit is used to extract features from the sample data at the analysis unit granularity, including statistical features and pattern features; the benchmark update unit is used to integrate and form the bloodline benchmark system of the analysis unit; the rule definition unit is used to define the calculation rules for bloodline feature dimensions; and the weight configuration unit is used to configure weights for each bloodline feature dimension.

[0049] The mapping optimization module includes a correlation strength calculation unit, an initial mapping unit, a constraint definition unit, and an optimization solution unit.

[0050] The association strength calculation unit is used to calculate the association strength between the measurement node and the analysis unit based on the association feature dimension calculation rules; the initial mapping unit is used to determine the initial mapping scheme based on the association strength; the constraint definition unit is used to define global association constraint rules, including node uniqueness constraints, unit capacity constraints and topology consistency constraints; the optimization solution unit is used to establish optimization objectives and solve them using optimization algorithms, and output the optimal mapping scheme.

[0051] The result generation module includes a mapping and solidification unit, a traceability archive generation unit, and a process report generation unit;

[0052] The mapping solidification unit is used to solidify the optimal mapping scheme into the final mapping relationship; the traceability archive generation unit is used to generate a bloodline traceability archive, recording the bloodline association strength and contribution ratio of each dimension for each mapping pair; the process report generation unit is used to generate an optimization process report, including the global total bloodline association strength and constraint satisfaction of the final scheme.

[0053] Compared with existing technologies, the beneficial effects of this invention are as follows: This invention, through an automated weighted fusion algorithm and mapping optimization process, can replace manual repetitive configuration of equipment-region association relationships, reducing labor costs in the short-cycle operation and maintenance of temporary building complexes; This invention, through multi-dimensional cross-validation calculation of association strength using metadata, power distribution topology, etc., can reduce single-dimensional misjudgments, improve the mapping accuracy of hydropower data and analysis units, and provide a reliable data foundation for subsequent metering statistics and energy consumption analysis; This invention, through flexible configuration of feature dimension weights and constraint coefficients, can adapt to the different scenarios of various types of temporary building complexes such as exhibition venues and construction sites, and has broad scenario adaptability. Attached Figure Description

[0054] Figure 1 This is a flowchart illustrating the hydropower data analysis system based on data lineage of the present invention. Detailed Implementation

[0055] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0056] Example: Figure 1 As shown, this invention provides a technical solution: a hydropower data analysis method based on data lineage, which includes the following steps:

[0057] Step 1: Obtain metering node, analysis unit, and deployment event data; define lineage characteristic dimensions and obtain corresponding raw supporting data;

[0058] The metering node refers to the hydropower metering equipment deployed in a single cycle, such as voltage transformers, current transformers, and power meters; the analysis unit refers to the analysis object in a fixed area, such as by dividing it into defined zones.

[0059] Step 2: Obtain the kinship benchmark system based on the analysis unit and measurement node data; define the calculation rules for kinship feature dimensions and configure the weights;

[0060] Step 3: Calculate the blood relation strength based on the blood relation feature dimension calculation rules; determine the initial mapping scheme based on the blood relation strength; define constraint rules to optimize the initial mapping scheme and obtain the optimal mapping scheme;

[0061] Step 4: Solidify the optimal mapping scheme into the final mapping relationship, generate a bloodline tracing file and an optimization process report, and provide feedback.

[0062] In step 1, the metering node is represented as M, and each node is bound to a unique identifier ID; the metering node's full-cycle runtime sequence data D(m) is obtained through the acquisition terminal, including continuous acquisition data of parameters such as voltage, current, power, flow rate, and load;

[0063] The analysis unit is represented by Z, and each unit is bound to a unique identifier ID; the analysis unit attribute data, including physical boundaries, area, and power distribution topology, is obtained through system archives;

[0064] Deployment event data is collected via mobile scanning and form submission, recording node deployment / removal information, including node ID, timestamp, metadata, and network access information;

[0065] Define the lineage feature dimension F, including metadata matching dimension f1 (matching key fields such as device affiliation area and measurement parameter type to initially judge the rationality of association), network topology dimension f2 (reflecting communication correlation based on network connection stability), time sequence pattern dimension f3 (verifying data consistency by matching operational patterns), power distribution topology dimension f4 (strengthening association logic based on physical topology relationship), and business time sequence dimension f5 (verifying adaptability by combining business scenario time sequence). The intensity value of each dimension is normalized to [0,1].

[0066] Collect the original supporting data corresponding to each lineage characteristic dimension, including regional registration information (including the list of equipment access for analysis units and information of responsible persons), gateway topology diagram (showing the coverage area of ​​the gateway and connected equipment), equipment operation curves (such as a 24-hour power change trend chart), power distribution topology relationship (clearly defining the affiliation of switches, lines and analysis units), and business schedule (recording the time of events such as inspection, scheduling, and maintenance).

[0067] In step 2, the process of obtaining the kinship benchmark system based on analysis unit and measurement node data (providing a reference standard at the analysis unit level for calculating association strength, avoiding the influence of randomness in single node data) specifically includes: selecting kinship mapping node data that has been manually verified and confirmed in historical periods as a sample set; and extracting features from the sample data at the analysis unit granularity.

[0068] Statistical characteristics: Calculate the periodic statistical indicators of the measurement node data under the same analysis unit, including the mean μ. z σ z ² and the percentage of peak interval ρ z ;

[0069] Pattern-based features: Analyze the changing patterns of time-series data at metering nodes and extract consistent time-series patterns P across periods. z (t);

[0070] The bloodline benchmark system B is integrated to form analysis unit Z. z ={μ z ,σ z ²,ρ z ,P z (t)}, and iterates and updates periodically;

[0071] The specific rules for calculating the bloodline feature dimension include:

[0072] Metadata matching dimension: f1(m,z)=N match (m,z) / N total (m,z); where N match (m,z) represents the number of key fields that match between the metadata of metering node m and the metadata of analysis unit z, where N is the number of fields that match between them. total(m,z) represents the preset total number of metadata key fields, N total (m,z)>0; For example: the preset key fields can be set to 5, including the region, equipment type, measurement parameters, installation location, and management department;

[0073] Network topology dimension: f2(m,z)=α×T stable (m,z) / T total (m,z); where α is the gateway home consistency weight, 0≤α≤1, T stable (m,z) represents the stable connection duration between the gateway of metering node m and the gateway of analysis unit z, T. total (m,z) represents the total duration within the statistical period, T total (m,z)>0;

[0074] Temporal pattern dimension: f3(m,z)=sim(D(m),P z (t)); where sim() is the temporal similarity calculation function, which can be implemented using the Dynamic Time Warping (DTW) algorithm;

[0075] Distribution topology dimension: f4(m,z)=β(m,z); where β(m,z) is the preset matching weight of the distribution topology level between metering node m and analysis unit z, for example: β=0.9 for the same branch, β=0.6 for adjacent branches, and β=0.1 for non-associated branches;

[0076] Business time sequence dimension: f5(m,z)=N aligned (m,z) / N events (z); where N aligned (m,z) represents the number of events in the runtime sequence of metering node m and the business events of analysis unit z that satisfy the preset alignment standard, N. events (z) represents the total number of business events in analysis unit z, N events (z)>0;

[0077] Assign weights W={w1,w2,w3,w4,w5} to each bloodline characteristic dimension. These weights can be adjusted according to the scenario. For example, if physical affinity is emphasized, increase the weight of w4.

[0078] In step 3, the calculation of bloodline association strength based on the bloodline feature dimension calculation rule (comprehensive multi-dimensional features to quantify the degree of association between nodes and analysis units) specifically includes: for each measurement node m∈M and candidate analysis unit z∈Z, the bloodline association strength is calculated through a weighted fusion algorithm: S(m→z)=σ(∑w i ·f i (m,z));

[0079] Where S(m→z) represents the kinship association strength from the measurement node m to the analysis unit z; σ(·) is the normalization function, which can be either the Sigmoid function or the Min-Max normalization; i∈{1,2,…,5};

[0080] The initial mapping scheme based on the strength of kinship specifically includes: for each measurement node m, arranging all analysis units z in descending order of S(m→z), and selecting the unit z with the highest comprehensive correlation strength. top1 This forms the initial mapping scheme M0;

[0081] Defining global lineage constraint rules (ensuring the mapping scheme is reasonable and feasible, and avoiding logical conflicts) specifically includes:

[0082] Node uniqueness constraint: A single metering node can only be mapped to one analysis unit;

[0083] Element capacity constraint: The total load L(z) of all nodes associated with element z must satisfy |L(z) - L base (z)|≤θ·L base (z); where L base (z) represents the preset baseline load; θ represents the constraint coefficient; node load data is obtained based on D(m);

[0084] Topology consistency constraint: Nodes under the same distribution topology are mapped to analysis units that are the same or adjacent in the same distribution topology level;

[0085] Establish the optimization objective: max∑S(m→z)·x mz -λ·∑C; The first term maximizes the global correlation strength, and the second term minimizes the constraint violation;

[0086] Where, x mz For 0-1 decision variables, x mz =1 indicates that m is mapped to z, x mz =0 indicates no mapping; C is the constraint violation penalty term;

[0087] The optimal mapping scheme M* that satisfies all global constraints is solved by using an integer programming algorithm or a heuristic greedy algorithm.

[0088] Meanwhile, the algorithm includes a relaxation mechanism (to avoid having no solution under strict constraints): when there is no solution under strict constraints, the constraint with the lowest penalty coefficient is relaxed first (such as appropriately expanding the fluctuation range of the capacity constraint) in order to seek a feasible and as optimal a compromise solution M*.

[0089] In step 4, the optimal mapping scheme M* obtained from the optimization solution is solidified into the final mapping relationship;

[0090] The generation of the bloodline tracing file specifically includes: when ∑wi ·f i When (m,z)≠0, record the S(m→z) and the contribution percentage of each dimension for each mapping (m→z): contrib(f i )=w i ·f i (m,z) / ∑w i ·f i (m,z) form a bloodline link;

[0091] The generated optimization process report specifically includes: the global total strength of bloodline association and constraint satisfaction of the final solution M*;

[0092] When the relaxation mechanism is enabled, the relaxed constraints and the degree of relaxation should be clearly reported.

[0093] When the optimization algorithm cannot find a feasible solution, it outputs a no-solution state and outputs the core conflict points that led to the no-solution (e.g., which specific constraints cannot be satisfied).

[0094] A hydropower data analysis system based on data lineage includes a data acquisition module, a benchmark construction module, a mapping optimization module, and a result generation module.

[0095] The data acquisition module is used to acquire data on metering nodes, analysis units, and deployment events; define kinship feature dimensions and acquire corresponding original supporting data; the benchmark construction module is used to obtain a kinship benchmark system based on the analysis unit and metering node data; define kinship feature dimension calculation rules and configure weights; the mapping optimization module is used to calculate kinship association strength based on the kinship feature dimension calculation rules; determine the initial mapping scheme based on the kinship association strength; define constraint rules to optimize the initial mapping scheme to obtain the optimal mapping scheme; the result generation module is used to solidify the optimal mapping scheme into the final mapping relationship, generate kinship traceability files and optimization process reports, and provide feedback.

[0096] The data acquisition module includes a metering node acquisition unit, an analysis unit acquisition unit, a deployment event collection unit, a supporting data collection unit, and a feature dimension definition unit.

[0097] The metering node acquisition unit is used to acquire the unique identifier and full-cycle runtime sequence data of the metering node; the analysis unit acquisition unit is used to acquire the unique identifier and attribute data of the analysis unit; the deployment event acquisition unit is used to collect deployment event data; the support data acquisition unit is used to collect the original support data corresponding to each lineage feature dimension; the feature dimension definition unit is used to define the lineage feature dimensions, including metadata matching dimension, network topology dimension, time sequence pattern dimension, power distribution topology dimension, and service time sequence dimension.

[0098] The benchmark construction module includes a sample selection unit, a feature extraction unit, a benchmark update unit, a rule definition unit, and a weight configuration unit;

[0099] The sample selection unit is used to select bloodline mapping node data that has been manually verified and confirmed in the historical period as a sample set; the feature extraction unit is used to extract features from the sample data at the analysis unit granularity, including statistical features and pattern features; the benchmark update unit is used to integrate and form the bloodline benchmark system of the analysis unit; the rule definition unit is used to define the calculation rules for bloodline feature dimensions; and the weight configuration unit is used to configure weights for each bloodline feature dimension.

[0100] The mapping optimization module includes a correlation strength calculation unit, an initial mapping unit, a constraint definition unit, and an optimization solution unit.

[0101] The association strength calculation unit is used to calculate the association strength between the measurement node and the analysis unit based on the association feature dimension calculation rules; the initial mapping unit is used to determine the initial mapping scheme based on the association strength; the constraint definition unit is used to define global association constraint rules, including node uniqueness constraints, unit capacity constraints and topology consistency constraints; the optimization solution unit is used to establish optimization objectives and solve them using optimization algorithms, and output the optimal mapping scheme.

[0102] The result generation module includes a mapping and solidification unit, a traceability archive generation unit, and a process report generation unit;

[0103] The mapping solidification unit is used to solidify the optimal mapping scheme into the final mapping relationship; the traceability archive generation unit is used to generate a bloodline traceability archive, recording the bloodline association strength and contribution ratio of each dimension for each mapping pair; the process report generation unit is used to generate an optimization process report, including the global total bloodline association strength and constraint satisfaction of the final scheme.

[0104] In this embodiment, the application scenario is a large temporary exhibition venue. The venue has a 15-day exhibition period and is divided into three core exhibition areas (A, B, and C) and supporting service areas. A total of 42 temporary water and electricity metering devices (including power meters, flow meters, voltage transformers, etc.) are deployed. The devices need to be flexibly moved according to the exhibition layout, and there are 3-5 new or removed devices every day.

[0105] Step 1: Multi-source data acquisition and feature dimension definition;

[0106] During the exhibition preparation phase, staff members connected to all metering nodes through IoT data collection terminals and assigned a unique ID to each device ("P05" represents power meter No. 5, "F38" represents flow meter No. 38, and "V12" represents voltage transformer No. 12). Real-time collection of full-cycle time-series data such as voltage, power, and flow was conducted. For example, the operating data of power meter No. 3 (P03) in exhibition area A included power records every 5 minutes during the exhibition period.

[0107] The venue is divided into 4 analysis units according to physical boundaries (Z1=Exhibition Area A, Z2=Exhibition Area B, Z3=Exhibition Area C, Z4=Service Area). The attribute data of each unit is obtained through system archives. The physical boundary of Z1 is the area of ​​booths 1-10 on the east side of the venue, and the power distribution topology belongs to the No. 1 main power supply path.

[0108] During equipment deployment / removal, staff can quickly collect deployment event data by scanning a QR code on their mobile devices, eliminating the need for manual recording of the unit. For example, on the 5th day of the exhibition, a flow meter (F38) was added to Zone B. The event data recorded after scanning the QR code is: "Node ID: F38, Timestamp: 06-12 14:20:00, Metadata: Cold water flow meter (measurement range 0-50m³ / h), Network access information: Gateway IP: XXX.XXX.2.56".

[0109] Simultaneously, five lineage characteristic dimensions (metadata matching, network topology, time sequence pattern, power distribution topology, and service time sequence) are defined, and corresponding original supporting data are collected synchronously: regional registration information (including a list of water and electricity equipment types allowed to be accessed in the Z1 exhibition area), gateway topology map (clearly indicating that the network segment only covers units Z2 and Z4), equipment operation curves (such as the curve characteristics of "high traffic during the exhibition period" for existing equipment in the Z2 exhibition area), power distribution topology relationship (the No. 1 main power supply path covers Z1, Z2, and Z4), and service schedule (recording events such as maintenance in the Z3 exhibition area on June 10 and daily operations from 9:00 to 18:00), providing multi-dimensional basis for determining the unit to which F38 belongs.

[0110] Step 2: Construction and rule configuration of the kinship benchmark system;

[0111] Twenty-eight sets of "metering node-analysis unit" mapping data, confirmed by maintenance personnel three days prior to the exhibition, were selected as the sample set (e.g., P05→Z1, F12→Z4). These samples were used only to establish a baseline; subsequent additions of nodes did not require manual confirmation. Features were extracted at the analysis unit granularity. For exhibition area Z2 (B exhibition area), the periodic mean μ of the power / flow data of the metering nodes in the samples was calculated. z2 =8.2kW, variance σ z2 ² = 0.28kW², the percentage of peak range (10-12kW) ρ z2=20%, and simultaneously extracted the cross-cycle time series pattern P of "high load / flow during the opening period (9:00-18:00) and low after closing". z2 (t), integrating to form the bloodline benchmark system B of Z2 z2 It is set to automatically iterate and update every day at midnight to adapt to changes in regional data characteristics in short-cycle scenarios.

[0112] The calculation rules for each dimension are defined as follows: The metadata matching dimension presets five key fields: region of origin, device type, measurement parameters, installation location, and management department; the gateway affiliation consistency weight α for the network topology dimension is set to 0.8; the time-series similarity function uses the DTW algorithm (used to compare the fit between node operating curves and regional baseline patterns); the distribution topology weight β is set to 0.9 for the same branch, 0.6 for adjacent branches, and 0.1 for unrelated branches; the business time-series dimension uses "whether the fluctuation of device operating data during the event period matches the regional business" as the alignment standard. Considering the frequent equipment migrations and network volatility at exhibitions, weights W={w1=0.2, w2=0.25, w3=0.25, w4=0.15, w5=0.15} are configured, focusing on strengthening the weights of network topology (reflecting the correlation of nodes accessing the regional network) and time-series patterns (reflecting the adaptability of node operating patterns to the regional scenario).

[0113] Step 3: Calculate the correlation strength and optimize the mapping scheme;

[0114] For each metering node and all candidate analysis units, a weighted fusion algorithm is used to calculate the lineage association strength (the higher the association strength, the greater the probability that the node belongs to that unit). For example, the newly added flow meter F38 has metadata (cold water flow meter, adapted to the water demand of Zone B) that matches four key fields with Z2 (Zone B) (f1=0.8); the gateway IP belongs to the Z2 coverage segment, and the stable connection time accounts for 92% of the statistical period (f2=0.8×0.92=0.736); the runtime sequence (the flow rate has been consistently high since access at 14:20, which matches the characteristics of the time period) matches the P of Z2. z2 (t) Similarity 0.88 (f3=0.88); the distribution topology is the adjacent branch of the main power supply path 1 where Z2 is located (f4=0.6); it aligns with Z2 in 2 out of the 3 business events of the day (9:00 start, 12:00 peak water usage, 18:00 closing) (f5=0.667). After weighted calculation, the correlation strength between F38 and Z2, S(F38→Z2)=0.79, is much higher than the correlation strength with Z1 (0.52), Z3 (0.41), and Z4 (0.63). Therefore, in the initial mapping scheme M0, it is preliminarily determined that F38 belongs to Z2.

[0115] The rationality was further verified by applying global constraint rules: ensuring that F38 belongs to only one cell (node ​​uniqueness constraint); setting the baseline load L for the Z2 expansion area. base =50kW, constraint coefficient θ=0.15, meaning the total load of all nodes belonging to Z2 must be between 42.5-57.5kW (unit capacity constraint); nodes on the same main power supply path must belong to Z1, Z2, or Z4 (topology consistency constraint). An optimization objective of "maximizing global association strength + minimizing constraint violations" is established, and a heuristic greedy algorithm is used to solve it. During the process, it is found that the total load of the currently associated nodes of Z2 is 56.8kW, which does not exceed the constraint upper limit, so there is no need to relax the rules. Finally, the optimal mapping scheme M* is output, clearly determining that F38 belongs to unit Z2.

[0116] Step 4: Mapping and Report Generation;

[0117] The optimal mapping scheme M* is solidified into the final mapping relationship, that is, "F38 belongs to Z2 (B exhibition area)", thus completing the automatic determination of the unit to which the metering node belongs. A lineage traceability file is generated for this mapping, recording "association strength S=0.79, contribution percentage of each dimension: f1=20.3%, f2=25.5%, f3=27.8%, f4=11.4%, f5=14.9%", clearly presenting the judgment basis and facilitating subsequent traceability.

[0118] The generated optimization process report shows that the overall global correlation strength of this solution is 86.3, all constraints are satisfied, and there are no relaxed items. On the 8th day of the exhibition, power meter P18 in booth Z2 was moved to booth Z3 due to booth layout adjustments. Staff only needed to scan the code to update the deployment event data (without manually modifying the unit it belongs to), and the system automatically re-executed the above steps: recalculating the correlation strength between P18 and each unit, and finding that its correlation strength with Z3 reached 0.81 (higher than Z2's 0.45). After constraint optimization, the mapping relationship was automatically updated to "P18 belongs to Z3", and the lineage tracing file was updated simultaneously.

[0119] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from its spirit or essential characteristics. Therefore, the embodiments should be considered in all respects as exemplary and non-limiting, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be included within the present invention. No reference numerals in the claims should be construed as limiting the scope of the claims.

Claims

1. A hydropower data analysis method based on data lineage, characterized in that: The method includes the following steps: Step 1: Obtain metering node, analysis unit, and deployment event data; define lineage characteristic dimensions and obtain corresponding raw supporting data; The metering node represents a hydropower metering device deployed in a single cycle; the analysis unit represents an analysis object with a fixed division area. Step 2: Obtain the kinship benchmark system based on the analysis unit and measurement node data; define the calculation rules for kinship feature dimensions and configure the weights; Step 3: Calculate the blood relation strength based on the blood relation feature dimension calculation rules; determine the initial mapping scheme based on the blood relation strength; define constraint rules to optimize the initial mapping scheme and obtain the optimal mapping scheme; Step 4: Solidify the optimal mapping scheme into the final mapping relationship, generate a bloodline tracing file and an optimization process report, and provide feedback; In step 2, the specific rules for defining the bloodline characteristic dimension calculation include: Metadata matching dimension: f1(m,z)=N match (m,z) / N total (m,z); where N match (m,z) represents the number of key fields that match between the metadata of metering node m and the metadata of analysis unit z, where N is the number of fields that match between them. total (m,z) represents the preset total number of metadata key fields, N total (m,z)>0; Network topology dimension: f2(m,z)=α×T stable (m,z) / T total (m,z); where α is the gateway home consistency weight, 0≤α≤1, T stable (m,z) represents the stable connection duration between the gateway of metering node m and the gateway of analysis unit z, T. total (m,z) represents the total duration within the statistical period, T total (m,z)>0; Temporal pattern dimension: f3(m,z)=sim(D(m),P z (t)); where sim() is the temporal similarity calculation function; Distribution topology dimension: f4(m,z)=β(m,z); where β(m,z) is the preset matching weight of the distribution topology level between metering node m and analysis unit z; Business time sequence dimension: f5(m,z)=N aligned (m,z) / N events (z); where N aligned (m,z) represents the number of events in the runtime sequence of metering node m and the business events of analysis unit z that satisfy the preset alignment standard, N. events (z) represents the total number of business events in analysis unit z, N events (z)>0; Assign weights W={w1,w2,w3,w4,w5} to each bloodline characteristic dimension.

2. The hydropower data analysis method based on data lineage as described in claim 1, characterized in that: In step 1, the metering node is represented as M, and each node is bound to a unique identifier ID; the full-cycle runtime sequence data D(m) of the metering node is obtained through the acquisition terminal; The analysis unit is represented by Z, and each unit is bound to a unique identifier ID; the analysis unit attribute data, including physical boundaries, area, and power distribution topology, is obtained through system archives; Deployment event data is collected via mobile scanning and form submission, recording node deployment / removal information, including node ID, timestamp, metadata, and network access information; Define the lineage feature dimension F, including metadata matching dimension f1, network topology dimension f2, time sequence pattern dimension f3, power distribution topology dimension f4, and service time sequence dimension f5. The intensity value of each dimension is normalized to [0,1]. Collect raw supporting data corresponding to each lineage characteristic dimension, including regional registration information, gateway topology diagram, equipment operation curve, power distribution topology relationship, and business schedule.

3. The hydropower data analysis method based on data lineage according to claim 2, characterized in that: In step 2, obtaining the kinship benchmark system based on analysis units and measurement node data specifically includes: selecting kinship mapping node data that has been manually verified and confirmed in historical periods as a sample set; and extracting features from the sample data at the analysis unit granularity. Statistical characteristics: Calculate the periodic statistical indicators of the measurement node data under the same analysis unit, including the mean μ. z σ z ² and the percentage of peak interval ρ z ; Pattern-based features: Analyze the changing patterns of time-series data at metering nodes and extract consistent time-series patterns P across periods. z (t); The bloodline benchmark system B is integrated to form analysis unit Z. z ={μ z ,σ z ²,ρ z ,P z (t)}, and iterates and updates periodically.

4. The hydropower data analysis method based on data lineage according to claim 3, characterized in that: In step 3, the calculation of kinship association strength based on the kinship feature dimension calculation rule specifically includes: for each measurement node m∈M and candidate analysis unit z∈Z, calculating the kinship association strength through a weighted fusion algorithm: S(m→z)=σ(∑w i ·f i (m,z)); Where S(m→z) represents the kinship association strength from the measurement node m to the analysis unit z; σ(·) is the normalization function, which can be either the Sigmoid function or the Min-Max normalization; i∈{1,2,…,5}; The initial mapping scheme based on the strength of kinship specifically includes: for each measurement node m, arranging all analysis units z in descending order of S(m→z), and selecting the unit z with the highest comprehensive correlation strength. top1 This forms the initial mapping scheme M0; Defining global lineage constraint rules specifically includes: Node uniqueness constraint: A single metering node can only be mapped to one analysis unit; Element capacity constraint: The total load L(z) of all nodes associated with element z must satisfy |L(z) - L base (z)|≤θ·L base (z); where L base (z) represents the preset baseline load; θ represents the constraint coefficient; node load data is obtained based on D(m); Topology consistency constraint: Nodes under the same distribution topology are mapped to analysis units that are the same or adjacent in the same distribution topology level; Establish the optimization objective: max∑S(m→z)·x mz -λ·∑C; where, x mz For 0-1 decision variables, x mz =1 indicates that m is mapped to z, x mz =0 indicates no mapping; C is the constraint violation penalty term; The algorithm employs either integer programming or a heuristic greedy algorithm to optimize the solution and outputs the optimal mapping scheme M* that satisfies all global constraints.

5. The hydropower data analysis method based on data lineage according to claim 4, characterized in that: In step 4, the optimal mapping scheme M* obtained from the optimization solution is solidified into the final mapping relationship; The generation of the bloodline tracing file specifically includes: when ∑w i ·f i When (m,z)≠0, record the S(m→z) and the contribution percentage of each dimension for each mapping (m→z): contrib(f i )=w i ·f i (m,z) / ∑w i ·f i (m,z) form a bloodline link; The generated optimization process report specifically includes: the global total strength of bloodline association and constraint satisfaction of the final solution M*.

6. A hydropower data analysis system based on data lineage, used to execute the hydropower data analysis method based on data lineage as described in any one of claims 1-5, characterized in that: The system includes a data acquisition module, a benchmark construction module, a mapping optimization module, and a result generation module; The data acquisition module is used to acquire data from metering nodes, analysis units, and deployment events; define lineage feature dimensions and acquire corresponding raw supporting data; The benchmark construction module is used to obtain a bloodline benchmark system based on analysis units and measurement node data; define the calculation rules for bloodline feature dimensions and configure weights; The mapping optimization module is used to calculate the strength of blood relation based on the blood relation feature dimension calculation rules; The initial mapping scheme is determined based on the strength of blood relations; Define constraint rules to optimize the initial mapping scheme and obtain the optimal mapping scheme; The result generation module is used to solidify the optimal mapping scheme into the final mapping relationship, generate a bloodline tracing file and an optimization process report, and provide feedback.

7. The hydropower data analysis system based on data lineage according to claim 6, characterized in that: The data acquisition module includes a metering node acquisition unit, an analysis unit acquisition unit, a deployment event collection unit, a supporting data collection unit, and a feature dimension definition unit. The metering node acquisition unit is used to acquire the unique identifier and full-cycle runtime sequence data of the metering node; the analysis unit acquisition unit is used to acquire the unique identifier and attribute data of the analysis unit; the deployment event collection unit is used to collect deployment event data. The supporting data acquisition unit is used to collect the original supporting data corresponding to each lineage feature dimension; the feature dimension definition unit is used to define the lineage feature dimensions, including metadata matching dimension, network topology dimension, time sequence pattern dimension, power distribution topology dimension and service time sequence dimension.

8. The hydropower data analysis system based on data lineage according to claim 7, characterized in that: The benchmark construction module includes a sample selection unit, a feature extraction unit, a benchmark update unit, a rule definition unit, and a weight configuration unit; The sample selection unit is used to select bloodline mapping node data that has been manually verified and confirmed in the historical period as a sample set; the feature extraction unit is used to extract features from the sample data at the analysis unit granularity, including statistical features and pattern features; the benchmark update unit is used to integrate and form the bloodline benchmark system of the analysis unit; the rule definition unit is used to define the calculation rules for bloodline feature dimensions; and the weight configuration unit is used to configure weights for each bloodline feature dimension.

9. The hydropower data analysis system based on data lineage according to claim 8, characterized in that: The mapping optimization module includes a correlation strength calculation unit, an initial mapping unit, a constraint definition unit, and an optimization solution unit. The association strength calculation unit is used to calculate the blood relationship association strength between the measurement node and the analysis unit based on the blood relationship feature dimension calculation rules; the initial mapping unit is used to determine the initial mapping scheme based on the blood relationship association strength. The constraint definition unit is used to define global lineage constraint rules, including node uniqueness constraints, unit capacity constraints, and topology consistency constraints; the optimization solution unit is used to establish optimization objectives and solve them using optimization algorithms, and output the optimal mapping scheme.

10. The hydropower data analysis system based on data lineage according to claim 9, characterized in that: The result generation module includes a mapping and solidification unit, a traceability archive generation unit, and a process report generation unit; The mapping solidification unit is used to solidify the optimal mapping scheme into the final mapping relationship; the traceability archive generation unit is used to generate a bloodline traceability archive, recording the bloodline association strength and contribution ratio of each dimension for each mapping pair; the process report generation unit is used to generate an optimization process report, including the global total bloodline association strength and constraint satisfaction of the final scheme.

Citation Information

Patent Citations

  • Multi-source field-level blood relationship tracking method and system

    CN121542323A

  • Intelligent data blood relationship tracking and visualization method based on graph calculation

    CN121579585A