Method and device for tracing and repairing water conservancy data quality abnormity and medium

By constructing a dynamic weighted Bayesian network and using Bayesian formulas to calculate the root causes of anomalies, the problem of insufficient precision and traceability in data quality control in the water conservancy trusted data space was solved. This enabled accurate tracing and repair of water conservancy data quality, improved the precision and reliability of data quality control, and adapted to the needs of water conservancy business.

CN122086882APending Publication Date: 2026-05-26SHANDONG FENGSHI INFORMATION TECH CO LTD
View PDF 9 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANDONG FENGSHI INFORMATION TECH CO LTD
Filing Date
2026-04-23
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing technologies in the water conservancy trusted data space suffer from problems such as fixed weights, poor traceability, lack of closed-loop mechanisms, and insufficient industry adaptability. This results in insufficient precision in data quality control and a lack of traceability, failing to meet the reliability requirements of water conservancy operations and the needs of data asset management.

Method used

A dynamic weighted Bayesian network integrating water conservancy business characteristics is constructed. Dynamic weights are calculated using the entropy weight method and the analytic hierarchy process. Combined with the priority of water conservancy business, the network enables accurate source tracing, hierarchical repair, and closed-loop verification of data quality anomalies. A three-layer Bayesian network topology and Bayesian formula are used to calculate the posterior probability of the root cause of the anomaly. Repair and verification are then carried out using a water conservancy industry-specific repair rule library.

Benefits of technology

It has enabled precise tracing of water conservancy data quality anomalies, improving the tracing accuracy rate to over 92% and the data quality qualification rate from 75% to over 95%, forming a closed-loop governance system that meets the data reliability and asset management needs of the water conservancy trusted data space and reduces operation and maintenance costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122086882A_ABST
    Figure CN122086882A_ABST
Patent Text Reader

Abstract

The invention relates to a method and equipment for tracing and repairing water conservancy data quality abnormity, and a medium, and belongs to the technical field of intelligent water conservancy. Collecting multi-source water conservancy data, extracting data features and generating a standardized feature vector set; constructing a three-layer Bayesian network topological structure of circulation link-business influence factor-anomaly type, training to obtain a dynamic weight Bayesian network model adaptive to the water conservancy scene, inputting the standardized feature vector into the dynamic weight Bayesian network model obtained by training, outputting quality anomaly root cause probability distribution, and obtaining the quality anomaly root cause probability distribution. The posterior probability of each abnormal root cause is calculated through a Bayesian formula, and a root cause output traceability result is judged; performing graded repair and verification according to different traceability results; and the repaired qualified data is synchronized to a data asset module of the water conservancy credible data space for service output and generation of a visual data quality report. According to the method, the traceability precision is remarkably improved, and the service adaptability is higher.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a method, equipment, and medium for tracing and repairing water conservancy data quality anomalies, and particularly to data quality control technology in the trusted water conservancy data space. Specifically, it is applied to the tracing, hierarchical repair, and closed-loop verification of quality anomalies in multi-source heterogeneous data within the trusted water conservancy data space, providing reliable data support for the "four predictions" (forecasting, early warning, rehearsal, and contingency plan) of water conservancy operations, and belongs to the field of smart water conservancy technology. Background Technology

[0002] With the deepening of smart water conservancy construction, the trusted water conservancy data space, as the core carrier for data aggregation, governance, and services, needs to integrate heterogeneous data from multiple sources, including hydrological monitoring, water conservancy projects, meteorological sharing, and IoT sensing. This data permeates the entire process of "collection-transmission-conversion-storage," and its quality directly determines the reliability of decisions in core business operations such as flood forecasting and water resource allocation. While Bayesian networks, a classic technique for probabilistic reasoning, have been applied to general data quality control scenarios, they exhibit significant limitations in their adaptability to the trusted water conservancy data space.

[0003] Existing related technologies mainly employ fixed-weight Bayesian networks for data quality anomaly identification. This method achieves preliminary classification of data quality problems by constructing a simple "data feature-anomaly type" mapping relationship. However, this approach has the following drawbacks or problems: ① Fixed weights: Traditional Bayesian networks use static weight configuration, which cannot adapt to the dynamic priorities of water conservancy business (such as water level data having higher priority than daily statistical data during flood control), resulting in insufficient accuracy of data quality control in different business scenarios; ② Poor targeting of source tracing: It can only identify surface anomalies such as missing data and outliers, but cannot accurately locate the root cause of the anomaly (such as data acquisition equipment failure, transmission interference, format conversion error, etc.), which is not conducive to subsequent targeted repair; ③ Lack of closed-loop mechanism: There is no dedicated verification process after abnormal data is repaired, which may lead to incomplete repair and make it difficult to meet the strict requirements of the water conservancy trusted data space for data reliability. ④ Insufficient industry adaptability: The Bayesian model for general scenarios does not take into account the characteristics of water conservancy data, such as "multiple circulation links, strong business correlation, and large differences in timeliness", and does not consider the multi-source data integration characteristics of the water conservancy trusted data space, resulting in poor data governance effect; ⑤ Lack of traceability: The entire "anomaly-source-repair" record has not been established, which fails to meet the traceability requirements for the management of water conservancy trusted data spatial data assets. Summary of the Invention

[0004] The purpose of this invention is to overcome the above-mentioned shortcomings and provide a method for tracing and repairing water conservancy data quality anomalies. By constructing a dynamic weighted Bayesian network that integrates water conservancy business characteristics, it is possible to achieve accurate tracing, hierarchical repair, and closed-loop verification of water conservancy trusted data space data quality anomalies.

[0005] The technical solution adopted in this invention is as follows: The method for tracing and repairing water conservancy data quality anomalies includes the following steps: S1. Collect multi-source water conservancy data, extract the basic characteristics, business characteristics, and circulation characteristics of the data, and generate a standardized feature vector set; S2. Construct a three-layer Bayesian network topology structure of "transfer links - business impact factors - anomaly types", and use the entropy weight method-analytic hierarchy process (EW-AHP) to calculate the dynamic weights of each transfer link under different water conservancy business priorities. Calculate the anomaly types using the dynamic weights. Train a dynamic weight Bayesian network model adapted to the water conservancy scenario using labeled historical sample data. S3. Input the standardized feature vector generated in step S1 into the trained dynamic weight Bayesian network model, output the probability distribution of the root cause of quality anomalies, calculate the posterior probability of each root cause of anomalies using the Bayesian formula, determine the root cause and output the source tracing result. S4. Based on different source tracing results, perform graded repairs and re-input the repaired data into a dynamic weighted Bayesian network to infer and calculate the anomaly probability for verification; S5. Repair qualified data and synchronize it to the data asset module of the water conservancy trusted data space for service output and generation of visual data quality reports.

[0006] In the above method, the water conservancy data mentioned in step S1 includes hydrological monitoring data, water conservancy project operation and maintenance data, water resource scheduling data, and meteorological shared data. The basic features are extracted from the data's inherent attribute features, including data type identifier, collection timestamp, numerical range, and accuracy level; the business features are extracted from features related to water conservancy business, including associated water conservancy objects, business scenario tags, and business priorities; the flow link features are extracted from the entire data flow link, including the collection link, transmission link, conversion link, and storage link.

[0007] The three-layer Bayesian network topology described in step S2 includes a parent node (data flow link), intermediate adjustment nodes (business impact factors), and child nodes (data quality anomaly types). The calculation process for dynamic weights is as follows: Objective weights were calculated using the entropy weight method (EW) based on historical data quality feedback from the water conservancy trusted data space. Subjective weights were calculated using the Analytic Hierarchy Process (AHP) and combined with the experience of water conservancy experts. Generate fusion weights: Combine objective weights and subjective weights in a 6:4 weighted ratio to obtain dynamic weights; The weight update cycle is synchronized with the data collection frequency.

[0008] In step S3, the posterior probability of each abnormal root cause is calculated using Bayes' theorem, as follows: , in, Let R be the posterior probability of the root cause R corresponding to anomaly type A. Let R be the conditional probability that the root cause R leads to the abnormality A. Let R be the prior probability of the root cause, and W be the dynamic weight. Let A be the marginal probability of anomaly type A.

[0009] The determination of the root cause is as follows: a root cause probability threshold is set, and when the probability of a certain root cause reaches the threshold, it is determined to be the main root cause; if there are multiple root causes with similar probabilities, further screening is carried out in combination with business impact factors; for scenarios where multiple anomaly types coexist, the joint root cause probability is calculated through the conditional dependency relationship between nodes, and a comprehensive tracing result is output.

[0010] Another objective of this invention is to provide a system for tracing and repairing anomalies in water conservancy data quality, including... Data access layer: used to acquire multi-source water conservancy data; Feature extraction layer: used to extract basic features, business features, and flow process features of data and generate a standardized feature vector set; Dynamic Weight Bayesian Network Model Layer: Through a three-layer Bayesian network topology structure of "transfer links - business impact factors - anomaly types", the dynamic weights of each transfer link under different water conservancy business priorities are calculated using the entropy weight method and the analytic hierarchy process. The anomaly types are calculated using the dynamic weights. A dynamic weight Bayesian network model adapted to the water conservancy scenario is trained using labeled historical sample data. The standardized feature vectors are input into the trained dynamic weight Bayesian network model, and the probability distribution of the root causes of quality anomalies is output. Source tracing and repair layer: used to calculate the posterior probability of each abnormal root cause using Bayesian formula, determine the root cause and output source tracing results; repair in stages according to different source tracing results, and re-input the repaired data into the dynamic weighted Bayesian network to infer and calculate the abnormal probability for verification; Asset Service Layer: This layer synchronizes the repaired and qualified data to the data asset module of the Water Resources Trusted Data Space for service output and to generate a visual data quality report.

[0011] The aforementioned source tracing and repair layer includes Anomaly tracing unit: used to output the probability distribution of the root cause of quality anomalies through Bayesian inference, determine the root cause, and output the tracing result; Repair Execution Unit: Built-in water conservancy industry-specific repair rule library, used to generate repair plans based on source tracing results and execute graded repairs; Repair and Validation Unit: Used to re-input the repaired data into the dynamic weighted Bayesian network model, verify the repair effect, and output qualified data.

[0012] A storage medium for tracing and repairing water conservancy data quality anomalies is provided. The storage medium stores a program, which, when executed by a processor, implements the steps of the method for tracing and repairing water conservancy data quality anomalies as described above.

[0013] The beneficial effects of this invention are: (1) Significantly improved source tracing accuracy: Through the water conservancy-specific three-layer Bayesian network topology and dynamic weight mechanism, the accuracy of anomaly root cause tracing is improved to over 92%, which is 15%-20% higher than the traditional fixed-weight Bayesian method. It can accurately locate specific root causes such as collection equipment failure and transmission interference. (2) Stronger business adaptability: Dynamic weighting synchronizes with the priority changes of water conservancy business, solves the data quality control needs in different scenarios such as flood control and water resource scheduling, and avoids the problem of "incompatibility" of general technologies; (3) Targeted repair: Based on the accurate source tracing results, a dedicated repair strategy is matched, and the data quality pass rate is increased from 75% of the traditional method to more than 95%, reducing invalid repair operations; (4) Form a closed-loop governance system: Construct a “source tracing-repair-verification” closed-loop mechanism. After repair, the data needs to be verified twice by a Bayesian network to ensure that the data quality meets the standards. (5) Meets the requirements of data asset management: It links the entire chain of "anomaly-source tracing-repair" records, supports data quality source tracing chain query, and adapts to the asset management needs of the water conservancy trusted data space; (6) Operation and maintenance efficiency optimization: By using the root cause early warning function (pushing maintenance early warnings for continuous high-probability abnormal links), the operation and maintenance costs of water conservancy equipment are reduced and the level of intelligent data management is improved. Attached Figure Description

[0014] Figure 1 This is a flowchart of the method of the present invention. Detailed Implementation

[0015] The present invention will be further described below with reference to specific embodiments.

[0016] Example 1: A method for tracing and repairing water conservancy data quality anomalies, including the following steps (e.g.) Figure 1 )as follows: S1. Collect multi-source water conservancy data, extract the basic features, business features, and circulation link features of the data, and generate a standardized feature vector set.

[0017] Based on the principle of "one source of data, multiple uses of one source", the system gathers data from the entire water conservancy region, including hydrological monitoring data (water level, flow rate, rainfall), water conservancy project operation and maintenance data (gate opening, equipment operating status), water resource scheduling data (water intake and use data, reservoir storage), and meteorological sharing data (typhoon path, rainfall forecast).

[0018] The basic features are the extracted attributes of the data itself, including data type identifier, collection timestamp, numerical range, precision level, etc. Business characteristics: Extract features related to water conservancy business, such as associated water conservancy objects, business scenario tags (e.g., "flood control emergency" and "water resources assessment"), and business priorities; Characteristics of the data flow process: Extracting the characteristics of the entire data flow process, including the acquisition stage (device number, acquisition frequency), the transmission stage (transmission protocol, latency indicators), the conversion stage (format conversion rules, field mapping relationships), and the storage stage (storage media, backup frequency); Feature standardization: One-Hot encoding is used to process categorical features (such as data type and business label), and Min-Max standardization is used to process numerical features (such as numerical range and precision level) to generate a standardized feature vector set.

[0019] S2. Construct a three-layer Bayesian network topology structure of "transfer links - business impact factors - anomaly types", and use the entropy weight method-analytic hierarchy process (EW-AHP) to calculate the dynamic weights of each transfer link under different water conservancy business priorities. Calculate the anomaly types using the dynamic weights. Train a dynamic weight Bayesian network model adapted to water conservancy scenarios using labeled historical sample data.

[0020] (1) Network topology design: Construct a three-layer Bayesian network topology of "flow link - business impact factor - anomaly type": Parent node: Data flow process (4 core processes: collection, transmission, conversion, and storage); Intermediate adjustment nodes: Business impact factors (set based on the priority of water conservancy business, such as flood control data timeliness weight 0.8, water resources data accuracy weight 0.9, and daily statistical data integrity weight 0.6). Child nodes: Data quality anomaly types (missing, outliers, inconsistencies, and duplicates). (2) Dynamic weight calculation: Objective weight calculation: The entropy weight method (EW) is adopted to calculate the objective weight based on the historical data quality feedback (such as the anomaly rate and repair success rate of each link) in the water conservancy trusted data space. Subjective weight calculation: The Analytic Hierarchy Process (AHP) is used, combined with the experience of water conservancy business experts to determine subjective weights (such as increasing the weight of the data collection process during flood control). Weighting generation: Objective weights and subjective weights are weighted and merged in a 6:4 ratio to obtain dynamic weights; the weight update cycle is synchronized with the data collection frequency (e.g., hydrological data is updated every 15 minutes). Utilizing dynamic weights to calculate anomalies: Dynamic weights are used as conditional probability correction coefficients from nodes in the flow stages of a Bayesian network to nodes of anomaly types, embedded in the probabilistic inference process of a three-layer network. Let the flow stages be... ( (Corresponding to acquisition, transmission, conversion, and storage respectively), the business impact factor is B, and the anomaly type is... ( (corresponding to missing, outlier, inconsistent, and duplicate values ​​respectively), then the exception type The probability calculation formula is: , In the formula: For the link Dynamic weights; For the next stage of business impact factor B Caused abnormality The basic conditional probability; For the link The prior probability.

[0021] Calculate the probability values ​​of the four anomaly types according to the above formula, and take the one with the highest probability as the main anomaly type of the current data; when multiple anomalies coexist, output the anomaly types with a probability higher than 30% together; in core business scenarios such as flood control and water resource scheduling, the dynamic weight automatically increases the impact intensity of the corresponding key links, making the anomaly calculation more in line with business priorities.

[0022] (3) Network training optimization: Sample preparation: Select historical data samples labeled with "transfer link - anomaly type - root cause label" (such as "collection link - outlier - sensor failure" and "transmission link - missing - network interruption") to construct a training set; Parameter optimization: The network conditional probability table is iteratively optimized using the Expectation-Maximization (EM) algorithm, with the iteration terminating when the model's source attribution accuracy is ≥92%. Model output: After training, output a dynamic weighted Bayesian network model adapted to water conservancy scenarios.

[0023] S3. Input the standardized feature vector generated in step S1 into the trained dynamic weight Bayesian network model, output the probability distribution of the root cause of quality anomalies, calculate the posterior probability of each abnormal root cause using the Bayesian formula, determine the root cause, and output the source tracing result.

[0024] The posterior probability of each anomalous root cause is calculated using Bayes' theorem, as follows: , in, Let R be the posterior probability of the root cause R corresponding to anomaly type A. Let R be the conditional probability that the root cause R leads to the abnormality A. Let R be the prior probability of the root cause, and W be the dynamic weight. Let A be the marginal probability of anomaly type A.

[0025] Root cause determination: Set a root cause probability threshold (≥80%). When the probability of a certain root cause reaches the threshold, it is determined to be the main root cause. If there are multiple root causes with similar probabilities, further screening is carried out in combination with business impact factors. For scenarios where multiple anomaly types coexist, the joint root cause probability is calculated through the conditional dependency relationship between nodes, and a comprehensive source tracing result is output.

[0026] S4. Based on different source tracing results, perform hierarchical repairs and re-input the repaired data into a dynamic weighted Bayesian network to infer and calculate the anomaly probability for verification.

[0027] (1) Remediation plan generation: The water conservancy industry-specific remediation rule library is called, and targeted remediation strategies are matched based on the source tracing results. Missing data due to equipment failure: The solution is to "complete the data by averaging the three nearest data collection periods + equipment maintenance early warning". Outlier data caused by transmission interference: adopt the "SSL encrypted retransmission verification + outlier removal" solution; Inconsistent data caused by format conversion errors: adopt the solution of "reprocessing standard conversion rules + field mapping validation"; (2) Tiered repair execution: Based on the data service priority (P1 level: flood control emergency and core dispatch data; P2 level: engineering operation and maintenance data; P3 level: daily statistical data), high priority data is repaired first to ensure that core water conservancy services are not affected; Numerical data: Interpolation is used for repair (cubic spline interpolation is used for hydrological time series data, and nearest neighbor mean interpolation is used for discrete monitoring data). Text-based data: Rule-based matching is used for repair; Spatiotemporal data: repaired using coordinate calibration; (3) Repair verification: Re-input the repaired data into the dynamic weighted Bayesian network and infer the probability of anomalies; if the probability of anomalies is ≤5%, the repair is deemed qualified; if the probability of anomalies is >5%, return to trace the source again and optimize the repair plan.

[0028] S5. Repair qualified data and synchronize it to the data asset module of the water conservancy trusted data space for service output and generation of visual data quality reports.

[0029] (1) Asset synchronization: Synchronize the repaired qualified data to the data asset module of the water conservancy trusted data space, update the data quality tags (such as "high quality after repair" and "original high quality"), and associate the full-link information of "abnormal source tracing results - repair plan - verification record"; (2) Service output: Provides standardized API interfaces to support water conservancy business systems (flood forecasting system, water resource scheduling system) in calling data with quality traceability information; (3) Report generation: Generate a visual data quality report, including data quality level distribution, root cause statistics, and comparison of repair effects, and supports export in PDF / Excel format.

[0030] Water Conservancy Scenario 1 - Handling Abnormal Water Level Data During Flood Control: During the typhoon, a hydrological station collected water level data every 15 minutes. After the trusted data space accessed the data, it was found that some data contained outliers.

[0031] Feature extraction: Extract the basic features of the data (value range exceeds the device range, collection timestamp is during the typhoon period), business features (business tag "flood control emergency", priority level P1), and circulation link features (collection device number S102, transmission protocol MQTT); Source tracing reasoning: When calculating dynamic weights, the weight of the flood control business impact factor is temporarily increased to 0.9. After inputting into the dynamic weight Bayesian network model, the output root cause probability of "data acquisition link - sensor failure" is 89%. Tiered repair: Since it belongs to P1 level data, repair is performed first, using the "average of the three nearest collection cycles to complete + push equipment maintenance warning to the operation and maintenance terminal" scheme; Verification output: After repair, the data is re-entered into the network with an anomaly probability of 3%, which is considered qualified. The data is synchronized to the data asset module and marked as "high quality after repair" for use by the flood forecasting system.

[0032] Water Conservancy Scenario 2 - Handling Missing Daily Water Resources Statistics: There are some missing data on daily water resource intake and use in a certain irrigation district. The business label is "water resource assessment" and the priority level is P3. Feature extraction: Extract basic features (missing data fields, collection timestamp is for daily time periods), business features (priority P3 level), and flow process features (transmission process uses HTTP protocol, latency index 1.2s); Source tracing reasoning: When calculating dynamic weights, a standard configuration is used. After inputting the dynamic weights Bayesian network model, the output probability of "transmission link - network interruption" as the root cause is 83%. Tiered repair: Repair is carried out in a priority order of P3, and the missing data is filled by the "SSL-encrypted retransmission verification" scheme. Verification output: After repair, the probability of data anomalies is 2%. If the data is qualified, it will be synchronized to the data asset module and a quality report will be generated for use in water resource statistical analysis.

[0033] Example 2: A system for tracing and repairing water conservancy data quality anomalies, including... Data access layer: used to acquire multi-source water conservancy data; Feature extraction layer: used to extract basic features, business features, and flow process features of data and generate a standardized feature vector set; Dynamic Weight Bayesian Network Model Layer: Through a three-layer Bayesian network topology structure of "transfer links - business impact factors - anomaly types", the dynamic weights of each transfer link under different water conservancy business priorities are calculated using the entropy weight method and the analytic hierarchy process. The anomaly types are calculated using the dynamic weights. A dynamic weight Bayesian network model adapted to the water conservancy scenario is trained using labeled historical sample data. The standardized feature vectors are input into the trained dynamic weight Bayesian network model, and the probability distribution of the root causes of quality anomalies is output. Source tracing and repair layer: used to calculate the posterior probability of each abnormal root cause using Bayesian formula, determine the root cause and output source tracing results; repair in stages according to different source tracing results, and re-input the repaired data into the dynamic weighted Bayesian network to infer and calculate the abnormal probability for verification; Asset Service Layer: This layer synchronizes the repaired and qualified data to the data asset module of the Water Resources Trusted Data Space for service output and to generate a visual data quality report.

[0034] Source tracing and repair layer includes Anomaly tracing unit: used to output the probability distribution of the root cause of quality anomalies through Bayesian inference, determine the root cause, and output the tracing result; Repair Execution Unit: Built-in water conservancy industry-specific repair rule library, used to generate repair plans based on source tracing results and execute graded repairs; Repair and Validation Unit: Used to re-input the repaired data into the dynamic weighted Bayesian network model, verify the repair effect, and output qualified data.

[0035] A storage medium for tracing and repairing water conservancy data quality anomalies is provided. The storage medium stores a program, which, when executed by a processor, implements the steps of the method for tracing and repairing water conservancy data quality anomalies as described in Example 1.

[0036] The above is a further description of the present invention in conjunction with specific embodiments, and the scope of protection of the present invention is not limited thereto.

Claims

1. A method for tracing and repairing anomalies in water conservancy data quality, characterized by: The steps include the following: S1. Collect multi-source water conservancy data, extract the basic characteristics, business characteristics, and circulation characteristics of the data, and generate a standardized feature vector set; S2. Construct a three-layer Bayesian network topology structure of "transfer links - business impact factors - anomaly types", and use the entropy weight method and the analytic hierarchy process to calculate the dynamic weights of each transfer link under different water conservancy business priorities. Use the dynamic weights to calculate the anomaly types. Use labeled historical sample data to train a dynamic weight Bayesian network model adapted to the water conservancy scenario. S3. Input the standardized feature vector generated in step S1 into the trained dynamic weight Bayesian network model, output the probability distribution of the root cause of quality anomalies, calculate the posterior probability of each root cause of anomalies using the Bayesian formula, determine the root cause and output the source tracing result. S4. Based on different source tracing results, perform graded repairs and re-input the repaired data into a dynamic weighted Bayesian network to infer and calculate the anomaly probability for verification; S5. Repair qualified data and synchronize it to the data asset module of the water conservancy trusted data space for service output and generation of visual data quality reports.

2. The method for tracing and repairing water conservancy data quality anomalies according to claim 1, characterized in that, The water conservancy data mentioned in step S1 includes hydrological monitoring data, water conservancy project operation and maintenance data, water resource allocation data, and meteorological shared data; The basic features are the extracted attributes of the data itself, including data type identifier, collection timestamp, numerical range, and precision level; the business features are the extracted features related to water conservancy business, including associated water conservancy objects, business scenario tags, and business priorities; the circulation link features are the extracted features of the entire data circulation chain, including the collection link, transmission link, conversion link, and storage link.

3. The method for tracing and repairing water conservancy data quality anomalies according to claim 1, characterized in that, The three-layer Bayesian network topology described in step S2 includes a parent node (data flow link), an intermediate adjustment node (business impact factor), and a child node (data quality anomaly type).

4. The method for tracing and repairing water conservancy data quality anomalies according to claim 1, characterized in that, The calculation process for the dynamic weights is as follows: Objective weights are calculated using the entropy weight method, based on historical data quality feedback from the water conservancy trusted data space. Subjective weights were calculated using the analytic hierarchy process and combined with the experience of water conservancy experts. Generate fusion weights: Combine objective weights and subjective weights in a 6:4 weighted ratio to obtain dynamic weights; The weight update cycle is synchronized with the data collection frequency.

5. The method for tracing and repairing water conservancy data quality anomalies according to claim 1, characterized in that, The anomaly calculation process using dynamic weights is as follows: Dynamic weights are used as conditional probability correction coefficients from nodes in the transition links of a Bayesian network to nodes of anomaly types, embedded in the probabilistic inference process of a three-layer network. Let the transition link be... , These correspond to acquisition, transmission, conversion, and storage, respectively. The business impact factor is B, and the anomaly type is [missing information]. , If these correspond to missing, outlier, inconsistent, and duplicate values ​​respectively, then exception type A... j The probability calculation formula is: , In the formula: For the link Dynamic weights; For the next stage of business impact factor B Caused abnormality The basic conditional probability; For the link The prior probability is calculated; the probability values ​​of the four anomaly types are calculated according to the above formula, and the one with the highest probability is taken as the main anomaly type of the current data; when multiple anomalies coexist, the anomaly types with a probability higher than 30% are output together.

6. The method for tracing and repairing water conservancy data quality anomalies according to claim 1, characterized in that, In step S3, the posterior probability of each abnormal root cause is calculated using Bayes' theorem, as follows: , in, Let R be the posterior probability of the root cause R corresponding to anomaly type A. Let R be the conditional probability that the root cause R leads to the abnormality A. Let R be the prior probability of the root cause, and W be the dynamic weight. Let A be the marginal probability of anomaly type A.

7. The method for tracing and repairing water conservancy data quality anomalies according to claim 1, characterized in that, The determination of the root cause is as follows: a root cause probability threshold is set, and when the probability of a certain root cause reaches the threshold, it is determined to be the main root cause; if there are multiple root causes with similar probabilities, further screening is carried out in combination with business impact factors; for scenarios where multiple anomaly types coexist, the joint root cause probability is calculated through the conditional dependency relationship between nodes, and a comprehensive tracing result is output.

8. A system for tracing and repairing anomalies in water conservancy data quality, characterized by: include Data access layer: used to acquire multi-source water conservancy data; Feature extraction layer: used to extract basic features, business features, and flow process features of data and generate a standardized feature vector set; Dynamic Weight Bayesian Network Model Layer: Through a three-layer Bayesian network topology structure of "transfer links - business impact factors - anomaly types", the dynamic weights of each transfer link under different water conservancy business priorities are calculated using the entropy weight method and the analytic hierarchy process. The anomaly types are calculated using the dynamic weights. A dynamic weight Bayesian network model adapted to the water conservancy scenario is trained using labeled historical sample data. The standardized feature vectors are input into the trained dynamic weight Bayesian network model, and the probability distribution of the root causes of quality anomalies is output. Source tracing and repair layer: used to calculate the posterior probability of each abnormal root cause using Bayesian formula, determine the root cause and output source tracing results; repair in stages according to different source tracing results, and re-input the repaired data into the dynamic weighted Bayesian network to infer and calculate the abnormal probability for verification; Asset Service Layer: This layer synchronizes the repaired and qualified data to the data asset module of the Water Resources Trusted Data Space for service output and to generate a visual data quality report.

9. The system for tracing and repairing water conservancy data quality anomalies according to claim 8, characterized in that, The aforementioned source tracing and repair layer includes Anomaly tracing unit: used to output the probability distribution of the root cause of quality anomalies through Bayesian inference, determine the root cause, and output the tracing result; Repair Execution Unit: Built-in water conservancy industry-specific repair rule library, used to generate repair plans based on source tracing results and execute graded repairs; Repair and Validation Unit: Used to re-input the repaired data into the dynamic weighted Bayesian network model, verify the repair effect, and output qualified data.

10. A storage medium for tracing and repairing water conservancy data quality anomalies, wherein the storage medium stores a program, characterized in that... When the program is executed by the processor, it implements the steps of the method for tracing and repairing water conservancy data quality anomalies as described in any one of claims 1-7.