Multi-source data base plate construction method and system for flood control of drainage basin
By constructing an entity relationship network, deploying a lightweight AI quality control module, and using drone swarms for emergency data collection, the problem of inconsistent multi-source data integration was solved, enabling real-time control and storage optimization of data quality, and improving the adaptability and efficiency of flood control decision-making.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-17
- Publication Date
- 2026-04-03
AI Technical Summary
Existing technologies lack dynamic confidence assessment mechanisms in data resource management, resulting in inconsistent results from multi-source data integration, insufficient real-time quality control mechanisms, delayed anomaly detection response, unoptimized storage solutions, a lack of intelligent triggering mechanisms for version control, and low efficiency of decision support systems.
An entity relationship network is constructed using ontology modeling technology in the field of flood control. Cross-source field similarity is calculated and weighted fusion is performed through semantic matching algorithm. A lightweight AI quality control module is deployed at the edge node for real-time detection, triggering emergency supplementary data collection by drone swarm. A four-dimensional spatiotemporal data cube is constructed and stored through blockchain. A flood control decision knowledge graph is generated and a digital twin is integrated to optimize the decision-making platform.
Significantly improves the accuracy and consistency of data integration, ensures the real-time and reliability of data quality control, saves storage resources, supports historical status retrospective analysis, enhances the adaptability and efficiency of flood control strategies, and achieves closed-loop decision execution.
Smart Images

Figure CN121786416A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data resource management technology, specifically a method and system for constructing a multi-source data base plate for watershed flood control. Background Technology
[0002] The field of data resource management technology mainly involves the collection, integration, storage, quality control, and intelligent application of multi-source heterogeneous data. Its core objective is to maximize the value of data assets through standardized processes and advanced algorithms. Specifically, this includes establishing a unified data model to resolve format differences, developing a real-time quality control mechanism to ensure data credibility, and constructing a spatiotemporal index to improve retrieval efficiency. Among these, a method and system for constructing a multi-source data base for watershed flood control refers to a method and system architecture for constructing a watershed-level flood control-specific data base by integrating multi-source heterogeneous data such as meteorological, hydrological, and geographic information, and employing semantic alignment, edge computing quality control, blockchain versioning management, and digital twin optimization technologies.
[0003] In data resource management, while unified data models address format differences, existing technologies rely on static rules to handle multi-source conflicts and lack dynamic confidence assessment mechanisms. This leads to inconsistent integration results, such as unresolved discrepancies between meteorological and hydrological data, causing errors in subsequent analysis. Real-time quality control mechanisms are underdeveloped, with processes largely concentrated on centralized servers. Edge devices have limited processing capabilities, resulting in delayed anomaly detection responses. For example, failure to promptly re-sample data during sudden water level changes reduces data reliability and delays decision-making. While spatiotemporal indexes improve retrieval efficiency, storage solutions lack optimized management of changes, resulting in large storage space requirements and a lack of intelligent triggering mechanisms for version control. Critical states are not captured promptly; for instance, snapshots of data during flood peaks are not automatically saved, hindering historical backtracking analysis. Decision support systems lack integration of historical knowledge, have rigid service composition strategies, and fail to build dynamic knowledge networks, hindering adaptive optimization and leading to inefficiency in emergency dispatch. For example, resource allocation is not based on past experience, prolonging response time and increasing risk. Summary of the Invention
[0004] To address the problem that existing technologies in data resource management, while using unified data models to resolve format differences, rely on static rules to handle multi-source conflicts and lack dynamic confidence assessment mechanisms, leading to inconsistent integration results, such as the ineffective reconciliation of deviations between meteorological and hydrological data, which causes errors in subsequent analysis, this application provides a method and system for constructing a multi-source data base for watershed flood control.
[0005] This application provides a method and system for constructing a multi-source data base plate for watershed flood control, which adopts the following technical solution: A method for constructing a multi-source data base for watershed flood control includes the following steps: S1: Based on multi-source heterogeneous data from meteorological satellites, hydrological stations, and geographic information, an entity relationship network is constructed using ontology modeling technology in the field of flood control. Cross-source field similarity is calculated using a semantic matching algorithm based on BERT, and conflicting data are weighted and fused using a long short-term memory network confidence evaluation model to generate a semantic fusion base. S2: Based on the real-time data stream of the semantic fusion base, a lightweight AI quality control module is deployed at the edge node of the watershed. Physical rule verification and isolated forest algorithm are used to detect outliers, triggering the drone cluster to perform emergency supplementary data collection through YOLOv7 visual recognition, generating enhanced quality inspection data; S3: Based on enhanced quality inspection data, a four-dimensional spatiotemporal data cube is constructed, and Delta encoding technology is used to store only the changes. When the water level changes abruptly, a snapshot is automatically triggered through the cumulative sum control chart algorithm, and the Hyperledger Fabric blockchain is used to store the data version hash and generate spatiotemporal version chain data. S4: Based on historical flood control records in spatiotemporal version chain data, a bidirectional long short-term memory-conditional random field entity extraction model is used to extract "event-measure-effect" triples. A flood control knowledge network is constructed through the graph database Neo4j, and the Q-learning algorithm is used to optimize the service composition strategy to generate a flood control decision knowledge graph. S5: Based on the flood control decision knowledge graph and real-time spatiotemporal version chain data, a watershed BIM model is constructed according to the CityGML 3.0 standard; SCADA sensor data is injected in real time through the Apache Kafka stream processing platform, and the acquisition frequency is dynamically adjusted by using a near-end strategy optimization reinforcement learning algorithm to output the flood diversion scheme to the gate control system and generate a twin optimization decision base.
[0006] Preferably, the S1-based generative semantic fusion substrate includes the following steps: S101: Based on raw data from meteorological satellites, hydrological stations, and geographic information, a flood control ontology model is generated using ontology modeling technology in the field of flood control. S102: Based on the flood control ontology model, a bidirectional encoder based on a converter is used to represent the semantic matching algorithm, which converts satellite remote sensing metadata, hydrological station JSON messages, and CAD topographic map attribute tables into vectors, calculates cross-source field cosine similarity, and generates a semantic alignment mapping table. S103: Based on the semantic alignment mapping table, a long short-term memory network confidence evaluation model is used for conflicting data. The input data accuracy level and time freshness features are used to output dynamic weighting coefficients and generate a semantic fusion substrate.
[0007] Preferably, the generation of enhanced quality inspection data based on S2 includes the following steps: S201: Real-time data stream based on semantic fusion platform, using physical rule verification engine, verifies the rationality of water level-flow relationship through Saint-Venant hydrodynamic equations, and generates physically compliant data stream; S202: Based on physical compliance data flow, deploy a lightweight isolated forest anomaly detection algorithm on edge RTU devices to identify water level gauge failure anomalies caused by sudden rainstorms and generate anomaly label datasets; S203: Based on the anomaly marker dataset, trigger the UAV swarm to use the YOLOv7 visual recognition algorithm to capture images of water surface markers, cross-validate them with satellite synthetic aperture radar images, and generate quality inspection enhancement data.
[0008] Preferably, the generation of spatiotemporal version chain data based on S3 includes the following steps: S301: Based on enhanced quality inspection data, construct a four-dimensional spatiotemporal data cube, and use Delta encoding technology to store only the change in flooding range to generate a spatiotemporal change matrix; S302: Based on the spatiotemporal change matrix, the cumulative sum control chart mutation detection algorithm is adopted to automatically trigger a full basin snapshot when the real-time water level exceeds the threshold, and generate a key state snapshot set; S303: Based on the key state snapshot set, the Hyperledger Fabric blockchain platform is used to generate Merkle tree hash values for data versions and store them on the chain to generate spatiotemporal version chain data.
[0009] Preferably, the generation of a flood control decision knowledge graph based on S4 includes the following steps: S401: Based on historical flood control reports in spatiotemporal version chain data, a bidirectional long short-term memory-conditional random field entity extraction model is used to identify the "breach location - material scheduling - loss reduction" triplet and generate a flood control entity triplet library. S402: Based on the flood control entity triplet library, a topology network is constructed through the Neo4j graph database to establish a relationship chain of "rainstorm event → service to be called" and generate a service relationship graph; S403: Based on the service relationship graph, the Q-learning reinforcement learning algorithm is used to optimize the service combination order and generate a flood control decision knowledge graph.
[0010] Preferably, the S5-based generative twin optimization decision platform includes the following steps: S501: Based on flood control decision-making knowledge graph and spatiotemporal version chain data, integrate LiDAR point cloud and oblique photography data according to CityGML 3.0 standard to generate a digital twin of the watershed; S502: Based on the watershed digital twin, SCADA sensor data streams are injected in real time through the Apache Kafka stream processing platform to generate a virtual-real synchronized data pipeline; S503: Based on a virtual-real synchronous data pipeline, it adopts a near-end strategy optimization reinforcement learning algorithm, uses the flood peak prediction accuracy as the reward function to dynamically adjust the UAV flight frequency, and generates an adaptive acquisition strategy. S504: A flood diversion scheme based on an adaptive acquisition strategy output, which uses a robot operating system to control the gate actuator to complete physical intervention and generate a twin-optimized decision base.
[0011] A multi-source data base construction system for watershed flood control includes the following modules: a data fusion quality control module, which constructs an entity relationship network based on original multi-source data using flood control ontology modeling technology, performs cross-source field alignment using a transducer-based bidirectional encoder representation semantic matching algorithm, and handles data conflicts using a long short-term memory network confidence evaluation model; deploys a physical rule verification engine and an isolated forest anomaly detection algorithm at edge nodes for real-time quality control, and finally drives an UAV cluster to perform emergency supplementary data collection using the YOLOv7 visual recognition algorithm to generate fused quality inspection data; The data fusion quality control module includes an ontology modeling submodule, a semantic alignment submodule, and a real-time quality control submodule; The version knowledge construction module constructs a four-dimensional spatiotemporal data cube based on fused quality inspection data and uses Delta encoding technology to store changes. It triggers key state snapshots through cumulative sum and control graph mutation detection algorithms and uses the Hyperledger Fabric blockchain platform for hash storage. Based on historical flood control data, it extracts triples using a bidirectional long short-term memory-conditional random field entity extraction model and constructs a knowledge network through the graph database Neo4j. Finally, it uses Q-learning reinforcement learning algorithm to optimize service composition strategies and generate a flood control knowledge base. The version knowledge construction module includes a spatiotemporal cube submodule, a snapshot management submodule, and a knowledge graph construction submodule; The twin decision optimization module, based on the flood control knowledge base, integrates LiDAR and oblique photography data to construct a digital model according to the CityGML 3.0 standard; injects SCADA sensor data streams in real time through the Apache Kafka streaming platform; uses a near-end strategy optimization reinforcement learning algorithm to dynamically adjust the acquisition frequency with flood peak prediction accuracy as the reward function; outputs the flood diversion plan to the gate control system to perform physical intervention, and generates the twin optimization decision base. The twin decision optimization module includes a twin construction submodule, a data synchronization submodule, and a strategy optimization submodule.
[0012] Preferably, the ontology modeling submodule generates a flood control ontology model based on the original multi-source data and using ontology modeling technology in the field of flood control. The semantic alignment submodule, based on the flood control ontology model, uses a converter-based bidirectional encoder to represent the semantic matching algorithm, calculates the cosine similarity of cross-source fields of satellite remote sensing metadata and hydrological station JSON messages, achieves semantic-level alignment, and generates a semantic alignment mapping table. The real-time quality control submodule, based on a semantic alignment mapping table, uses a physical rule verification engine and an isolated forest anomaly detection algorithm to identify data anomalies and triggers a drone swarm to use the YOLOv7 visual recognition algorithm for emergency supplementary data collection, generating fused quality inspection data.
[0013] Preferably, the spatiotemporal cube submodule constructs a four-dimensional spatiotemporal data cube based on fused quality inspection data, and uses Delta encoding technology to store only the change in flooding range, generating a spatiotemporal change matrix; The snapshot management submodule, based on the spatiotemporal change matrix and using the cumulative sum control chart mutation detection algorithm, automatically triggers a key state snapshot of the entire basin when the real-time water level exceeds the threshold, generating a key state snapshot set; The knowledge graph construction submodule extracts "event-measure-effect" triples from historical flood control reports based on key state snapshot sets and a bidirectional long short-term memory-conditional random field entity extraction model. It then constructs a knowledge network using the graph database Neo4j and finally uses the Q-learning reinforcement learning algorithm to optimize the service composition strategy and generate a flood control knowledge base.
[0014] Preferably, the digital twin construction submodule, based on the flood control knowledge base, integrates LiDAR point cloud and UAV oblique photography data according to the CityGML 3.0 standard to construct a watershed digital model and generate a watershed digital twin; The data synchronization submodule, based on the watershed digital twin, injects SCADA sensor data streams in real time through the Apache Kafka stream processing platform, establishes a physical and virtual data interaction channel, and generates a virtual-physical synchronized data pipeline. The strategy optimization submodule, based on the virtual-real synchronous data pipeline, adopts a near-end strategy optimization reinforcement learning algorithm, dynamically adjusts the data acquisition frequency with the flood peak prediction accuracy as the reward function, and outputs the flood diversion scheme to the gate control system, generating a twin optimization decision base.
[0015] In summary, this application includes at least one of the following beneficial technical effects: An entity relationship network is constructed through ontology modeling in the flood control field. A transducer-based semantic matching algorithm is used to calculate cross-source field similarity. A long short-term memory network confidence assessment model is employed to weightedly fuse conflicting data, significantly improving the accuracy and consistency of data integration and reducing semantic bias and conflict risks during multi-source heterogeneous data fusion. In real-time data stream processing, lightweight artificial intelligence is deployed at edge nodes, combining physical rule verification and the isolated forest algorithm to detect outliers. This triggers an unmanned aerial vehicle (UAV) swarm to perform emergency data collection via visual recognition, ensuring the real-time nature and reliability of data quality control, shortening anomaly response time, and enhancing data credibility. A four-dimensional spatiotemporal data cube is constructed, using variable storage. The technology retains only the differences, automatically triggers key state snapshots by combining cumulative and control chart algorithms, and uses blockchain to store data version hashes, efficiently saving storage resources and enhancing data traceability. It supports historical state backtracking, extracts event-measure-effect triples from historical flood control records, builds a knowledge network using a graph database, and applies reinforcement learning algorithms to optimize service combinations, forming an intelligent decision support framework. This improves the adaptability and efficiency of flood control strategies, integrates real-time data to build digital models according to standards, injects sensor data through a stream processing platform, dynamically adjusts the acquisition frequency using reinforcement learning, outputs optimized solutions to control physical equipment, achieves closed-loop decision execution, and enhances the scientific nature and operability of flood control response.
[0016] By constructing an entity relationship network through ontology modeling in the flood control field, a transducer-based bidirectional encoder is used to represent the semantic matching algorithm to calculate cross-source field similarity. A long short-term memory network confidence assessment model is employed to handle data conflicts, significantly improving the accuracy and consistency of data integration and reducing semantic bias and conflict risks during the fusion of multi-source heterogeneous data. In real-time data processing, a physical rule verification engine and an isolated forest anomaly detection algorithm are deployed at edge nodes, combined with visual recognition-driven UAV swarm emergency data collection, ensuring the real-time performance and reliability of data quality control, shortening anomaly response time, and enhancing data credibility. A four-dimensional spatiotemporal data cube is constructed, employing a change quantity storage technique to retain only the differing parts, combined with cumulative... The control chart mutation detection algorithm automatically triggers key state snapshots and performs hash storage through a blockchain platform, efficiently saving storage resources and enhancing data traceability. It supports historical state retrospective analysis, extracts event-measure-effect triples from historical flood control records, constructs a knowledge network using a graph database, and applies reinforcement learning algorithms to optimize service combination strategies, forming an intelligent decision support framework. This improves the adaptability and efficiency of flood control strategies. The integrated digital model is constructed according to standards, and sensor data is injected in real time through a stream processing platform. Reinforcement learning is used to dynamically adjust the acquisition frequency with flood peak prediction accuracy as the reward function, outputting flood diversion schemes to control physical equipment, realizing closed-loop decision execution, and enhancing the scientific nature and operability of flood control response. Attached Figure Description
[0017] Figure 1 This is a schematic diagram of the main steps of this application; Figure 2 This is a detailed schematic diagram of S1 in this application; Figure 3 This is a detailed schematic diagram of S2 in this application; Figure 4 This is a detailed schematic diagram of S3 in this application; Figure 5 This is a detailed schematic diagram of S4 in this application; Figure 6 This is a detailed schematic diagram of S5 in this application; Figure 7 The system block diagram is for application purposes. Detailed Implementation
[0018] The following is in conjunction with the appendix Figures 1-7 This application will be described in further detail.
[0019] See Figure 1 A method for constructing a multi-source data base for watershed flood control includes the following steps: S1: Based on multi-source heterogeneous data from meteorological satellites, hydrological stations, and geographic information, an entity relationship network is constructed using ontology modeling technology in the field of flood control. Cross-source field similarity is calculated using a semantic matching algorithm based on BERT, and conflicting data are weighted and fused using a long short-term memory network confidence evaluation model to generate a semantic fusion base. S2: Based on the real-time data stream of the semantic fusion base, a lightweight AI quality control module is deployed at the edge node of the watershed. Physical rule verification and isolated forest algorithm are used to detect outliers, triggering the drone cluster to perform emergency supplementary data collection through YOLOv7 visual recognition, generating enhanced quality inspection data; S3: Based on enhanced quality inspection data, a four-dimensional spatiotemporal data cube is constructed, and Delta encoding technology is used to store only the changes. When the water level changes abruptly, a snapshot is automatically triggered through the cumulative sum control chart algorithm, and the Hyperledger Fabric blockchain is used to store the data version hash and generate spatiotemporal version chain data. S4: Based on historical flood control records in spatiotemporal version chain data, a bidirectional long short-term memory-conditional random field entity extraction model is used to extract "event-measure-effect" triples. A flood control knowledge network is constructed through the graph database Neo4j, and the Q-learning algorithm is used to optimize the service composition strategy to generate a flood control decision knowledge graph. S5: Based on the flood control decision knowledge graph and real-time spatiotemporal version chain data, a watershed BIM model is constructed according to the CityGML 3.0 standard; SCADA sensor data is injected in real time through the Apache Kafka stream processing platform, and the acquisition frequency is dynamically adjusted by using a near-end strategy optimization reinforcement learning algorithm to output the flood diversion scheme to the gate control system and generate a twin optimization decision base.
[0020] An entity relationship network is constructed through ontology modeling in the flood control field. A transducer-based semantic matching algorithm is used to calculate cross-source field similarity. A long short-term memory network confidence assessment model is employed to weightedly fuse conflicting data, significantly improving the accuracy and consistency of data integration and reducing semantic bias and conflict risks during multi-source heterogeneous data fusion. In real-time data stream processing, lightweight artificial intelligence is deployed at edge nodes, combining physical rule verification and the isolated forest algorithm to detect outliers. This triggers an unmanned aerial vehicle (UAV) swarm to perform emergency data collection via visual recognition, ensuring the real-time nature and reliability of data quality control, shortening anomaly response time, and enhancing data credibility. A four-dimensional spatiotemporal data cube is constructed, using variable storage. The technology retains only the differences, automatically triggers key state snapshots by combining cumulative and control chart algorithms, and uses blockchain to store data version hashes, efficiently saving storage resources and enhancing data traceability. It supports historical state backtracking, extracts event-measure-effect triples from historical flood control records, builds a knowledge network using a graph database, and applies reinforcement learning algorithms to optimize service combinations, forming an intelligent decision support framework. This improves the adaptability and efficiency of flood control strategies, integrates real-time data to build digital models according to standards, injects sensor data through a stream processing platform, dynamically adjusts the acquisition frequency using reinforcement learning, outputs optimized solutions to control physical equipment, achieves closed-loop decision execution, and enhances the scientific nature and operability of flood control response.
[0021] See Figure 2 The S1-based semantic fusion substrate includes the following steps: S101: Based on raw data from meteorological satellites, hydrological stations, and geographic information, a flood control ontology model is generated using ontology modeling technology in the field of flood control. Based on raw data from meteorological satellites, hydrological stations, and geographic information systems, this raw data captures cloud images and water vapor distribution information through satellite remote sensing equipment. Hydrological monitoring stations record water level and flow changes in real time, and geographic information systems provide watershed topography and land use maps. In flood control scenarios, such as the Yangtze River basin flood control project, satellite data includes infrared band images, hydrological station data is stored in time-series format, and geographic information data involves digital elevation models. Ontology modeling techniques in the flood control field are employed. Specifically, the Protégé software tool is used to manually define core entities such as rainfall, water level, and inundation area, and establish relationship networks such as rainfall intensity affecting water level changes, and water level changes relating to inundation ranges. Entity attribute settings refer to flood control standard documents. The intensity attribute is defined in millimeters per hour, the water level entity is defined in elevation in meters, and the geographic entity is defined in slope in degrees. During modeling, the consistency of entities from different data sources is compared. For example, satellite rainfall data is compared with hydrological station measured data. If the satellite rainfall value is higher than the hydrological station value, it is determined whether the data difference is within an acceptable range. The acceptable range is based on historical flood event experience; a rainfall deviation of less than 20% is considered consistent; otherwise, a conflict is marked. The entity relationship weights are adjusted, with weight coefficients set based on expert knowledge. The weight of rainfall on water level is set to 0.6, and the weight of topography on inundation is set to 0.4. Entity definitions are input through the software interface, relationship mapping is performed, and a structured ontology model is generated, including entity classes, attribute tables, and relationship graphs.
[0022] S102: Based on the flood control ontology model, a bidirectional encoder based on a converter is used to represent the semantic matching algorithm, which converts satellite remote sensing metadata, hydrological station JSON messages, and CAD topographic map attribute tables into vectors, calculates cross-source field cosine similarity, and generates a semantic alignment mapping table. Based on a flood control ontology model, which includes definitions of rainfall, water level, geographic entities, and their relationships, the ontology model provides a semantic framework for flood control applications such as the Yellow River flood control system. Satellite remote sensing metadata, obtained from Sentinel satellites, includes band reflectance values. Hydrological station JSON messages store water level and flow rates. CAD topographic map attribute tables contain coordinate and elevation information. A transducer-based bidirectional encoder is used to represent the semantic matching algorithm. Specifically, a Python script calls a pre-trained model to convert metadata, messages, and attribute table text fields into numerical vectors with a dimension of 300. The cosine phase of the cross-source fields is then calculated. Similarity is calculated for fields such as satellite rainfall and the ontology rainfall entity, and hydrological station water level and the ontology water level entity. The calculation process involves vector dot product and modulus operation. The dot product result is divided by the modulus product. The similarity value ranges from zero to one. A value higher than 0.7 is considered high similarity and triggers automatic mapping. A value lower than 0.3 is considered low similarity and requires manual review. In the example, the similarity between the satellite band field and the ontology rainfall field is calculated to be 0.85, the similarity between the hydrological station JSON water level field and the ontology water level field is 0.9, and the similarity between the CAD elevation field and the ontology geographic field is 0.75. The mapping relationship table is output through batch processing by script, generating a semantic alignment mapping table.
[0023] S103: Based on the semantic alignment mapping table, a long short-term memory network confidence evaluation model is used for conflicting data. The input data accuracy level and time freshness features are used to output dynamic weighting coefficients and generate a semantic fusion substrate.
[0024] Based on a semantic alignment mapping table that records field similarity correspondences, in the Pearl River Basin flood control case, the mapping table identifies conflicting data, such as discrepancies between satellite rainfall values and hydrological station rainfall values. A Long Short-Term Memory (LSTM) network confidence assessment model is used for the conflicting data. Input data accuracy level features are defined based on the data source type: satellite data (high accuracy) is assigned to level one, hydrological station data (medium accuracy) to level two, and geographic data (low accuracy) to level three. The time freshness feature is based on the difference in collection time; the freshness value is calculated as the current time minus the collection time, in hours. Values less than three hours are considered high freshness, and values between three and twelve hours are considered medium freshness. For data exceeding twelve hours, the specific process uses the TensorFlow framework to build the model. The input layer receives the precision level and freshness value, the hidden layer processes the time series, and the output is a dynamic weighted coefficient with a range of zero to one. The model is trained on historical flood data, including conflicting data such as 100 mm of satellite rainfall and 80 mm of hydrological station rainfall. The precision level is 1:2, and the freshness value is 2 hours for satellite data and 4 hours for hydrological station data. The model calculates the weights as 0.7 for satellite data and 0.3 for hydrological station data. The weighted fusion is then performed to calculate the weighted average value, multiplying the satellite weights by the hydrological weights. The result generates fused data and a semantic fusion template.
[0025] See Figure 3The generation of enhanced quality inspection data based on S2 includes the following steps: S201: Real-time data stream based on semantic fusion platform, using physical rule verification engine, verifies the rationality of water level-flow relationship through Saint-Venant hydrodynamic equations, and generates physically compliant data stream; The real-time data stream of the semantic fusion platform includes water level and flow rate numerical sequences. This data comes from fused satellite meteorological, hydrological station monitoring, and geographic information systems. In the Yangtze River flood control scenario, water level data is represented by river height in meters, and flow rate data is represented by water volume per second in cubic meters per second. The data stream is timestamped and updated every second. A physical rule verification engine is used, specifically involving the invocation of the Saint-Venant hydrodynamic equations, which include continuity and momentum equations to describe the relationship between water level and flow rate. To verify the rationality, the current water level value is extracted, and the theoretical flow rate is calculated. The theoretical flow rate is estimated based on the water level value using an exponential function. The parameters of the exponential function are set according to the shape of the river cross-section; for rectangular channels, the parameter is set to 0.5, and for trapezoidal channels... The river channel is set to 0.6. The measured flow rate is compared with the theoretical flow rate, and the percentage deviation is calculated. The deviation is equal to the absolute value of the measured value minus the theoretical value, divided by the theoretical value, and multiplied by 100. The deviation range is from 0 to 100. A deviation of less than 5% is considered low deviation and reasonable. A deviation of 5% to 10% is moderate and requires verification. A deviation of more than 10% is high deviation and unreasonable. The reasonableness threshold is set to 8%. Based on historical flood event statistics, the threshold is adjusted to 10% during periods of high flood incidence. In the example, when the water level is high, such as 5 meters, the theoretical flow rate is calculated to be 100 cubic meters per second, while the measured flow rate is 110 cubic meters per second. The deviation is 10%, which is higher than the threshold and is marked as unreasonable. The data status is adjusted. Data that passes the verification retains its original value, while data that fails is marked as needing correction, and a physically compliant data stream is generated.
[0026] S202: Based on physical compliance data flow, deploy a lightweight isolated forest anomaly detection algorithm on edge RTU devices to identify water level gauge failure anomalies caused by sudden rainstorms and generate anomaly label datasets; The physically compliant data stream contains verified water level and flow time series data, stored as an array with one data point per second. In the Yellow River flood control case, the water level data series shows height changes, and the flow data series shows volume changes. The data stream is transmitted to an edge RTU device with limited memory, deploying a lightweight isolated forest anomaly detection algorithm. The algorithm constructs a binary tree structure, randomly segments data features, generates path lengths (shorter path lengths indicate a higher probability of anomalies), calculates anomaly scores ranging from 0 to 1, with scores above 0.7 considered high anomalies, scores between 0.3 and 0.7 considered moderate (requiring manual inspection), and scores below 0.3 considered low anomalies. Normally, the anomaly score threshold is set to 0.65. Based on equipment processing capacity and historical anomaly frequency, the threshold is lowered to 0.6 during flood season. When identifying level gauge failure anomalies caused by sudden heavy rain, the water level sequence is analyzed. If a sharp drop or rise occurs in the sequence, the rate of change is calculated. The rate of change equals the current value minus the previous value divided by the time interval. A rate of change higher than 0.1 meters per second is considered a significant change, marking a potential failure. In the example, the water level sequence suddenly drops from 5 meters to 4 meters, with a rate of change of 0.2 meters per second. The algorithm calculates an anomaly score of 0.75, which is higher than the threshold, triggering anomaly marking. A dataset containing timestamps, anomaly types, and confidence levels is generated, creating an anomaly-marked dataset.
[0027] S203: Based on the anomaly marker dataset, trigger the UAV swarm to use the YOLOv7 visual recognition algorithm to capture images of water surface markers, cross-validate them with satellite synthetic aperture radar images, and generate quality inspection enhancement data.
[0028] Anomaly marker datasets contain the time, location, and type information of anomalies, such as water level gauge failure markers at specific coordinates and times. In the Pearl River flood control project, the dataset identifies the location of anomalies, such as a section of the river. Coordinates are based on GPS data, and time is based on UTC timestamps. A drone swarm is triggered, consisting of multiple drones equipped with high-definition cameras, to fly to the anomaly point coordinates. The YOLOv7 visual recognition algorithm captures images of water surface markers. The algorithm processes image frames, detects preset markers such as poles or buoys, calculates the pixel coordinates of the marker positions, and converts them to the actual water level elevation. The elevation equals the baseline height minus the pixel offset multiplied by a scaling factor, which is based on the camera focal length and distance settings, with 0.1 meters per pixel. Afterwards, cross-validation is performed with satellite synthetic aperture radar imagery. Satellite imagery provides water surface elevation data, which is analyzed through radar reflection signals. The elevation calculated by the UAV is compared with the satellite elevation, and the absolute value of the difference is calculated. A difference less than 0.1 meters is considered low difference and verifies consistency. A difference between 0.1 and 0.3 meters is moderate and requires further verification. A difference greater than 0.3 meters is high difference and inconsistent. The verification threshold is set at 0.2 meters, based on sensor accuracy standards. In the example, the UAV image captures a marker, and the calculated water level elevation is 4.5 meters. The satellite image elevation is 4.6 meters, with a difference of 0.1 meters, which is below the threshold. The verification passes, and the data is marked as reliable. Enhanced data is generated, including the original data, verification results, and confidence level, to produce quality inspection enhanced data.
[0029] See Figure 4 The generation of spatiotemporal version chain data based on S3 includes the following steps: S301: Based on enhanced quality inspection data, construct a four-dimensional spatiotemporal data cube, and use Delta encoding technology to store only the change in flooding range to generate a spatiotemporal change matrix; The enhanced quality inspection data includes verified inundation range data sequences. These data come from real-time monitoring of the flood control system. In the Yangtze River basin flood control scenario, the inundation range represents the inundated area in square kilometers. The time series is stored at second intervals. The data cube construction involves defining four dimensions: longitude, latitude, elevation, and time. The longitude range covers the basin from the starting point of the east longitude to the ending point, the latitude range covers the starting point of the north latitude to the ending point, and the elevation ranges from zero meters to the highest point. The time interval is one second. When Delta encoding technology is applied, the inundation range value at the current moment is compared with the value at the previous moment to calculate the change. The change is equal to the current value minus the previous value. If the absolute value of the change is less than a threshold, it is considered as no change and is not stored. The threshold setting refers to historical inundation fluctuations. The threshold is adjusted during periods of high flood incidence. For example, a change percentage of less than 1% is considered a low change and is ignored. Changes from 1% to 5% are stored as medium changes, and changes greater than 5% must be stored. In the example, the inundation range at the previous moment was 100 square kilometers, and the current inundation range is 102 square kilometers, with a change of 2 square kilometers and a percentage of 2. If the change is higher than the low threshold, the change is stored. The generated matrix only contains the change value, saving storage resources.
[0030] S302: Based on the spatiotemporal change matrix, the cumulative sum control chart mutation detection algorithm is adopted to automatically trigger a full basin snapshot when the real-time water level exceeds the threshold, and generate a key state snapshot set; The spatiotemporal variation matrix contains a sequence of inundation range changes. In the Yellow River flood control case, the matrix provides a sequence of water level change data, one data point per second. A cumulative sum control chart abrupt change detection algorithm is used. The algorithm calculates the cumulative sum of the water level sequence, which equals the cumulative deviation between the sequence value and the target value. The target water level is set as the warning level, based on historical flood event statistics, for example, a warning level of 50 meters. The deviation at each time point is calculated; the deviation equals the measured water level minus the target value. The cumulative sum equals the current deviation. The upper control limit is set to 0.5 meters. The lower limit is set to 0.5 meters. A snapshot is triggered when the cumulative sum exceeds the upper limit or falls below the lower limit. Simultaneously, a snapshot is triggered directly when the real-time water level exceeds the threshold. The water level threshold is set based on the flood control plan standard, with the warning water level as the benchmark. A water level exceeding the warning water level by 0.5 meters is considered a high water level and must be triggered. The real-time water level is obtained by extracting the water level value from the matrix and calculating the deviation. For example, if the water level is 51 meters, the deviation is 1 meter. The cumulative sum is calculated for instance sequence deviations of 0.1 meter, 0.2 meter, and 0.3 meter. A cumulative sum of 0.6 meters triggers a snapshot when it exceeds the upper limit of 0.5 meters. A full-basin snapshot captures the inundation range data of all points and generates a snapshot set.
[0031] S303: Based on the key state snapshot set, the Hyperledger Fabric blockchain platform is used to generate Merkle tree hash values for data versions and store them on the chain to generate spatiotemporal version chain data.
[0032] The critical state snapshot set contains data on the inundation range of the entire river basin at critical moments of flooding. In the Pearl River flood control project, each snapshot stores data for all coordinate points at a given time. Utilizing the Hyperledger Fabric blockchain platform, the platform configures smart contracts and channels to generate Merkle tree hash values for data versions. During Merkle tree construction, data is processed in blocks, and a hash value is calculated for each block. The hash algorithm uses standard cryptographic functions, and the root hash is calculated hierarchically. On-chain notarization involves writing the root hash into the blockchain ledger. Data versions include timestamps and hash values, linked to form a chain structure. In this example, snapshot data is divided into multiple blocks, and the hash value of each block is calculated to construct the tree structure. After the root hash is generated, it is submitted to the blockchain via a smart contract to generate version chain data.
[0033] See Figure 5 The generation of a flood control decision knowledge graph based on S4 includes the following steps: S401: Based on historical flood control reports in spatiotemporal version chain data, a bidirectional long short-term memory-conditional random field entity extraction model is used to identify the "breach location - material scheduling - loss reduction" triplet and generate a flood control entity triplet library. Historical flood control reports in the spatiotemporal version chain data cover decisions and response measures under different weather conditions. These reports provide a rich data source for flood control decision-making. A bidirectional long short-term memory-conditional random field model is adopted, which can capture the contextual relationships of time series data and identify entity triples through feature extraction. The triples consist of "breach location", "material dispatch", and "loss reduction". In practical applications, the analysis of historical flood control reports extracts the "breach location" which may be in a certain area, such as a specific section of a river, and stores it as a node. The analysis of flood control material dispatch information identifies the corresponding materials and dispatch time, determines loss reduction measures, and sorts the generated entity triples based on relevance, records the most frequently occurring decision patterns, and further constructs a flood control entity triple library to facilitate subsequent queries and decision support. In the example, if a historical report shows that the water level reached a high-risk threshold at a specific breach location during a flood, the timeliness information on dispatching sandbags will be extracted and integrated into the triple.
[0034] S402: Based on the flood control entity triplet library, a topology network is constructed through the Neo4j graph database to establish a relationship chain of "rainstorm event → service to be called" and generate a service relationship graph; The flood control entity triplet database contains multiple historical decision-making information nodes. Before constructing the topology network, all valid node information needs to be extracted from the entity triplet database using the Neo4j graph database. The extracted information includes the location of the breach, the corresponding material dispatch records, and the resulting losses. The relationship chain is established based on the association between events and services. Service information involves services related to emergency response, such as drainage, damming, and early warning. Through the graph database function, a "rainstorm event → service to be called" relationship chain is created, and a weight is assigned to each node value. The weight setting is based on the effectiveness of historical decisions. The decision effect evaluation is obtained through backtracking calculations of historical data. If historical data shows that calling a certain service is effective, the weight is set to a higher value, and vice versa. In the example, if the dispatch record of rainstorm event A shows that the success rate of service calls is 80%, then the weight of that service is set to 0.8. Therefore, a complete service relationship graph can be constructed, which stores the frequent paths and dependencies between various nodes.
[0035] S403: Based on the service relationship graph, the Q-learning reinforcement learning algorithm is used to optimize the service combination order and generate a flood control decision knowledge graph.
[0036] The service relationship graph integrates the relationships between emergency services throughout history. Using a Q-learning reinforcement learning algorithm, it optimizes the service combination order. First, the state space is defined as different service combinations, such as combinations of services A, B, and C. A reward function is formulated, and rewards are calculated based on actual response efficiency and service quality evaluation information, using historical data feedback. This involves collecting successful and failed patterns from historical decisions to train the model and adjust the service call order to select the optimal service combination for specific rainstorm events. Through continuous trial and error and learning, the optimal combination is generated. The key is to evaluate the effectiveness of each combination in simulated operation. A time factor is introduced to calculate the critical value of the service combination. In the example, assuming the historical effectiveness of the service combination (A, B, C) is 70%, after algorithm optimization it reaches 75%, indicating that the optimization is effective. The final flood control decision knowledge graph summarizes this series of optimization results, forming a visualized decision path.
[0037] See Figure 6 The S5-based generative twin optimization decision platform includes the following steps: S501: Based on flood control decision-making knowledge graph and spatiotemporal version chain data, integrate LiDAR point cloud and oblique photography data according to CityGML 3.0 standard to generate a digital twin of the watershed; Flood control decision-making knowledge graph and spatiotemporal version chain data integrate historical flood decision-making information and geographic spatiotemporal records. In the Yangtze River basin flood control scenario, the knowledge graph provides the relationship between events and measures, the spatiotemporal version chain stores changes in inundation range, lidar point clouds capture terrain elevation features, and oblique photogrammetry data obtains building and surface texture features, all according to CityGML. The 3.0 standard defines the hierarchical structure of the 3D model. The integration process involves coordinate system alignment, matching point cloud coordinates with photographic coordinates, comparing coordinate deviations during the matching process, and calculating the deviation as the absolute value of the difference between the point cloud coordinates and the photographic coordinates. The deviation range is divided into low, medium, and high intervals. Low deviations are considered acceptable as they are less than the threshold, medium deviations require manual verification, and high deviations must be adjusted. The threshold setting refers to historical data accuracy reports. For example, the terrain elevation deviation threshold is set to allow small fluctuations within a unit length. In practice, if the point cloud elevation value is high and the photographic elevation value is low, the data can be directly fused after the deviation value is calculated if it is in the low interval; otherwise, it needs to be re-acquired or corrected. The fused data is used to construct the 3D model through software tools such as a GIS platform. In the example, the point cloud elevation shows a higher slope in the river section, while the photographic elevation shows a lower slope. The deviation is judged to be medium and needs to be verified and adjusted to generate a complete digital model.
[0038] S502: Based on the watershed digital twin, SCADA sensor data streams are injected in real time through the Apache Kafka stream processing platform to generate a virtual-real synchronized data pipeline; A watershed digital twin, comprising a 3D geographic model, is used in the Yellow River flood control case to simulate the river and its surrounding environment. Data topics are configured via the Apache Kafka streaming platform, defining sensor data types. SCADA sensor data streams include water level, flow rate, and meteorological parameters. The real-time injection process involves data stream classification, with water level and flow rate data streams processed separately. Data stream rate monitoring is implemented, with rate ranges categorized as low, medium, and high. Low rates are considered high-risk due to data latency, medium rates require optimization, and high rates are considered normal. Threshold settings are based on network bandwidth standards; for example, the water level data rate threshold is set as the number of transmissions per unit time. In practice, if the water level data stream rate is calculated as the transmission frequency, an alarm is triggered if it falls within the low range. During data injection, the real-time value is compared with the model value for consistency. Consistency is calculated as the absolute value of the model value minus the sensor value. Consistency ranges are categorized as high and low; high consistency is directly synchronized, while low consistency requires verification. Data streams are processed through Kafka consumer groups. In this example, if the sensor water level value is high and the model value is low, the model is recalibrated when consistency is low, generating a data pipeline.
[0039] S503: Based on a virtual-real synchronous data pipeline, it adopts a near-end strategy optimization reinforcement learning algorithm, uses the flood peak prediction accuracy as the reward function to dynamically adjust the UAV flight frequency, and generates an adaptive acquisition strategy. A virtual-real synchronized data pipeline connects sensor data with a virtual model. In the Pearl River flood control project, this pipeline transmits real-time water level and flow information. A near-end strategy optimization reinforcement learning algorithm defines the state space state as the UAV flight frequency. The flight frequency range is divided into low, medium, and high frequencies. Low frequencies are considered insufficient data collection, while high frequencies are considered a waste of resources. Threshold settings are based on historical flood peak events. Flood peak prediction accuracy is used as a reward function, with accuracy values divided into high and low. High accuracy is rewarded, while low accuracy is penalized. Accuracy is calculated as the absolute value of the predicted value minus the actual value. Accuracy intervals are defined: high accuracy is considered to have small errors, while low accuracy has large errors. The dynamic adjustment process analyzes the effect of the current frequency, compares accuracy values at different frequencies, calculates the accuracy difference, maintains the frequency if the difference is small, and adjusts the frequency if the difference is large. The adjustment magnitude is based on the learning rate parameter. The learning rate is set with reference to the algorithm's training history; for example, the initial frequency is set as the base value. In practice, if the current frequency is low, the accuracy is low, and the calculated difference is large, the frequency is increased. Feedback data is collected during algorithm operation. In the example, after the UAV flight frequency is adjusted from low to medium frequency, the accuracy improves, which is determined to be a high-precision maintenance strategy, generating a data collection plan.
[0040] S504: A flood diversion scheme based on an adaptive acquisition strategy output, which uses a robot operating system to control the gate actuator to complete physical intervention and generate a twin-optimized decision base.
[0041] The adaptive acquisition strategy outputs a flood diversion plan, which provides flood response instructions. In the Yangtze River flood control scenario, the plan specifies gate operation parameters and configures control instructions through the robot operating system. The instructions include gate opening degree and timing. The process of controlling the gate actuator involves sending execution signals. The signal strength monitoring intensity value range is divided into weak, medium, and strong. Weak intensity is considered as execution failure risk, and medium intensity is normal. The threshold setting is based on mechanical performance standards. For example, the opening signal strength threshold is set as the base value. When physical intervention is completed, the consistency between the actual opening degree and the plan opening degree is compared. The consistency calculation is the absolute value of the plan value minus the actual value. The consistency interval is divided into high and low. High consistency means success, and low consistency means failure. Instructions are sent through the operating system interface. In the example, if the plan opening degree is high and the actual opening degree is low, the signal is resent until the consistency is high, and a decision base is generated.
[0042] See Figure 7 A multi-source data base construction system for watershed flood control includes the following modules: a data fusion quality control module, which constructs an entity relationship network based on original multi-source data using flood control ontology modeling technology, performs cross-source field alignment using a transducer-based bidirectional encoder representation semantic matching algorithm, and handles data conflicts using a long short-term memory network confidence evaluation model; deploys a physical rule verification engine and an isolated forest anomaly detection algorithm at edge nodes for real-time quality control, and finally drives an UAV cluster to perform emergency supplementary data collection using the YOLOv7 visual recognition algorithm to generate fused quality inspection data; The data fusion quality control module includes an ontology modeling submodule, a semantic alignment submodule, and a real-time quality control submodule; The version knowledge construction module constructs a four-dimensional spatiotemporal data cube based on fused quality inspection data and uses Delta encoding technology to store changes. It triggers key state snapshots through cumulative sum and control graph mutation detection algorithms and uses the Hyperledger Fabric blockchain platform for hash storage. Based on historical flood control data, it extracts triples using a bidirectional long short-term memory-conditional random field entity extraction model and constructs a knowledge network through the graph database Neo4j. Finally, it uses Q-learning reinforcement learning algorithm to optimize service composition strategies and generate a flood control knowledge base. The version knowledge construction module includes a spatiotemporal cube submodule, a snapshot management submodule, and a knowledge graph construction submodule; The twin decision optimization module, based on the flood control knowledge base, integrates LiDAR and oblique photography data to construct a digital model according to the CityGML 3.0 standard; injects SCADA sensor data streams in real time through the Apache Kafka streaming platform; uses a near-end strategy optimization reinforcement learning algorithm to dynamically adjust the acquisition frequency with flood peak prediction accuracy as the reward function; outputs the flood diversion plan to the gate control system to perform physical intervention, and generates the twin optimization decision base. The twin decision optimization module includes a twin construction submodule, a data synchronization submodule, and a strategy optimization submodule.
[0043] By constructing an entity relationship network through ontology modeling in the flood control field, a transducer-based bidirectional encoder is used to represent the semantic matching algorithm to calculate cross-source field similarity. A long short-term memory network confidence assessment model is employed to handle data conflicts, significantly improving the accuracy and consistency of data integration and reducing semantic bias and conflict risks during the fusion of multi-source heterogeneous data. In real-time data processing, a physical rule verification engine and an isolated forest anomaly detection algorithm are deployed at edge nodes, combined with visual recognition-driven UAV swarm emergency data collection, ensuring the real-time performance and reliability of data quality control, shortening anomaly response time, and enhancing data credibility. A four-dimensional spatiotemporal data cube is constructed, employing change quantity storage technology to retain only the differing parts, combined with cumulative... The product sum control chart mutation detection algorithm automatically triggers critical state snapshots and performs hash-based notarization through a blockchain platform, efficiently saving storage resources and enhancing data traceability. It supports historical state retrospective analysis, extracts event-measure-effect triples from historical flood control records, constructs a knowledge network using a graph database, and applies reinforcement learning algorithms to optimize service composition strategies, forming an intelligent decision support framework. This improves the adaptability and efficiency of flood control strategies. Integrated digital models are built according to standards, and sensor data is injected in real-time through a stream processing platform. Reinforcement learning is used to dynamically adjust the acquisition frequency with flood peak prediction accuracy as the reward function, outputting flood diversion schemes to control physical equipment, achieving closed-loop decision execution, and enhancing the scientific and operational nature of flood control response. The ontology modeling submodule generates a flood control ontology model based on original multi-source data and using ontology modeling technology in the field of flood control. Based on meteorological satellite cloud images, hydrological station monitoring records, and geographic information data, this paper uses ontology modeling technology in the field of flood control to define core entities (such as rainfall, water level, and inundation area) and their attribute relationships. By constructing a structured semantic network through expert knowledge, the paper clarifies the logic of rainfall intensity affecting water level changes and water level changes relating to inundation range, thus forming a standardized flood control ontology model.
[0044] The semantic alignment submodule, based on the flood control ontology model, uses a converter-based bidirectional encoder to represent the semantic matching algorithm, calculates the cosine similarity of cross-source fields of satellite remote sensing metadata and hydrological station JSON messages, achieves semantic-level alignment, and generates a semantic alignment mapping table. Based on the flood control ontology model, heterogeneous fields such as satellite remote sensing metadata (band reflectivity) and hydrological station JSON messages (water level values) are converted into vectors. The cosine similarity between cross-source fields is calculated using a pre-trained semantic matching algorithm. A similarity threshold is set to automatically map highly relevant fields (such as satellite rainfall and ontology rainfall entities), while low-similarity fields are manually reviewed, and a field alignment mapping table is output.
[0045] The real-time quality control submodule, based on a semantic alignment mapping table, uses a physical rule verification engine and an isolated forest anomaly detection algorithm to identify data anomalies and triggers a drone swarm to use the YOLOv7 visual recognition algorithm for emergency supplementary data collection, generating fused quality inspection data.
[0046] Based on a semantic alignment mapping table, a physical rule engine is used to verify the rationality of the water level-flow relationship. Combined with the isolated forest algorithm, data anomalies (such as sudden drops in water level) are detected at edge nodes. After anomaly identification, a drone swarm is triggered to capture images of water surface markers through visual recognition algorithms. The images are then cross-validated with satellite imagery to generate fused quality inspection data.
[0047] The spatiotemporal cube submodule, based on fused quality inspection data, constructs a four-dimensional spatiotemporal data cube and uses Delta encoding technology to store only the changes in the flooding range, generating a spatiotemporal change matrix; Using fused quality inspection data as input, a four-dimensional data cube of longitude, latitude, elevation, and time is constructed. A change storage technique is applied to record only the differences in inundation extent between adjacent time points (such as area increase or decrease), ignoring low-fluctuation data, thus generating a highly compressed spatiotemporal change matrix.
[0048] The snapshot management submodule, based on the spatiotemporal change matrix and using the cumulative sum control chart mutation detection algorithm, automatically triggers a key state snapshot of the entire basin when the real-time water level exceeds the threshold, generating a key state snapshot set; Based on the water level sequence of the spatiotemporal change matrix, a cumulative sum control chart algorithm is used to monitor real-time water level deviations. When the water level exceeds the flood control plan threshold (such as the warning water level), a key state snapshot of the entire basin is automatically triggered, capturing the inundation data of all coordinate points at that moment and forming a historical backtracking snapshot set.
[0049] The knowledge graph construction submodule extracts "event-measure-effect" triples from historical flood control reports based on key state snapshot sets and a bidirectional long short-term memory-conditional random field entity extraction model. It then constructs a knowledge network using the graph database Neo4j and finally uses the Q-learning reinforcement learning algorithm to optimize the service composition strategy and generate a flood control knowledge base.
[0050] Extracting triples such as "rainstorm event - sandbag dispatch - loss reduction" from historical flood control reports linked to snapshots, and constructing a topological relationship network using a graph database (e.g., rainstorm events pointing to the drainage services that need to be invoked). Utilizing reinforcement learning algorithms to optimize the service combination order, a flood control knowledge base adapted to different flood disaster scenarios is formed.
[0051] The digital twin construction submodule, based on the flood control knowledge base, integrates LiDAR point cloud and UAV oblique photography data according to the CityGML 3.0 standard to construct a watershed digital model and generate a watershed digital twin; By integrating flood control knowledge base with lidar point cloud (terrain elevation) and UAV oblique photography (building texture), and fusing multi-source data according to 3D geographic information standards, a digital twin of the watershed is constructed to realize flood inundation simulation and risk assessment.
[0052] The data synchronization submodule, based on the watershed digital twin, injects SCADA sensor data streams in real time through the Apache Kafka stream processing platform, establishes a physical and virtual data interaction channel, and generates a virtual-physical synchronized data pipeline. SCADA sensor data (water level, flow rate) is injected into the digital twin in real time through a stream processing platform to dynamically calibrate the virtual model. Data consistency verification rules are set, and an alarm is triggered when the deviation between the real-time value and the model value exceeds the limit, generating a two-way synchronous data pipeline between the physical and virtual systems.
[0053] The strategy optimization submodule, based on the virtual-real synchronous data pipeline, adopts a near-end strategy optimization reinforcement learning algorithm, dynamically adjusts the data acquisition frequency with the flood peak prediction accuracy as the reward function, and outputs the flood diversion scheme to the gate control system, generating a twin optimization decision base.
[0054] Based on a virtual-real synchronized data pipeline, the frequency of UAV flights is dynamically optimized using flood peak prediction error as a feedback indicator. After outputting the flood diversion plan, the gate opening degree (e.g., opening percentage) is adjusted through the control system to complete the closed-loop decision-making from data acquisition to physical intervention.
[0055] The above are all preferred embodiments of this application, and are not intended to limit the scope of protection of this application. Therefore, all equivalent changes made in accordance with the structure, shape and principle of this application should be covered within the scope of protection of this application.
Claims
1. A method for constructing a multi-source data base for watershed flood control, characterized in that, Includes the following steps: S1: Based on multi-source heterogeneous data from meteorological satellites, hydrological stations, and geographic information, an entity relationship network is constructed using ontology modeling technology in the field of flood control. Cross-source field similarity is calculated using a semantic matching algorithm based on BERT, and conflicting data are weighted and fused using a long short-term memory network confidence evaluation model to generate a semantic fusion base. S2: Based on the real-time data stream of the semantic fusion base, a lightweight AI quality control module is deployed at the edge node of the watershed. Physical rule verification and isolated forest algorithm are used to detect outliers, triggering the drone cluster to perform emergency supplementary data collection through YOLOv7 visual recognition, generating enhanced quality inspection data; S3: Based on enhanced quality inspection data, a four-dimensional spatiotemporal data cube is constructed, and Delta encoding technology is used to store only the changes. When the water level changes abruptly, a snapshot is automatically triggered through the cumulative sum control chart algorithm, and the Hyperledger Fabric blockchain is used to store the data version hash and generate spatiotemporal version chain data. S4: Based on historical flood control records in spatiotemporal version chain data, a bidirectional long short-term memory-conditional random field entity extraction model is used to extract "event-measure-effect" triples. A flood control knowledge network is constructed through the graph database Neo4j, and the Q-learning algorithm is used to optimize the service composition strategy to generate a flood control decision knowledge graph. S5: Based on the flood control decision knowledge graph and real-time spatiotemporal version chain data, a watershed BIM model is constructed according to the CityGML 3.0 standard; SCADA sensor data is injected in real time through the Apache Kafka stream processing platform, and the acquisition frequency is dynamically adjusted by using a near-end strategy optimization reinforcement learning algorithm to output the flood diversion scheme to the gate control system and generate a twin optimization decision base.
2. The method for constructing a multi-source data base for watershed flood control according to claim 1, characterized in that: The S1-based generative semantic fusion platform includes the following steps: S101: Based on raw data from meteorological satellites, hydrological stations, and geographic information, a flood control ontology model is generated using ontology modeling technology in the field of flood control. S102: Based on the flood control ontology model, a bidirectional encoder based on a converter is used to represent the semantic matching algorithm, which converts satellite remote sensing metadata, hydrological station JSON messages, and CAD topographic map attribute tables into vectors, calculates cross-source field cosine similarity, and generates a semantic alignment mapping table. S103: Based on the semantic alignment mapping table, a long short-term memory network confidence evaluation model is used for conflicting data. The input data accuracy level and time freshness features are used to output dynamic weighting coefficients and generate a semantic fusion substrate.
3. The method for constructing a multi-source data base for watershed flood control according to claim 1, characterized in that: The generation of enhanced quality inspection data based on S2 includes the following steps: S201: Real-time data stream based on semantic fusion platform, using physical rule verification engine, verifies the rationality of water level-flow relationship through Saint-Venant hydrodynamic equations, and generates physically compliant data stream; S202: Based on physical compliance data flow, deploy a lightweight isolated forest anomaly detection algorithm on edge RTU devices to identify water level gauge failure anomalies caused by sudden rainstorms and generate anomaly label datasets; S203: Based on the anomaly marker dataset, trigger the UAV swarm to use the YOLOv7 visual recognition algorithm to capture images of water surface markers, cross-validate them with satellite synthetic aperture radar images, and generate quality inspection enhancement data.
4. The method for constructing a multi-source data base for watershed flood control according to claim 1, characterized in that: The generation of spatiotemporal version chain data based on S3 includes the following steps: S301: Based on enhanced quality inspection data, construct a four-dimensional spatiotemporal data cube, and use Delta encoding technology to store only the change in flooding range to generate a spatiotemporal change matrix; S302: Based on the spatiotemporal change matrix, using the cumulative sum control chart abrupt change detection algorithm, when the real-time water level exceeds the threshold, a full basin snapshot is automatically triggered to generate a set of key state snapshots; S303: Based on the key state snapshot set, the Hyperledger Fabric blockchain platform is used to generate Merkle tree hash values for data versions and store them on the chain to generate spatiotemporal version chain data.
5. The method for constructing a multi-source data base for watershed flood control according to claim 1, characterized in that: The generation of a flood control decision knowledge graph based on S4 includes the following steps: S401: Based on historical flood control reports in spatiotemporal version chain data, a bidirectional long short-term memory-conditional random field entity extraction model is used to identify the "breach location - material scheduling - loss reduction" triplet and generate a flood control entity triplet library. S402: Based on the flood control entity triplet library, a topology network is constructed through the Neo4j graph database to establish a relationship chain of "rainstorm event → service to be called" and generate a service relationship graph; S403: Based on the service relationship graph, the Q-learning reinforcement learning algorithm is used to optimize the service combination order and generate a flood control decision knowledge graph.
6. The method for constructing a multi-source data base for watershed flood control according to claim 1, characterized in that: The S5-based generative twin optimization decision platform includes the following steps: S501: Based on flood control decision-making knowledge graph and spatiotemporal version chain data, integrate LiDAR point cloud and oblique photography data according to CityGML 3.0 standard to generate a digital twin of the watershed; S502: Based on the watershed digital twin, SCADA sensor data streams are injected in real time through the Apache Kafka stream processing platform to generate a virtual-real synchronized data pipeline; S503: Based on a virtual-real synchronous data pipeline, it adopts a near-end strategy optimization reinforcement learning algorithm, uses the flood peak prediction accuracy as the reward function to dynamically adjust the UAV flight frequency, and generates an adaptive acquisition strategy. S504: A flood diversion scheme based on an adaptive acquisition strategy output, which uses a robot operating system to control the gate actuator to complete physical intervention and generate a twin-optimized decision base.
7. A multi-source data baseboard construction system for watershed flood control, characterized in that, It includes the following modules: a data fusion quality control module, which constructs an entity relationship network based on original multi-source data using flood control ontology modeling technology, performs cross-source field alignment using a transducer-based bidirectional encoder to represent semantic matching algorithm, and uses a long short-term memory network confidence evaluation model to handle data conflicts; A physical rule verification engine and an isolated forest anomaly detection algorithm are deployed at the edge nodes for real-time quality control. Finally, the YOLOv7 visual recognition algorithm drives the drone cluster to collect emergency supplementary data and generate fused quality inspection data. The data fusion quality control module includes an ontology modeling submodule, a semantic alignment submodule, and a real-time quality control submodule; The version knowledge construction module constructs a four-dimensional spatiotemporal data cube based on fused quality inspection data and uses Delta encoding technology to store changes. It triggers key state snapshots through cumulative sum and control graph mutation detection algorithms and uses the Hyperledger Fabric blockchain platform for hash storage. Based on historical flood control data, it extracts triples using a bidirectional long short-term memory-conditional random field entity extraction model and constructs a knowledge network through the graph database Neo4j. Finally, it uses Q-learning reinforcement learning algorithm to optimize service composition strategies and generate a flood control knowledge base. The version knowledge construction module includes a spatiotemporal cube submodule, a snapshot management submodule, and a knowledge graph construction submodule; The twin decision optimization module, based on a flood control knowledge base, integrates LiDAR and oblique photography data to construct a digital model according to the CityGML 3.0 standard; and injects SCADA sensor data streams in real time through the Apache Kafka streaming platform. The near-end strategy optimization reinforcement learning algorithm is used to dynamically adjust the acquisition frequency with flood peak prediction accuracy as the reward function; the flood diversion plan is output to the gate control system to perform physical intervention, and a twin optimization decision base is generated; The twin decision optimization module includes a twin construction submodule, a data synchronization submodule, and a strategy optimization submodule.
8. A multi-source data baseboard construction system for watershed flood control according to claim 7, characterized in that: The ontology modeling submodule generates a flood control ontology model based on original multi-source data and using ontology modeling technology in the field of flood control. The semantic alignment submodule, based on the flood control ontology model, uses a converter-based bidirectional encoder to represent the semantic matching algorithm, calculates the cosine similarity of cross-source fields of satellite remote sensing metadata and hydrological station JSON messages, achieves semantic-level alignment, and generates a semantic alignment mapping table. The real-time quality control submodule, based on a semantic alignment mapping table, uses a physical rule verification engine and an isolated forest anomaly detection algorithm to identify data anomalies and triggers the UAV cluster to use the YOLOv7 visual recognition algorithm for emergency supplementary data collection, generating fused quality inspection data.
9. A multi-source data baseboard construction system for watershed flood control according to claim 7, characterized in that: The spatiotemporal cube submodule, based on fused quality inspection data, constructs a four-dimensional spatiotemporal data cube and uses Delta encoding technology to store only the changes in the flooding range, generating a spatiotemporal change matrix; The snapshot management submodule, based on the spatiotemporal change matrix, adopts the cumulative sum and control chart mutation detection algorithm. When the real-time water level exceeds the threshold, it automatically triggers a snapshot of the key state of the entire basin and generates a set of key state snapshots. The knowledge graph construction submodule extracts "event-measure-effect" triples from historical flood control reports based on key state snapshot sets and a bidirectional long short-term memory-conditional random field entity extraction model. It then constructs a knowledge network using the graph database Neo4j and finally uses the Q-learning reinforcement learning algorithm to optimize the service composition strategy and generate a flood control knowledge base.
10. A multi-source data baseboard construction system for watershed flood control according to claim 7, characterized in that: The digital twin construction submodule, based on the flood control knowledge base, integrates LiDAR point cloud and UAV oblique photography data according to the CityGML 3.0 standard to construct a watershed digital model and generate a watershed digital twin; The data synchronization submodule, based on the watershed digital twin, injects SCADA sensor data streams in real time through the Apache Kafka stream processing platform, establishes a physical and virtual data interaction channel, and generates a virtual-physical synchronized data pipeline. The strategy optimization submodule, based on the virtual-real synchronous data pipeline, adopts a near-end strategy optimization reinforcement learning algorithm, dynamically adjusts the data acquisition frequency with the flood peak prediction accuracy as the reward function, and outputs the flood diversion scheme to the gate control system, generating a twin optimization decision base.