Reservoir real-time and offline data analysis method and system based on Spark engine
By adopting real-time data analysis method based on Spark engine in reservoir engineering, dynamically adjusting computing resources and model complexity, the problem that reservoir prediction models in the existing technology cannot respond at high speed is solved, and more efficient reservoir monitoring and management is achieved.
Patent Information
- Application Number
- CN202510614275.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-05-13
AI Technical Summary
The existing reservoir engineering monitoring and management methods have problems such as fixed design and inability to dynamically adjust the calculation frequency and model complexity, which leads to inaccurate prediction results in extreme weather conditions, affecting emergency management.
The real-time and offline data analysis method of reservoirs based on Spark engine is adopted, and by collecting monitoring data and meteorological data, distinguishing real-time data from offline data, using digital twin technology to build three-dimensional scenes, dynamically map real-time data, and performing in parallel through Spark engine, dynamically adjusting computing resources and model complexity.
Real-time data processing and efficient resource allocation are realized, real-time and accuracy of reservoir monitoring and management are improved, ensuring rapid response and scientific decision-making in extreme cases.
Smart Images

Figure CN120144968A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of water conservancy data processing, and particularly relates to a method and system for real-time and offline data analysis of reservoirs based on the Spark engine. Background Art
[0002] With the rapid development of Internet, big data, and Internet of Things technologies, the speed and scale of data generation have increased significantly, and the demand for data processing and analysis in various industries has been rising continuously. In this context, the analysis of real-time data and offline data has become a research hotspot. Real-time data is time-sensitive and can support decision-making in a timely manner. Offline data is historical static data, which is used for in-depth analysis and model construction. The combination of the two can improve the accuracy and response speed of analysis.
[0003] At the present stage, in the monitoring and management of reservoir projects, real-time data is usually used to set thresholds to trigger, and offline data is used to train static models for decision-making. This method has certain problems. On the one hand, current reservoir prediction models often adopt fixed designs and cannot dynamically adjust the calculation frequency and model complexity according to different environments and changes. In extreme weather conditions, many models cannot respond quickly and efficiently, resulting in inaccurate prediction results and affecting emergency management. On the other hand, traditional methods often adopt linear analysis and averaging processing methods. Model prediction relies on historical data for inertial reasoning. The computing power utilization differences of different analysis objects are small, and the overall response speed is relatively average. Especially during critical periods such as the flood season, efficient resource allocation and management cannot be ensured. Summary of the Invention
[0004] The purpose of the present invention is to provide a method and system for real-time and offline data analysis of reservoirs based on the Spark engine to solve the problems raised in the above background art.
[0005] To solve the above technical problems, the present invention provides a method for real-time and offline data analysis of reservoirs based on the Spark engine, including the following steps: S100. Collect monitoring data and meteorological data, and distinguish real-time data and offline data.
[0006] S200. Analyze the water storage relationship and precipitation relationship of the reservoir according to the offline data.
[0007] S300. Analyze the water level relationship between reservoirs, and comprehensively predict the development trend based on real-time data and meteorological data.
[0008] S400. Divide the observation area according to the development trend and allocate computing power, and execute in parallel using the Spark engine.
[0009] In S100, the monitoring data includes rainfall logs, water regime logs, and project operation logs. The rainfall logs include precipitation records at different times, and each precipitation record includes the precipitation amounts of each reservoir. The water regime logs include hydrological records at different times, and each hydrological record includes the water storage amounts of each reservoir and the water level values at different locations. The water level values are collected by monitoring devices installed near the reservoir dams. The project operation logs include GIS maps at different times. The GIS maps include the topographic structures, relative positions, and water flow directions of each reservoir. The meteorological data refers to the predicted precipitation amounts of each reservoir in a future period of time.
[0010] Monitoring data such as rainfall, water regime, and project operation exist in different storage media. The Spark engine realizes the unified parsing of monitoring data such as rainfall, water regime, and project operation into a two-dimensional tabular data structure for access and processing through distributed in-memory computing and data partitioning, greatly improving the data guarantee rate and playing an active role in communication guarantee during critical periods such as the flood season.
[0011] Take the precipitation amount, water storage amount, water level value, and GIS map at the latest time as real-time data, and the data at other times as offline data. Use digital twin technology to construct a three-dimensional scene based on the GIS map at the latest time, and dynamically map the real-time data into the three-dimensional scene.
[0012] Specifically, it includes the following steps: First, collect the precipitation amount, water storage amount, and water level value at the latest time in real-time, and obtain the corresponding GIS map information. These data are updated in real-time through sensors. Preprocess the collected real-time data, including data cleaning, format conversion, and standardization, to ensure the consistency and availability of the data.
[0013] Second, use digital twin technology to construct a three-dimensional scene based on the latest GIS map information. This process involves converting the two-dimensional data in the geographic information system into a three-dimensional model to form a visual virtual environment.
[0014] Then, dynamically map the processed precipitation amount, water storage amount, and water level value into the three-dimensional scene. Through programming interfaces or visualization tools, bind the data to the corresponding parts of the three-dimensional model, so that the water bodies and terrain in the scene can reflect the data changes in real-time.
[0015] Finally, display the dynamically mapped three-dimensional scene through a visualization platform, enabling users to intuitively observe the impact of real-time data on the environment and supporting decision-making and management.
[0016] Through the above steps, the effective combination of real-time data and the three-dimensional scene can be achieved, improving the monitoring and management capabilities of environmental changes.
[0017] S200 includes the following steps: S201. Analyze the upstream and downstream relationships of the reservoir based on the water flow direction in the 3D scene. All upstream reservoirs directly connected to reservoir Q1 are regarded as its influencing objects. Establish a water storage set and a precipitation set. Obtain the precipitation records when the precipitation of reservoir Q1 and all its influencing objects is zero, and put the times of these precipitation records into the water storage set. Obtain the precipitation records when the precipitation of reservoir Q1 is not zero and the precipitation of all its influencing objects is zero, and put the times of these precipitation records into the precipitation set.
[0018] S202. Obtain the hydrological records whose time intervals with any element in the water storage set are less than the time threshold w. Take the water storage volume of reservoir Q1 in each hydrological record as the dependent variable, and take the water storage volumes of all influencing objects of reservoir Q1 as independent variables and package them as samples. All samples are input into the polynomial regression model for training to obtain the water storage relationship expression XQ.
[0019] S203. Obtain the hydrological records whose time intervals with any element in the precipitation set are less than w. Substitute the water storage volumes of all influencing objects of reservoir Q1 in each hydrological record into the expression XQ to calculate e, and subtract e from the water storage volume of reservoir Q1 to obtain the water increment of the corresponding hydrological record. Each hydrological record matches the precipitation record with a time interval less than w. Take the water increment of each hydrological record as the dependent variable, and take the precipitation of reservoir Q1 in the matching precipitation record as the independent variable and package them as samples. All samples are input into the polynomial regression model for training to obtain the precipitation relationship expression. Analyze the influencing objects and relationship expressions of each reservoir.
[0020] S300 includes the following steps: S301. Take the downstream reservoir Q2 as the child node, and the upstream reservoir of Q2 as the parent node of Q2. Establish a tree - like relationship diagram according to the parent - child node relationship. Obtain the predicted precipitation of each reservoir at time t1 in the meteorological data. Substitute the predicted precipitation of the root node Qg into the precipitation relationship expression to calculate u, and add u to the current water storage volume of Qg to obtain the predicted water storage volume Lg of Qg.
[0021] There may be multiple tree - like relationships in the tree - like relationship diagram, that is, there are multiple root nodes at the same time. Each root node only has child nodes and no parent nodes. When the predicted precipitation of the root node is zero, the predicted water storage volume takes the current water storage volume.
[0022] S302. Substitute Lg into the water storage relationship expression of the child node Qz of Qg to calculate p, and then substitute the predicted precipitation of the child node Qz into the precipitation relationship expression to calculate the result f. Add f and p to obtain the predicted water storage volume of the child node Qz. Calculate the predicted water storage volumes of each reservoir progressively in the calculation direction from the parent node to the child node.
[0023] The predicted water storage volume of each child node is the sum of the results of the precipitation relationship expression and the water storage relationship expression, and the calculation is carried out in turn according to the parent - child node order.
[0024] S303. Analyze the reservoir to which the monitoring device belongs based on its location. Take the water level value of the monitoring device d1 in each hydrological record in the water regime log as the dependent variable, and the water storage volume of the reservoir to which the monitoring device d1 belongs as the independent variable and package them as samples. Input all the samples into the polynomial regression model for training to obtain the water level relationship expression of the monitoring device d1, and analyze the water level relationship expressions of each monitoring device. Substitute the predicted water storage volume into the water level relationship expressions of each monitoring device in the reservoir to calculate the predicted water level of each monitoring device.
[0025] S400 includes the following steps: S401. Analyze the positional distances between the monitoring device n and other monitoring devices in the same reservoir and calculate the average value to obtain the average distance H n , sum up the average distances of all monitoring devices in the same reservoir and then calculate the average value H ave , substitute it into the formula to calculate the trend index TR of each reservoir: ; In the formula, is the number of all monitoring devices in the reservoir, is the predicted water level of the th monitoring device, is the current water level of the th monitoring device, is the warning water level of the
[0026] Monitoring devices are usually installed near the reservoir dam. Under natural environmental conditions, the dam heights are inconsistent, so the warning water level values of different monitoring devices are different. The specific values are set in advance by the management personnel according to the actual situation, or are self-adjusted by the system according to the virtual dam mapped in real time in the three-dimensional scene.
[0027] The location of the monitoring device is set in advance by the management personnel. Usually, the more concentrated the monitoring devices are, the more important the location is. By calculating the average distance of the monitoring devices, the importance of the installation location to the reservoir can be analyzed. The shorter the average distance is, the more monitoring devices are set near that location, and the more important it is to the reservoir.
[0028] S402. Regard the area where the reservoir with a trend index greater than the threshold c in the three-dimensional scene is located as the observation area; obtain the remaining available computing power CP of the data center m and the trend index TR of the reservoir corresponding to each observation area v , and calculate the maximum allocated computing power CP of each observation area respectively max : ; In the formula, TR sumIt is the sum of the trend indices of the reservoirs corresponding to all observation areas.
[0029] Adjust the calculation frequency and the complexity of the model according to the trend index, which indirectly affects the use of resources. The reservoirs corresponding to the observation areas operate models that are more complex and require more computing resources, as well as more frequent tasks, while the reservoirs in non-observation areas operate simple models and low-frequency tasks.
[0030] In such a scenario, the reservoirs in the same batch processing job can be divided into different groups according to the trend index, and different complexity models and different numbers of runs are applied to each group. In Spark, for the data prediction tasks of each reservoir, different prediction models are selected and called according to its trend index, and at the same time, the calculation frequency is adjusted as much as possible. Each reservoir in the high-risk group needs to be processed multiple times in the same batch or scheduled at shorter intervals.
[0031] S403. Use the Spark engine to dynamically update the observation areas and allocate the maximum computing power, and monitor the total computing power used by the observation areas in real time; on the premise of not exceeding their respective maximum allocated computing power, call more complex prediction models and adjust the faster calculation frequency for each observation area in parallel, so as to improve the prediction frequency and accuracy of each observation area.
[0032] Calling more complex prediction models and adjusting the faster calculation frequency can pre-enact the dynamic evolution process of dam break in extreme situations of the reservoir, display the corresponding dam break flow process and the corresponding dam break parameter data. On the one hand, it can effectively guide the emergency rescue work of the reservoir in extreme situations, and on the other hand, it can inversely guide the preparation of the reservoir flood control plan and improve the accuracy level of the plan.
[0033] Based on the dynamic mechanism of the Spark engine, the list of observation areas can be monitored and updated in real time, and the resource supply can be automatically adjusted according to the preset maximum allocated computing power threshold. When the trend index of the reservoir in a certain observation area breaks through the threshold, the system will trigger an emergency expansion strategy: temporarily call the idle nodes of the cluster or enable elastic expansion of cloud services, increase the number of GPU nodes and increase the prediction frequency to ensure that the prediction tasks of the reservoirs in this observation area can obtain up to 100% of the maximum allocated computing power again. At the same time, dynamically reduce the resource quota of non-observation areas to maintain the global load balance.
[0034] The present invention also provides a real-time and offline data analysis system for reservoirs based on the Spark engine, including a data acquisition module, a fusion analysis module, a trend prediction module and a resource allocation module.
[0035] The data acquisition module is used to collect monitoring data and meteorological data, and distinguish real-time data and offline data.
[0036] The fusion analysis module is used to analyze the water storage relationship and precipitation relationship of the reservoir according to the offline data, and obtain the relationship expression.
[0037] The trend prediction module is used to analyze the water level relationship of the reservoir, and predict the development trend of the water storage volume and water level of each reservoir through real-time data and meteorological data.
[0038] The resource allocation module is used to divide the observation area and allocate computing power, and is executed by parallel scheduling using the Spark engine.
[0039] The data acquisition module includes a monitoring data acquisition unit and a meteorological data acquisition unit.
[0040] The monitoring data acquisition unit is used to collect rainfall logs, water regime logs, and engineering situation logs.
[0041] The rainfall log includes precipitation records at different times, and each precipitation record includes the precipitation of each reservoir.
[0042] The water regime log includes hydrological records at different times, and each hydrological record includes the water storage volume of each reservoir and the water level values at different locations.
[0043] The engineering situation log includes GIS maps at different times.
[0044] The meteorological data acquisition unit is used to collect the predicted precipitation of each reservoir within a certain period of time in the future.
[0045] The precipitation, water storage volume, water level value, and GIS map at the latest time are used as real-time data, and the data at other times are used as offline data. The digital twin technology is used to construct a three-dimensional scene based on the GIS map at the latest time, and the real-time data is dynamically mapped into the three-dimensional scene.
[0046] The fusion analysis module includes a time analysis unit and a relationship construction unit.
[0047] The time analysis unit is used to establish a water storage set and a precipitation set for each reservoir.
[0048] First, analyze the upstream and downstream relationships of the reservoirs according to the water flow direction in the three-dimensional scene, and all the upstream reservoirs directly connected to reservoir Q1 are used as its influencing objects.
[0049] Secondly, establish a water storage set and a precipitation set, obtain the precipitation records where the precipitation of reservoir Q1 and all its influencing objects is zero, and put the times of these precipitation records into the water storage set.
[0050] Finally, obtain the precipitation records where the precipitation of reservoir Q1 is not zero and the precipitation of all its influencing objects is zero, and put the times of these precipitation records into the precipitation set.
[0051] The relationship construction unit is used to analyze the relationship expressions of each reservoir.
[0052] First, obtain the hydrological records with a time interval less than the time threshold w from any element in the water storage set. Take the water storage volume of reservoir Q1 in each hydrological record as the dependent variable, and the water storage volumes of all influencing objects of reservoir Q1 as independent variables and package them as samples. All samples are input into the polynomial regression model for training to obtain the water storage relationship expression XQ.
[0053] Secondly, obtain the hydrological records with a time interval less than w from any element in the precipitation set. Substitute the water storage volumes of all influencing objects of reservoir Q1 in each hydrological record into the expression XQ to calculate e, and subtract e from the water storage volume of reservoir Q1 to obtain the water increment of the corresponding hydrological record.
[0054] Finally, match each hydrological record with precipitation records with a time interval less than w. Take the water increment of each hydrological record as the dependent variable, and the precipitation volume of reservoir Q1 in the matched precipitation record as the independent variable and package them as samples. All samples are input into the polynomial regression model for training to obtain the precipitation relationship expression.
[0055] Analyze the influencing objects and relationship expressions of each reservoir.
[0056] The trend prediction module includes a water storage prediction unit and a water level prediction unit.
[0057] The water storage prediction unit is used to calculate the predicted water storage volume of each reservoir.
[0058] First, take the downstream reservoir Q2 as a child node, and the upstream reservoir of Q2 as the parent node of Q2, and establish a tree-like relationship diagram according to the parent-child node relationship.
[0059] Secondly, obtain the predicted precipitation volumes of each reservoir at time t1 in the meteorological data. Substitute the predicted precipitation volume of the root node Qg into the precipitation relationship expression to calculate u, and add u to the current water storage volume of Qg to obtain the predicted water storage volume Lg of Qg.
[0060] Then, substitute Lg into the water storage relationship expression of the child node Qz of Qg to calculate p, and then substitute the predicted precipitation volume of the child node Qz into the precipitation relationship expression to calculate the result f, and add f to p to obtain the predicted water storage volume of the child node Qz.
[0061] Finally, calculate the predicted water storage volumes of each reservoir in a progressive manner in the calculation direction from the parent node to the child node.
[0062] The water level prediction unit is used to calculate the predicted water levels of each monitoring device.
[0063] First, analyze the belonging reservoir according to the location of the monitoring device. Take the water level value of the monitoring device d1 in each hydrological record in the water regime log as the dependent variable, and the water storage volume of the reservoir to which the monitoring device d1 belongs as the independent variable and package them as samples.
[0064] Secondly, all samples are input into the polynomial regression model for training to obtain the water level relationship expression of the monitoring device d1, and the water level relationship expressions of each monitoring device are analyzed.
[0065] Finally, the predicted water storage volumes are respectively substituted into the water level relationship expressions of each monitoring device in the reservoir to calculate the predicted water levels of each monitoring device.
[0066] The resource allocation module includes a computing power allocation unit and a parallel scheduling unit.
[0067] The computing power allocation unit is used to calculate the maximum allocated computing power for each observation area.
[0068] First, analyze the position distances between the monitoring device n and other monitoring devices in the same reservoir and calculate the average value to obtain the average distance H n , and calculate the average value H after summing the average distances of all monitoring devices in the same reservoir ave .
[0069] Secondly, according to the formula: Calculate the trend index of each reservoir. The area where the reservoir with a trend index greater than the threshold c in the three-dimensional scene is used as the observation area. Among them, is the total number of all monitoring devices in the reservoir, is the predicted water level of the th monitoring device, is the current water level of the th monitoring device, is the warning water level of the
[0070] Finally, obtain the remaining available computing power CP of the data center m and the trend index TR of the reservoir corresponding to each observation area v , according to the formula: Calculate the maximum allocated computing power for each observation area respectively. Among them, TR sum is the total sum of the trend indices of the reservoirs corresponding to all observation areas.
[0071] The parallel scheduling unit uses the Spark engine to dynamically update the observation areas and the maximum allocated computing power, and real-time monitors the total computing power utilized by the observation areas. On the premise of not exceeding their respective maximum allocated computing powers, it parallelly calls more complex prediction models and adjusts faster computing frequencies for each observation area, so as to improve the prediction frequency and accuracy of each observation area.
[0072] The Spark engine execution includes the following steps: Based on the preset trend index formula, the real-time trend index of each reservoir is calculated in parallel on the full-scale data with the help of Spark SQL. The system-level risk threshold is dynamically updated through sliding window statistics, and the high, medium, and low risk level intervals are adaptively delimited. After marking the results, a list of reservoirs with priority labels is generated, providing a decision-making basis for computing power scheduling.
[0073] Build a quantitative mapping model from risk levels to computing resources: Set high computing power density and high prediction frequency for high-risk reservoirs. For medium-risk tasks, a shared resource pool is adopted, and the number of CPU cores is allocated on demand. Low-risk batch jobs are merged and executed, and time slice rotation is used to reduce resource occupancy. Through the Spark configuration parameterization template, the linkage adjustment of resource specifications and risk levels is realized.
[0074] Intelligently slice the reservoir data set according to risk labels: The high-risk group uses fine-grained partitioning to maximize parallelism, and the medium- and low-risk groups aggregate partitions by geographical area to improve batch processing efficiency. Through a custom Spark scheduler plug-in, an independent resource pool is preempted for high-risk tasks in the YARN queue, and task priority weights are set to ensure that high-risk and time-sensitive tasks obtain computing resources first.
[0075] Integrate heterogeneous hardware resources at the Spark computing layer: The high-risk reservoir tasks are scheduled to the GPU cluster to accelerate the inference of LSTM / CNN models, the medium-risk tasks use CPU multi-threading to execute the gradient boosting tree model, and the low-risk batches call the pre-trained lightweight regression model. Through memory optimization strategies, the data Shuffle overhead is reduced, and the checkpoint mechanism is combined to ensure the fault tolerance of long-term prediction tasks.
[0076] The prediction results are written into the distributed storage in real time and visualized through the Dashboard. Establish a closed-loop feedback mechanism: Use Spark ML to analyze the execution effects of historical tasks, adjust the weight coefficients in the maximum allocated computing power formula by setting, and optimize the parameter sensitivity by combining the actual flood event data to improve the adaptability of the dynamic scheduling strategy.
[0077] Deploy a cluster health monitoring subsystem to track the resource utilization rate, queue backlog, and completion rate of tasks at each risk level in real time. Preset the fuse threshold to support manually increasing the computing power quota of specified reservoirs. Through A / B testing, compare the system indicators under different scheduling strategies, and continuously optimize the resource allocation algorithm to ensure the prediction accuracy and timeliness of key targets during flood peaks.
[0078] Compared with the prior art, the beneficial effects achieved by the present invention are: 1. Real-time data processing: Different from traditional offline monitoring systems, this solution can process monitoring data and meteorological data in real time based on the Spark engine, improving the timeliness of data analysis. This means that in critical periods such as the flood season, water situation changes can be responded to more quickly, enabling faster decision-making.
[0079] 2. Data fusion and unified parsing: Through distributed in-memory computing and data partitioning, this solution unifies the parsing of rainfall, water level, and project situation data into two-dimensional tabular data, eliminating the problem of data silos and improving the availability and guarantee rate of data. This efficient data integration makes decision-making more grounded.
[0080] 3. Dynamic resource allocation: This solution dynamically adjusts computing resources by setting a trend index to ensure that high-risk reservoirs receive a higher computing frequency and support from complex models, optimizing the resource utilization efficiency. At the same time, the system has a flexible resource expansion mechanism and can quickly increase computing resources when needed.
[0081] 4. Enhanced decision-making support: By presenting a dynamically mapped three-dimensional scene on a visualization platform, users can intuitively observe the impact of real-time data on the environment, supporting scientific decision-making and management. This intuitive data presentation reduces potential risks caused by information asymmetry. Description of the Drawings
[0082] The drawings are used to provide a further understanding of the present invention and form a part of the specification. They are used together with the embodiments of the present invention to explain the present invention and do not constitute a limitation to the present invention. In the drawings: Figure 1 is a schematic flowchart of the method in the present invention.
[0083] Figure 2 is a schematic diagram of the observation area planning of the method in the present invention.
[0084] Figure 3 is a schematic structural diagram of the system in the present invention.
[0085] Figure 4 is a schematic diagram of the result output of the system in the present invention. Detailed Embodiments
[0086] The present invention will be further described below with reference to the drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present invention and cannot be used to limit the protection scope of the present invention.
[0087] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0088] Please refer to Figure 1 , the present invention provides a method for real-time and offline data analysis of reservoirs based on the Spark engine, including the following steps: S100. Collect monitoring data and meteorological data, and distinguish real-time data and offline data. In S100, the monitoring data includes rainfall logs, water regime logs, and engineering situation logs. The rainfall logs include precipitation records at different times, and each precipitation record includes the precipitation amounts of each reservoir. The water regime logs include hydrological records at different times, and each hydrological record includes the water storage amounts of each reservoir and the water level values at different positions, and the water level values are collected by monitoring devices installed near the reservoir dams. The engineering situation logs include GIS maps at different times. The GIS maps include the topographic structures, relative positions, and water flow directions of each reservoir. The meteorological data refers to the predicted precipitation amounts of each reservoir in a future period of time.
[0089] Monitoring data such as rainfall, water regime, and engineering situation exist in different storage media. The Spark engine realizes the unified parsing of monitoring data such as rainfall, water regime, and engineering situation into a two-dimensional tabular data structure for access and processing through distributed memory computing and data partitioning, greatly improving the data guarantee rate and playing an active role in communication guarantee during critical periods such as flood seasons.
[0090] Take the precipitation amount, water storage amount, water level value, and GIS map at the latest time as real-time data, and the data at other times as offline data. Use digital twin technology to construct a three-dimensional scene based on the GIS map at the latest time, and dynamically map the real-time data into the three-dimensional scene.
[0091] Specifically, it includes the following steps: S101. Collect the precipitation amount, water storage amount, and water level value at the latest time in real time, and obtain the corresponding GIS map information. These data are updated in real time through sensors. Preprocess the collected real-time data, including data cleaning, format conversion, and standardization, to ensure the consistency and availability of the data.
[0092] S102. Use digital twin technology to construct a three-dimensional scene based on the latest GIS map information. This process involves converting the two-dimensional data in the geographic information system into a three-dimensional model to form a visual virtual environment.
[0093] S103. Dynamically map the processed precipitation, water storage volume, and water level values into a three-dimensional scene. Through a programming interface or visualization tool, bind the data to the corresponding parts of the three-dimensional model so that the water bodies and terrain in the scene can reflect data changes in real time.
[0094] S104. Display the dynamically mapped three-dimensional scene through a visualization platform, enabling users to intuitively observe the impact of real-time data on the environment and supporting decision-making and management.
[0095] Through the above steps, the effective combination of real-time data and the three-dimensional scene can be achieved, improving the monitoring and management capabilities of environmental changes.
[0096] S200. Analyze the water storage relationship and precipitation relationship of the reservoir based on offline data analysis. Specifically, S200 includes the following steps: S201. Analyze the upstream and downstream relationships of the reservoir according to the water flow direction in the three-dimensional scene. All upstream reservoirs directly connected to reservoir Q 1 are taken as its influencing objects. Establish a water storage set and a precipitation set, obtain precipitation records where the precipitation of reservoir Q 1 and all its influencing objects is zero, and put the time of these precipitation records into the water storage set. Obtain precipitation records where the precipitation of reservoir Q 1 is not zero and the precipitation of all its influencing objects is zero, and put the time of these precipitation records into the precipitation set.
[0097] S202. Obtain hydrological records with a time interval less than the time threshold w from any element in the water storage set. Take the water storage volume of reservoir Q 1 in each hydrological record as the dependent variable, and the water storage volumes of all influencing objects of reservoir Q 1 are taken as independent variables and packaged as samples. All samples are input into a polynomial regression model for training to obtain the water storage relationship expression X Q .
[0098] S203. Obtain hydrological records with a time interval less than w from any element in the precipitation set. Substitute the water storage volumes of all influencing objects of reservoir Q 1 in each hydrological record into the expression X Q to calculate e. Subtract e from the water storage volume of reservoir Q 1 to obtain the increased water volume corresponding to the hydrological record. Each hydrological record is matched with precipitation records with a time interval less than w. The increased water volume of each hydrological record is used as the dependent variable, and the precipitation of reservoir Q 1 in the matched precipitation record is used as the independent variable and packaged as samples. All samples are input into a polynomial regression model for training to obtain the precipitation relationship expression. Analyze the influencing objects and relationship expressions of each reservoir.
[0099] S300. Analyze the water level relationship between reservoirs, and predict the development trend by integrating real-time data and meteorological data. Specifically, S300 includes the following steps: S301. Take the downstream reservoir Q 2 as a child node, and the upstream reservoir of Q 2 as the parent node of Q 2 to establish a tree-like relationship diagram according to the parent-child node relationship. Obtain the predicted precipitation of each reservoir at time t 1 in the meteorological data. Substitute the predicted precipitation of the root node Q g into the precipitation relationship expression to calculate u, and add u to the current water storage of Q g to obtain the predicted water storage L g of Q g .
[0100] There may be multiple tree-like relationships in the tree-like relationship diagram, that is, there are multiple root nodes at the same time, and each root node only has child nodes and no parent nodes. When the predicted precipitation of the root node is zero, the predicted water storage takes the current water storage value.
[0101] S302. Substitute L g continuously into the water storage relationship expression of the child node Q g of Q z to calculate p, and then substitute the predicted precipitation of the child node Q z into the precipitation relationship expression to calculate the result f, and add f to p to obtain the predicted water storage of the child node Q z . Calculate the predicted water storage of each reservoir progressively in the calculation direction from the parent node to the child node.
[0102] The predicted water storage of each child node is the sum of the results of the precipitation relationship expression and the water storage relationship expression, and the calculation is carried out sequentially according to the parent-child node order.
[0103] S303. Analyze the affiliated reservoir according to the location of the monitoring device, and take the water level value of the monitoring device d 1 in each hydrological record in the water regime log as the dependent variable, and the water storage of the reservoir to which the monitoring device d 1 belongs as the independent variable and package them as samples. All samples are input into the polynomial regression model for training to obtain the water level relationship expression of the monitoring device d 1 . Analyze the water level relationship expressions of each monitoring device. Substitute the predicted water storage into the water level relationship expressions of each monitoring device in the reservoir to calculate the predicted water level of each monitoring device.
[0104] S400. Divide the observation area according to the development trend and allocate computing power, and execute in parallel using the Spark engine. Specifically, S400 includes the following steps: S401. Analyze the position distances between the monitoring device n and other monitoring devices in the same reservoir, calculate the average value to obtain the average distance H n , sum the average distances of all monitoring devices in the same reservoir and then calculate the average value H ave , thereby calculating the trend index TR of each reservoir: The calculation formula for the trend index is: ; In the formula, is the number of all monitoring devices in the reservoir, is the predicted water level of the th monitoring device, is the current water level of the th monitoring device, is the warning water level of the th monitoring device.
[0105] Monitoring devices are usually installed near the reservoir dam. Under natural environmental conditions, the dam heights are inconsistent, so the warning water level values of different monitoring devices are different. The specific values are set in advance by the management personnel according to the actual situation, or are self-adjusted by the system according to the virtual dam mapped in real time in the 3D scene.
[0106] The position of the monitoring device is set in advance by the management personnel. Usually, the more concentrated the monitoring devices are, the more important the location is. By calculating the average distance of the monitoring devices, the importance of the installation location to the reservoir can be analyzed. The shorter the average distance, the more monitoring devices are set near the location, and the more important it is to the reservoir.
[0107] S402. Take the area where the reservoir with a trend index greater than the threshold c in the 3D scene as the observation area. Obtain the remaining available computing power of the data center and the trend index of the reservoir corresponding to each observation area , substitute them into the formula to calculate the maximum allocated computing power of each observation area : ; In the formula, is the sum of the trend indexes of the reservoirs corresponding to all observation areas.
[0108] Adjust the calculation frequency and the complexity of the model according to the trend index, thereby indirectly affecting the use of resources. The reservoirs corresponding to the observation areas run more complex models and require more computing resources and more frequent tasks, while the reservoirs in non-observation areas run simple models and low-frequency tasks.
[0109] In such a scenario, the reservoirs in the same batch processing job can be divided into different groups according to the trend index, and each group applies models of different complexity and different running times. In Spark, for the data prediction task of each reservoir, different prediction models are selected and called according to its trend index, and the calculation frequency is adjusted as much as possible. In the high-risk group, each reservoir needs to be processed multiple times in the same batch or scheduled in a shorter interval.
[0110] S403, please refer to Figure 2 , the Spark engine is used to dynamically update the observation area and the maximum allocated computing power, and the total computing power used in the observation area is monitored in real time; without exceeding the respective maximum allocated computing power, more complex prediction models are called in parallel for each observation area and a faster calculation frequency is adjusted, thereby improving the prediction frequency and accuracy of each observation area.
[0111] Calling more complex prediction models and adjusting faster calculation frequencies can preview the dynamic evolution of dam breaches under extreme conditions, and display the corresponding dam breach flow process and corresponding dam breach parameter data. On the one hand, it can effectively guide emergency rescue work in reservoirs under extreme conditions, and on the other hand, it can reversely guide the preparation of reservoir flood control plans and improve the accuracy of the plans.
[0112] The dynamic mechanism based on the Spark engine can monitor and update the list of observation areas in real time, and automatically adjust the resource supply according to the preset maximum allocated computing power threshold. When the reservoir trend index of a certain observation area exceeds the threshold, the system will trigger an emergency expansion strategy: temporarily call idle nodes in the cluster or enable elastic expansion of cloud services, increase the number of GPU nodes and increase the prediction frequency to ensure that the prediction tasks of the reservoirs in the observation area can obtain up to 100% of the maximum allocated computing power again. At the same time, dynamically reduce the resource quota of non-observation areas to maintain global load balancing.
[0113] See also Figure 3 Based on the same invention, the present invention provides a real-time and offline reservoir data analysis system based on the Spark engine, including a data acquisition module, a fusion analysis module, a trend prediction module and a resource allocation module.
[0114] The data acquisition module is used to collect monitoring data and meteorological data, and distinguish between real-time data and offline data. The data acquisition module includes a monitoring data acquisition unit and a meteorological data acquisition unit.
[0115] The monitoring data collection unit is used to collect rainfall logs, water logs and work logs.
[0116] The rainfall log includes precipitation records at different times, and each precipitation record includes the precipitation amounts of each reservoir. The water regime log includes hydrological records at different times, and each hydrological record includes the water storage amounts of each reservoir and the water level values at different locations. The project situation log includes GIS maps at different times.
[0117] The meteorological data acquisition unit is used to collect the predicted precipitation amounts of each reservoir for a period of time in the future.
[0118] The precipitation amount, water storage amount, water level value, and GIS map at the latest time are used as real-time data, and the data at other times are used as offline data. The digital twin technology is used to construct a three-dimensional scene based on the GIS map at the latest time, and the real-time data is dynamically mapped into the three-dimensional scene.
[0119] The fusion analysis module is used to analyze the water storage relationship and precipitation relationship of the reservoir based on the offline data and obtain the relationship expression. The fusion analysis module includes a time analysis unit and a relationship construction unit.
[0120] The time analysis unit is used to establish a water storage set and a precipitation set for each reservoir.
[0121] First, analyze the upstream and downstream relationships of the reservoirs according to the water flow direction in the three-dimensional scene. All upstream reservoirs directly connected to reservoir Q 1 are used as its influencing objects.
[0122] Second, establish a water storage set and a precipitation set, and obtain precipitation records where the precipitation amounts of reservoir Q 1 and all its influencing objects are zero. Put the times of these precipitation records into the water storage set.
[0123] Finally, obtain precipitation records where the precipitation amount of reservoir Q 1 is not zero and the precipitation amounts of all its influencing objects are zero. Put the times of these precipitation records into the precipitation set.
[0124] The relationship construction unit is used to analyze the relationship expressions of each reservoir.
[0125] First, obtain hydrological records whose time intervals from any element in the water storage set are less than the time length threshold w. Take the water storage amount of reservoir Q in each hydrological record 1 as the dependent variable, and the water storage amounts of all influencing objects of reservoir Q 1 are used as independent variables and packed into samples. All samples are input into the polynomial regression model for training to obtain the water storage relationship expression X Q .
[0126] Second, obtain hydrological records whose time intervals from any element in the precipitation set are less than w. Substitute the water storage amounts of all influencing objects of reservoir Q in each hydrological record into the expression X 1 Q Calculate e from it, and subtract e from the water storage volume of reservoir Q 1 to obtain the increased water volume corresponding to the hydrological record.
[0127] Finally, each hydrological record is matched with precipitation records with a time interval less than w. The increased water volume of each hydrological record is used as the dependent variable, and the precipitation volume of reservoir Q 1 in the matched precipitation records is used as the independent variable and packaged as a sample. All samples are input into the polynomial regression model for training to obtain the precipitation relationship expression. Analyze the influencing objects and relationship expressions of each reservoir.
[0128] The trend prediction module is used to analyze the water level relationship of the reservoir, and predict the development trend of the water storage volume and water level of each reservoir through real-time data and meteorological data. The trend prediction module includes a water storage prediction unit and a water level prediction unit.
[0129] The water storage prediction unit is used to calculate the predicted water storage volume of each reservoir.
[0130] First, take the downstream reservoir Q 2 as the child node, and the upstream reservoir of Q 2 as the parent node of Q 2 to establish a tree-like relationship diagram according to the parent-child node relationship.
[0131] Secondly, obtain the predicted precipitation volume of each reservoir at time t 1 in the meteorological data. Substitute the predicted precipitation volume of the root node Q g into the precipitation relationship expression to calculate u, and add u to the current water storage volume of Q g to obtain the predicted water storage volume L g of Q g .
[0132] Then, continue to substitute L g into the water storage relationship expression of the child node Q g of Q z to calculate p, and then substitute the predicted precipitation volume of the child node Q z into the precipitation relationship expression to calculate the result f, and add f to p to obtain the predicted water storage volume of the child node Q z .
[0133] Finally, calculate the predicted water storage volume of each reservoir progressively in the calculation direction from the parent node to the child node.
[0134] The water level prediction unit is used to calculate the predicted water level of each monitoring device.
[0135] First, analyze the affiliated reservoir according to the location of the monitoring device. Take the water level value of the monitoring device d 1 in each hydrological record in the water regime log as the dependent variable, and the water storage volume of the reservoir to which the monitoring device d 1 belongs as the independent variable and package it as a sample.
[0136] Secondly, all samples are input into the polynomial regression model for training to obtain the water level relationship expression of the monitoring device d 1 and analyze the water level relationship expressions of each monitoring device.
[0137] Finally, substitute the predicted water storage volume into the water level relationship expressions of each monitoring device in the reservoir to calculate the predicted water level of each monitoring device.
[0138] The resource allocation module is used to divide the observation area and allocate computing power, and is executed in parallel scheduling by the Spark engine. The resource allocation module includes a computing power allocation unit and a parallel scheduling unit.
[0139] The computing power allocation unit is used to calculate the maximum allocated computing power for each observation area.
[0140] First, analyze the position distances between the monitoring device n and other monitoring devices in the same reservoir and calculate the average value to obtain the average distance , and calculate the average value after summing the average distances of all monitoring devices in the same reservoir .
[0141] Secondly, according to the formula: Calculate the trend index of each reservoir. The area where the reservoir with a trend index greater than the threshold c in the three-dimensional scene is used as the observation area. Among them, is the total number of all monitoring devices in the reservoir, is the predicted water level of the th monitoring device, is the current water level of the th monitoring device, is the warning water level of the th monitoring device.
[0142] Finally, obtain the remaining available computing power of the data center and the trend index of the reservoir corresponding to each observation area , and according to the formula: Calculate the maximum allocated computing power for each observation area respectively. Among them, is the total sum of the trend indexes of the reservoirs corresponding to all observation areas.
[0143] The parallel scheduling unit uses the Spark engine to dynamically update the observation area and the maximum allocated computing power, and real-time monitors the total computing power utilized by the observation area. On the premise of not exceeding their respective maximum allocated computing powers, it parallelly calls more complex prediction models and adjusts faster computing frequencies for each observation area, so as to improve the prediction frequency and accuracy of each observation area.
[0144] Please refer to Figure 4 , the execution of the Spark engine includes the following steps: Based on a preset trend index formula, the real-time trend index of each reservoir is calculated in parallel on the full volume of data with the help of Spark SQL. The system-level risk threshold is dynamically updated through sliding window statistics, and the high, medium, and low risk level intervals are adaptively delimited. After the results are marked, a list of reservoirs with priority labels is generated, providing a decision-making basis for computing power scheduling.
[0145] Construct a quantitative mapping model from risk levels to computing resources: Set high computing power density (such as each task exclusive GPU node) and high prediction frequency (minute level) for high-risk reservoirs. Shared resource pools are used for medium-risk tasks, and the number of CPU cores is allocated on demand. Low-risk batch jobs are merged and executed, and time slice rotation is used to reduce resource occupancy. Through the Spark configuration parameterization template, the linkage adjustment of resource specifications (such as the number of Executors, memory quota) and risk levels is realized.
[0146] Intelligently slice the reservoir data set according to risk labels: The high-risk group uses fine-grained partitioning (such as each partition with a single reservoir's data) to maximize parallelism, and the medium- and low-risk groups aggregate partitions by geographical area to improve batch processing efficiency. Through a custom Spark scheduler plugin, a separate resource pool is preempted for high-risk tasks in the YARN queue, and task priority weights are set (such as multi-pool configuration of Fair Scheduler), ensuring that high-risk and time-sensitive tasks can obtain computing resources first.
[0147] Integrate heterogeneous hardware resources at the Spark computing layer: The tasks of high-risk reservoirs are scheduled to the GPU cluster to accelerate the inference of LSTM / CNN models, medium-risk tasks use CPU multi-threading to execute gradient boosting tree models, and low-risk batches call pre-trained lightweight regression models. Through memory optimization strategies (such as off-heap memory allocation, RDD persistence strategy), the data Shuffle overhead is reduced, and the checkpoint mechanism is combined to ensure the fault tolerance of long-term prediction tasks.
[0148] The prediction results are written into distributed storage (such as HBase + HDFS) in real time and presented visually through the Dashboard. Establish a closed-loop feedback mechanism: Use Spark ML to analyze the execution effects of historical tasks (prediction error, response latency), adjust the weight coefficients in the maximum allocated computing power formula by setting, and optimize the parameter sensitivity in combination with actual flood event data to improve the adaptability of the dynamic scheduling strategy.
[0149] Deploy a cluster health monitoring subsystem to track the resource utilization rate, queue backlog, and completion rate of tasks at each risk level in real time. Preset a fuse threshold (such as an alarm is automatically triggered when the high-risk task delay exceeds 5 minutes), and support manually increasing the computing power quota of specified reservoirs. Through A / B testing, compare the system indicators under different scheduling strategies, and continuously optimize the resource allocation algorithm to ensure the prediction accuracy and timeliness of key targets during flood peaks.
[0150] Example 1: Assume reservoir Q 8 There are a total of 4 monitoring devices, A1, A2, A3, and A4, inside. Their average distances are: 26m, 24m, 36m, and 18m respectively. Their water level information is as follows: A1: Predicted water level: 180cm; Current water level: 150cm; Warning water level: 250cm; A2: Predicted water level: 160cm; Current water level: 140cm; Warning water level: 210cm; A3: Predicted water level: 150cm; Current water level: 140cm; Warning water level: 180cm; A4: Predicted water level: 170cm; Current water level: 160cm; Warning water level: 240cm; Substitute into the formula to calculate reservoir Q 8 's trend index: ; Then the trend index of reservoir Q 8 is 0.32.
[0151] In summary, this technical solution realizes precise computing power flow allocation with risk as the weight through a multi-level algorithm, meeting the core requirements of key target key guarantee in intelligent water conservancy. It significantly improves the real-time performance, accuracy, and emergency response ability of reservoir monitoring and management, and has obvious advantages compared with the existing technology.
[0152] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device.
[0153] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
[0154] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.
Claims
1. A real-time and offline reservoir data analysis method based on Spark engine, characterized by: The method comprises the following steps: S100, collecting monitoring data and meteorological data, and distinguishing between real-time data and offline data; S200, analyzing the relationship between water storage and precipitation of the reservoir according to the offline data; S300, analyze the water level relationship between reservoirs, and predict the development trend by integrating real-time data and meteorological data; S400, divide the observation area and allocate computing power according to the development trend, and use the Spark engine for parallel execution.
2. The method for real-time and offline reservoir data analysis based on Spark engine according to claim 1, characterized in that: In S100, the monitoring data includes rainfall log, water log and work log; The rainfall log includes precipitation records at different times, and each precipitation record includes the precipitation amount of each reservoir; The water regime log includes hydrological records at different times. Each hydrological record includes the water storage capacity of each reservoir and the water level values at different locations. The water level values are collected by monitoring equipment installed near the reservoir dam; The work log includes GIS maps at different times; GIS maps include the topographic structure, relative position and water flow direction of each reservoir; meteorological data refers to the predicted precipitation of each reservoir in the future period; The latest precipitation, water storage, water level and GIS map are used as real-time data, and the data at other times are used as offline data. Digital twin technology is used to build a three-dimensional scene based on the latest GIS map, and the real-time data is dynamically mapped to the three-dimensional scene.
3. A method for real-time and offline reservoir data analysis based on Spark engine according to claim 2, characterized in that: S200 includes the following steps: S201, analyzing the upstream and downstream relationship of the reservoir according to the water flow direction in the three-dimensional scene, and taking all upstream reservoirs directly connected to the reservoir Q1 as its impact objects; S202, establishing a water storage set and a precipitation set, obtaining precipitation records in which the precipitation of reservoir Q1 and all its affected objects is zero, and putting the time of these precipitation records into the water storage set; S203, obtaining precipitation records in which the precipitation of reservoir Q1 is not zero and the precipitation of all its affected objects is zero, and putting the time of these precipitation records into the precipitation set; S204, analyzing the water storage set and the precipitation set, establishing a water storage relationship expression and a precipitation relationship expression for reservoir Q1, and analyzing the influencing objects and relationship expressions of other reservoirs.
4. The method for real-time and offline reservoir data analysis based on Spark engine according to claim 3 is characterized in that: In S204, establishing a relational expression includes the following steps: S2041. Obtain hydrological records whose time interval with any element in the water storage set is less than the time threshold w; take the water storage capacity of reservoir Q1 in each hydrological record as the dependent variable, and the water storage capacity of all influencing objects of reservoir Q1 as independent variables and package them into samples; input all samples into the polynomial regression model for training, and obtain the water storage relationship expression X Q ; S2042. Obtain hydrological records with a time interval less than w from any element in the precipitation set, and substitute the water storage capacity of all affected objects of reservoir Q1 in each hydrological record into the expression X Q Calculate e in the calculation, subtract e from the water storage of reservoir Q1 to get the water increase corresponding to the hydrological record; S2043. Each hydrological record is matched with precipitation records with a time interval less than w. The water increase of each hydrological record is used as the dependent variable. The precipitation of reservoir Q1 in the matched precipitation record is used as the independent variable and packaged as samples. All samples are input into the polynomial regression model for training to obtain the precipitation relationship expression.
5. The method for real-time and offline reservoir data analysis based on Spark engine according to claim 3, characterized in that: S300 includes the following steps: S301, taking the downstream reservoir Q2 as a child node and the upstream reservoir of Q2 as the parent node of Q2, and establishing a tree relationship diagram according to the parent-child node relationship; S302, obtain the predicted precipitation of each reservoir at time t1 in the meteorological data, and set the root node Q g Substitute the predicted precipitation into the precipitation relationship expression to calculate u, Q g The current water storage plus u gives Q g The predicted water storage capacity L g ; S303, L g Continue to substitute Q g The child node Q z Calculate p from the water storage relationship expression, and then add the child node Q z Substitute the predicted precipitation into the precipitation relationship expression to calculate the result f; S304, add f to p to obtain child node Q z The predicted water storage capacity of each reservoir is calculated in turn according to the calculation direction from the parent node to the child node. S305. Analyze the water level relationship expression of each monitoring device according to the water situation log, substitute the predicted water storage capacity into the water level relationship expression of each monitoring device in the reservoir, and calculate the predicted water level of each monitoring device.
6. The method for real-time and offline reservoir data analysis based on Spark engine according to claim 5, characterized in that: The establishment of water level relationship expression includes the following steps: Analyze the reservoir to which each monitoring device belongs, take the water level value of monitoring device d1 in each hydrological record in the water regime log as the dependent variable, and the water storage capacity of the reservoir to which monitoring device d1 belongs as the independent variable and package them as samples; all samples are input into the polynomial regression model for training to obtain the water level relationship expression of monitoring device d1; and so on, analyze the water level relationship expression of each monitoring device.
7. The method for real-time and offline reservoir data analysis based on Spark engine according to claim 5, characterized in that: S400 includes the following steps: S401, analyzing the distance between monitoring device n and other monitoring devices in the same reservoir and calculating the average value to obtain the average distance H n , and calculate the average value H by summing the average distances of all monitoring devices in the reservoir ave , thereby calculating the trend index TR of each reservoir; S402: The area where the reservoir with a trend index greater than a threshold value c in the three-dimensional scene is located is used as the observation area; and the remaining available computing power CP of the data center is obtained. m And the trend index TR of each observation area corresponding to the reservoir v , calculate the maximum allocated computing power CP of each observation area respectively max ; S403, using the Spark engine to dynamically update the observation area and the maximum allocated computing power, and monitor the total amount of computing power used in the observation area in real time; Without exceeding the maximum allocated computing power of each area, more complex prediction models are called in parallel for each observation area and faster calculation frequencies are adjusted to improve the prediction frequency and accuracy of each observation area.
8. The method for real-time and offline reservoir data analysis based on Spark engine according to claim 7, characterized in that: In S401, the trend index calculation formula is: ; In the formula, is the number of all monitoring devices in the reservoir, For the The predicted water level of each monitoring device, For the The current water level of each monitoring device, For the The warning water level of a monitoring device.
9. The method for real-time and offline reservoir data analysis based on Spark engine according to claim 7, characterized in that: In S402, the maximum allocated computing power calculation formula is: ; In the formula, TR sum It is the sum of trend indices of all corresponding reservoirs in the observation area.
10. A real-time and offline reservoir data analysis system based on Spark engine, characterized by: It includes data acquisition module, fusion analysis module, trend prediction module and resource allocation module; The data acquisition module is used to collect monitoring data and meteorological data, and distinguish between real-time data and offline data; The fusion analysis module is used to analyze the water storage relationship and precipitation relationship of the reservoir based on offline data and obtain the relationship expression; The trend prediction module is used to analyze the water level relationship of the reservoir and predict the development trend of the water storage capacity and water level of each reservoir through real-time data and meteorological data; The resource allocation module is used to divide the observation area and allocate computing power, and uses the Spark engine for parallel scheduling and execution.
Citation Information
Patent Citations
Flink-based hydrological sensor data analysis system and construction method thereof
CN114817361A
Water resource optimization scheduling management method and system based on digital twinning
CN118536773A
Hydro-meteorological time sequence-based runoff simulation method and system
WO2021217776A1
Tailing pond risk monitoring and early-warning system based on internet of things
WO2023061039A1
KR20240039858A