A Method and System for Real-time and Offline Data Analysis of Reservoirs Based on the Spark Engine
Through the combination of Spark engine and digital twin technology, the reservoir data analysis method is dynamically adjusted, and the accuracy and resource allocation problems of the reservoir prediction model in extreme weather are solved, rapid response and efficient resource utilization are achieved, and decision-making support capabilities for reservoir management are improved.
Patent Information
- Application Number
- CN202510614275.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-05-13
AI Technical Summary
The existing reservoir prediction model cannot dynamically adjust the calculation frequency and model complexity according to environmental changes, resulting in inaccurate prediction results in extreme weather conditions, inefficient response, and improper resource allocation, affecting emergency management.
The real-time and offline data analysis method of reservoirs based on Spark engine is adopted, and the monitoring data is uniformly analyzed and monitored through distributed memory computing and data partitioning, combined with digital twin technology to build three-dimensional scenarios, dynamically adjust computing resources, realize the effective combination of real-time data and environmental changes, and execute complex models in parallel to improve prediction frequency and accuracy.
It achieves rapid response and accurate prediction in extreme weather conditions, optimizes resource utilization efficiency, and improves decision-making support capabilities and emergency management levels.
Smart Images

Figure CN120144968B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of water conservancy data processing, and particularly relates to a method and system for real-time and offline data analysis of reservoirs based on the Spark engine. Background Art
[0002] With the rapid development of Internet, big data, and Internet of Things technologies, the speed and scale of data generation have increased significantly, and the demand for data processing and analysis in various industries has been rising continuously. Against this background, the analysis of real-time data and offline data has become a research hotspot. Real-time data is time-sensitive and can support decision-making in a timely manner. Offline data is historical static data and is used for in-depth analysis and model construction. The combination of the two can improve the accuracy and response speed of analysis.
[0003] At present, in the monitoring and management of reservoir projects, real-time data is usually used to set thresholds to trigger, and offline data is used to train static models for decision-making. This method has certain problems. On the one hand, current reservoir prediction models often adopt fixed designs and cannot dynamically adjust the calculation frequency and model complexity according to different environments and changes. In extreme weather conditions, many models cannot respond quickly and efficiently, resulting in inaccurate prediction results and affecting emergency management. On the other hand, traditional methods often adopt linear analysis and averaging processing methods. Model prediction relies on historical data for inertial reasoning. The computing power utilization differences of different analysis objects are small, and the overall response speed is relatively average. Especially during critical periods such as the flood season, efficient resource allocation and management cannot be ensured. Summary of the Invention
[0004] The purpose of the present invention is to provide a method and system for real-time and offline data analysis of reservoirs based on the Spark engine to solve the problems raised in the above background art.
[0005] To solve the above technical problems, the present invention provides a method for real-time and offline data analysis of reservoirs based on the Spark engine, including the following steps:
[0006] S100. Collect monitoring data and meteorological data, and distinguish real-time data and offline data.
[0007] S200. Analyze the water storage relationship and precipitation relationship of the reservoir according to the offline data.
[0008] S300. Analyze the water level relationship between reservoirs, and comprehensively predict the development trend based on real-time data and meteorological data.
[0009] S400. Divide the observation area according to the development trend and allocate computing power, and execute in parallel using the Spark engine.
[0010] In S100, the monitoring data includes rainfall logs, water regime logs, and project operation logs. The rainfall logs include precipitation records at different times, and each precipitation record includes the precipitation amounts of each reservoir. The water regime logs include hydrological records at different times, and each hydrological record includes the water storage amounts of each reservoir and the water level values at different locations. The water level values are collected by monitoring devices installed near the reservoir dams. The project operation logs include GIS maps at different times. The GIS maps include the topographic structures, relative positions, and water flow directions of each reservoir. The meteorological data refers to the predicted precipitation amounts of each reservoir in a future period of time.
[0011] Monitoring data such as rainfall, water regime, and project operation exist on different storage media. The Spark engine realizes the unified parsing of monitoring data such as rainfall, water regime, and project operation into a two-dimensional tabular data structure through distributed in-memory computing and data partitioning for access and processing, greatly improving the data guarantee rate and playing an active role in communication guarantee during critical periods such as the flood season.
[0012] Take the precipitation amount, water storage amount, water level value, and GIS map at the latest time as real-time data, and the data at other times as offline data. Use digital twin technology to construct a three-dimensional scene based on the GIS map at the latest time, and dynamically map the real-time data into the three-dimensional scene.
[0013] Specifically, it includes the following steps:
[0014] First, collect the precipitation amount, water storage amount, and water level value at the latest time in real-time, and obtain the corresponding GIS map information. These data are updated in real-time through sensors. Preprocess the collected real-time data, including data cleaning, format conversion, and standardization, to ensure the consistency and availability of the data.
[0015] Second, use digital twin technology to construct a three-dimensional scene based on the latest GIS map information. This process involves converting the two-dimensional data in the geographic information system into a three-dimensional model to form a visual virtual environment.
[0016] Then, dynamically map the processed precipitation amount, water storage amount, and water level value into the three-dimensional scene. Through programming interfaces or visualization tools, bind the data to the corresponding parts of the three-dimensional model, so that the water bodies and terrain in the scene can reflect the data changes in real-time.
[0017] Finally, display the dynamically mapped three-dimensional scene through a visualization platform, enabling users to intuitively observe the impact of real-time data on the environment and supporting decision-making and management.
[0018] Through the above steps, the effective combination of real-time data and the three-dimensional scene can be achieved, improving the monitoring and management capabilities of environmental changes.
[0019] S200 includes the following steps:
[0020] S201. Analyze the upstream and downstream relationships of the reservoir based on the water flow direction in the three-dimensional scene. All upstream reservoirs directly connected to reservoir Q1 are regarded as its influencing objects. Establish a water storage set and a precipitation set, obtain precipitation records where the precipitation of reservoir Q1 and all its influencing objects is zero, and put the times of these precipitation records into the water storage set. Obtain precipitation records where the precipitation of reservoir Q1 is not zero and the precipitation of all its influencing objects is zero, and put the times of these precipitation records into the precipitation set.
[0021] S202. Obtain hydrological records whose time intervals from any element in the water storage set are less than the time threshold w. Take the water storage volume of reservoir Q1 in each hydrological record as the dependent variable, and the water storage volumes of all influencing objects of reservoir Q1 are taken as independent variables and packed as samples. All samples are input into the polynomial regression model for training to obtain the water storage relationship expression XQ.
[0022] S203. Obtain hydrological records whose time intervals from any element in the precipitation set are less than w. Substitute the water storage volumes of all influencing objects of reservoir Q1 in each hydrological record into the expression XQ to calculate e, and subtract e from the water storage volume of reservoir Q1 to obtain the water increase volume of the corresponding hydrological record. Each hydrological record matches a precipitation record with a time interval less than w. Take the water increase volume of each hydrological record as the dependent variable, and the precipitation of reservoir Q1 in the matched precipitation record as the independent variable and pack as samples. All samples are input into the polynomial regression model for training to obtain the precipitation relationship expression. Analyze the influencing objects and relationship expressions of each reservoir.
[0023] S300 includes the following steps:
[0024] S301. Take the downstream reservoir Q2 as the child node, and the upstream reservoir of Q2 as the parent node of Q2, and establish a tree-like relationship graph according to the parent-child node relationship. Obtain the predicted precipitation of each reservoir at time t1 in the meteorological data. Substitute the predicted precipitation of the root node Qg into the precipitation relationship expression to calculate u, and add u to the current water storage volume of Qg to obtain the predicted water storage volume Lg of Qg.
[0025] There may be multiple tree-like relationships in the tree-like relationship graph, that is, there are multiple root nodes at the same time, and each root node only has child nodes and no parent nodes. When the predicted precipitation of the root node is zero, the predicted water storage volume takes the current water storage volume.
[0026] S302. Substitute Lg into the water storage relationship expression of the child node Qz of Qg to calculate p, and then substitute the predicted precipitation of the child node Qz into the precipitation relationship expression to calculate the result f. Add f to p to obtain the predicted water storage volume of the child node Qz. Calculate the predicted water storage volumes of each reservoir in a progressive manner according to the calculation direction from the parent node to the child node.
[0027] The predicted water storage volume of each child node is the sum of the results of the precipitation relationship expression and the water storage relationship expression, and the calculations are performed sequentially in the order of parent and child nodes.
[0028] S303. Analyze the reservoir to which the monitoring device belongs based on its location. Take the water level value of the monitoring device d1 in each hydrological record in the water regime log as the dependent variable, and the water storage volume of the reservoir to which the monitoring device d1 belongs as the independent variable and package them as samples. All samples are input into the polynomial regression model for training to obtain the water level relationship expression of the monitoring device d1, and analyze the water level relationship expressions of each monitoring device. Substitute the predicted water storage volume into the water level relationship expressions of each monitoring device in the reservoir to calculate the predicted water level of each monitoring device.
[0029] S400 includes the following steps:
[0030] S401. Analyze the location distances between the monitoring device n and other monitoring devices in the same reservoir and calculate the average value to obtain the average distance H n , sum the average distances of all monitoring devices in the same reservoir and then calculate the average value H ave , substitute it into the formula to calculate the trend index TR of each reservoir:
[0031] ;
[0032] In the formula, is the number of all monitoring devices in the reservoir, is the predicted water level of the th monitoring device, is the current water level of the th monitoring device, is the warning water level of the
[0033] Monitoring devices are usually installed near the reservoir dam. Under natural environmental conditions, the dam heights are inconsistent, so the warning water level values of different monitoring devices are different. The specific values are set in advance by the management personnel according to the actual situation, or are self-adjusted by the system according to the virtual dam mapped in real time in the three-dimensional scene.
[0034] The location of the monitoring device is set in advance by the management personnel. Usually, the more concentrated the monitoring devices are, the more important the location is. By calculating the average distance of the monitoring devices, the importance of the installation location to the reservoir can be analyzed. The shorter the average distance, the more monitoring devices are set near that location, and the more important it is to the reservoir.
[0035] S402. Take the area where the reservoir with a trend index greater than the threshold c in the three-dimensional scene as the observation area; obtain the remaining available computing power CP of the data center m and the trend index TR of the reservoir corresponding to each observation area v, calculate the maximum allocated computing power CP of each observation area respectively max :
[0036] ;
[0037] In the formula, TR sum It is the sum of trend indices of all corresponding reservoirs in the observation area.
[0038] The frequency of calculations and the complexity of the model are adjusted according to the trend index, which indirectly affects the use of resources. Reservoirs in the observation area run more complex models that require more computing resources and more frequent tasks, while reservoirs in the non-observation area run simple models and low-frequency tasks.
[0039] In such a scenario, the reservoirs in the same batch processing job can be divided into different groups according to the trend index, and each group applies models of different complexity and different running times. In Spark, for the data prediction task of each reservoir, different prediction models are selected and called according to its trend index, and the calculation frequency is adjusted as much as possible. In the high-risk group, each reservoir needs to be processed multiple times in the same batch or scheduled in a shorter interval.
[0040] S403. Use the Spark engine to dynamically update the observation area and the maximum allocated computing power, and monitor the total computing power used by the observation area in real time; without exceeding the respective maximum allocated computing power, call more complex prediction models in parallel for each observation area and adjust the faster calculation frequency, thereby improving the prediction frequency and accuracy of each observation area.
[0041] Calling more complex prediction models and adjusting faster calculation frequencies can preview the dynamic evolution of dam breaches under extreme conditions, and display the corresponding dam breach flow process and corresponding dam breach parameter data. On the one hand, it can effectively guide emergency rescue work in reservoirs under extreme conditions, and on the other hand, it can reversely guide the preparation of reservoir flood control plans and improve the accuracy of the plans.
[0042] The dynamic mechanism based on the Spark engine can monitor and update the list of observation areas in real time, and automatically adjust the resource supply according to the preset maximum allocated computing power threshold. When the reservoir trend index of a certain observation area exceeds the threshold, the system will trigger an emergency expansion strategy: temporarily call idle nodes in the cluster or enable elastic expansion of cloud services, increase the number of GPU nodes and increase the prediction frequency to ensure that the prediction tasks of the reservoirs in the observation area can obtain up to 100% of the maximum allocated computing power again. At the same time, dynamically reduce the resource quota of non-observation areas to maintain global load balancing.
[0043] The present invention also provides a real-time and offline reservoir data analysis system based on the Spark engine, comprising a data acquisition module, a fusion analysis module, a trend prediction module and a resource allocation module.
[0044] The data acquisition module is used to collect monitoring data and meteorological data, and distinguish between real-time data and offline data.
[0045] The fusion analysis module is used to analyze the water storage relationship and precipitation relationship of the reservoir based on offline data, and obtain the relationship expression.
[0046] The trend prediction module is used to analyze the water level relationship of the reservoir, and predict the development trend of the water storage volume and water level of each reservoir through real-time data and meteorological data.
[0047] The resource allocation module is used to divide the observation area and allocate computing power, and is executed in parallel scheduling using the Spark engine.
[0048] The data acquisition module includes a monitoring data acquisition unit and a meteorological data acquisition unit.
[0049] The monitoring data acquisition unit is used to collect rainfall logs, water regime logs, and engineering situation logs.
[0050] The rainfall log includes precipitation records at different times, and each precipitation record includes the precipitation of each reservoir.
[0051] The water regime log includes hydrological records at different times, and each hydrological record includes the water storage volume of each reservoir and the water level values at different locations.
[0052] The engineering situation log includes GIS maps at different times.
[0053] The meteorological data acquisition unit is used to collect the predicted precipitation of each reservoir within a certain period of time in the future.
[0054] The precipitation, water storage volume, water level value, and GIS map at the latest time are used as real-time data, and the data at other times are used as offline data. The digital twin technology is used to construct a three-dimensional scene based on the GIS map at the latest time, and the real-time data is dynamically mapped into the three-dimensional scene.
[0055] The fusion analysis module includes a time analysis unit and a relationship construction unit.
[0056] The time analysis unit is used to establish a water storage set and a precipitation set for each reservoir.
[0057] First, analyze the upstream and downstream relationship of the reservoir according to the water flow direction in the three-dimensional scene, and all the upstream reservoirs directly connected to reservoir Q1 are used as its influencing objects.
[0058] Secondly, establish a water storage set and a precipitation set, obtain the precipitation records where the precipitation of reservoir Q1 and all its influencing objects is zero, and put the time of these precipitation records into the water storage set.
[0059] Finally, obtain the precipitation records where the precipitation of reservoir Q1 is not zero and the precipitation of all its influencing objects is zero, and put the times of these precipitation records into the precipitation set.
[0060] The relationship construction unit is used to analyze the relationship expressions of each reservoir.
[0061] First, obtain the hydrological records whose time intervals from any element in the water storage set are less than the time length threshold w. Take the water storage volume of reservoir Q1 in each hydrological record as the dependent variable, and the water storage volumes of all influencing objects of reservoir Q1 as independent variables and package them as samples. All samples are input into the polynomial regression model for training to obtain the water storage relationship expression XQ.
[0062] Second, obtain the hydrological records whose time intervals from any element in the precipitation set are less than w. Substitute the water storage volumes of all influencing objects of reservoir Q1 in each hydrological record into the expression XQ to calculate e, and subtract e from the water storage volume of reservoir Q1 to obtain the water increase volume of the corresponding hydrological record.
[0063] Finally, each hydrological record is matched with the precipitation records whose time intervals are less than w. Take the water increase volume of each hydrological record as the dependent variable, and the precipitation of reservoir Q1 in the matched precipitation record as the independent variable and package them as samples. All samples are input into the polynomial regression model for training to obtain the precipitation relationship expression.
[0064] Analyze the influencing objects and relationship expressions of each reservoir.
[0065] The trend prediction module includes a water storage prediction unit and a water level prediction unit.
[0066] The water storage prediction unit is used to calculate the predicted water storage volume of each reservoir.
[0067] First, take the downstream reservoir Q2 as the child node, and the upstream reservoir of Q2 as the parent node of Q2, and establish a tree-like relationship diagram according to the parent-child node relationship.
[0068] Second, obtain the predicted precipitation of each reservoir at time t1 in the meteorological data. Substitute the predicted precipitation of the root node Qg into the precipitation relationship expression to calculate u, and add u to the current water storage volume of Qg to obtain the predicted water storage volume Lg of Qg.
[0069] Then, continue to substitute Lg into the water storage relationship expression of the child node Qz of Qg to calculate p, and then substitute the predicted precipitation of the child node Qz into the precipitation relationship expression to calculate the result f, and add f to p to obtain the predicted water storage volume of the child node Qz.
[0070] Finally, calculate the predicted water storage volume of each reservoir step by step in the calculation direction from the parent node to the child node.
[0071] The water level prediction unit is used to calculate the predicted water level of each monitoring device.
[0072] First, analyze the reservoir to which the monitoring device belongs according to its location. Take the water level value of the monitoring device d1 in each hydrological record in the water regime log as the dependent variable, and package the water storage capacity of the reservoir to which the monitoring device d1 belongs as the independent variable into a sample.
[0073] Secondly, input all samples into the polynomial regression model for training to obtain the water level relationship expression of the monitoring device d1, and analyze the water level relationship expressions of each monitoring device.
[0074] Finally, substitute the predicted water storage capacity into the water level relationship expressions of each monitoring device in the reservoir to calculate the predicted water level of each monitoring device.
[0075] The resource allocation module includes a computing power allocation unit and a parallel scheduling unit.
[0076] The computing power allocation unit is used to calculate the maximum allocated computing power for each observation area.
[0077] First, analyze the positional distances between the monitoring device n and other monitoring devices in the same reservoir and calculate the average value to obtain the average distance H n , and calculate the average value H after summing up the average distances of all monitoring devices in the same reservoir ave .
[0078] Secondly, according to the formula: Calculate the trend index of each reservoir. Take the area where the reservoir with a trend index greater than the threshold c in the three-dimensional scene is located as the observation area. Among them, is the total number of all monitoring devices in the reservoir, is the predicted water level of the th monitoring device, is the current water level of the th monitoring device, is the warning water level of the th monitoring device.
[0079] Finally, obtain the remaining available computing power CP of the data center m and the trend index TR of the reservoir corresponding to each observation area v , and according to the formula: Calculate the maximum allocated computing power for each observation area respectively. Among them, TR sum is the total sum of the trend indices of the reservoirs corresponding to all observation areas.
[0080] The parallel scheduling unit uses the Spark engine to dynamically update the observation area and the maximum allocated computing power, and real-time monitors the total computing power utilized by the observation area. On the premise of not exceeding their respective maximum allocated computing powers, it parallelly calls more complex prediction models and adjusts faster computing frequencies for each observation area, so as to improve the prediction frequency and accuracy of each observation area.
[0081] The Spark engine execution includes the following steps:
[0082] Based on a preset trend index formula, use Spark SQL to parallelly calculate the real-time trend index of each reservoir on the full volume of data. Dynamically update the system-level risk threshold through sliding window statistics, and adaptively delimit high, medium, and low risk level intervals. After the results are marked, a list of reservoirs with priority labels is generated, providing a decision-making basis for computing power scheduling.
[0083] Build a quantitative mapping model from risk levels to computing resources: Set high computing power density and high prediction frequency for high-risk reservoirs. For medium-risk tasks, use a shared resource pool and allocate CPU cores as needed. Merge and execute low-risk batch jobs, and use time slice rotation to reduce resource occupancy. Through the Spark configuration parameterization template, realize the linkage adjustment of resource specifications and risk levels.
[0084] Intelligently slice the reservoir data set according to risk labels: The high-risk group uses fine-grained partitioning to maximize parallelism, and the medium- and low-risk groups aggregate partitions by geographical area to improve batch processing efficiency. Through a custom Spark scheduler plugin, preempt an independent resource pool for high-risk tasks in the YARN queue, set task priority weights, and ensure that high-risk and time-sensitive tasks obtain computing resources first.
[0085] Integrate heterogeneous hardware resources at the Spark computing layer: Schedule high-risk reservoir tasks to the GPU cluster to accelerate LSTM / CNN model inference, use CPU multi-threading to execute gradient boosting tree models for medium-risk tasks, and call pre-trained lightweight regression models for low-risk batches. Reduce data Shuffle overhead through memory optimization strategies, and combine checkpoint mechanisms to ensure the fault tolerance of long-term prediction tasks.
[0086] Write the prediction results to distributed storage in real time and visualize them through the Dashboard. Establish a closed-loop feedback mechanism: Use Spark ML to analyze the execution effects of historical tasks, optimize parameter sensitivity by setting and adjusting the weight coefficients in the maximum allocated computing power formula, and combine actual flood event data to improve the adaptability of dynamic scheduling strategies.
[0087] Deploy a cluster health monitoring subsystem to track the resource utilization rate, queue backlog, and completion rate of tasks at each risk level in real time. Preset a fuse threshold and support manually increasing the computing power quota of specified reservoirs. Compare system metrics under different scheduling strategies through A / B testing, and continuously optimize the resource allocation algorithm to ensure the prediction accuracy and timeliness of key targets during flood peaks.
[0088] Compared with the prior art, the beneficial effects achieved by the present invention are:
[0089] 1. Real-time data processing: Different from traditional offline monitoring systems, this solution can process monitoring data and meteorological data in real time based on the Spark engine, improving the timeliness of data analysis. This means that during critical periods such as the flood season, changes in water conditions can be responded to more quickly, enabling faster decision-making.
[0090] 2. Data fusion and unified parsing: Through distributed in-memory computing and data partitioning, this solution unifies the parsing of rainfall, water level, and project operation data into two-dimensional tabular data, eliminating the problem of data islands and improving the availability and guarantee rate of data. This efficient data integration makes decision-making more well-founded.
[0091] 3. Dynamic resource allocation: This solution dynamically adjusts computing resources by setting a trend index to ensure that high-risk reservoirs receive a higher computing frequency and support from complex models, optimizing the utilization efficiency of resources. At the same time, the system has a flexible resource expansion mechanism and can quickly increase computing resources when needed.
[0092] 4. Enhanced decision-making support: By displaying a dynamically mapped three-dimensional scene on a visualization platform, users can intuitively observe the impact of real-time data on the environment, supporting scientific decision-making and management. This intuitive data presentation method reduces potential risks caused by information asymmetry. Description of the Drawings
[0093] The drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation to the present invention. In the drawings:
[0094] Figure 1 is a schematic flowchart of the method in the present invention.
[0095] Figure 2 is a schematic diagram of the observation area planning of the method in the present invention.
[0096] Figure 3 is a schematic diagram of the structure of the system in the present invention.
[0097] Figure 4 is a schematic diagram of the result output of the system in the present invention. Detailed Embodiments
[0098] The present invention will be further described below with reference to the drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present invention and cannot be used to limit the protection scope of the present invention.
[0099] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0100] Please refer to Figure 1 , the present invention provides a method for real-time and offline data analysis of reservoirs based on the Spark engine, including the following steps:
[0101] S100. Collect monitoring data and meteorological data, and distinguish real-time data and offline data. In S100, the monitoring data includes rainfall logs, water regime logs, and engineering situation logs. The rainfall logs include precipitation records at different times, and each precipitation record includes the precipitation amounts of each reservoir. The water regime logs include hydrological records at different times, and each hydrological record includes the water storage amounts of each reservoir and the water level values at different positions, and the water level values are collected by monitoring devices installed near the reservoir dams. The engineering situation logs include GIS maps at different times. The GIS maps include the topographic structures, relative positions, and water flow directions of each reservoir. The meteorological data refers to the predicted precipitation amounts of each reservoir within a certain period of time in the future.
[0102] Monitoring data such as rainfall, water regime, and engineering situation exist in different storage media. The Spark engine realizes the unified parsing of monitoring data such as rainfall, water regime, and engineering situation into a two-dimensional tabular data structure for access and processing through distributed memory computing and data partitioning, greatly improving the data guarantee rate and playing an active role in communication guarantee during critical periods such as flood seasons.
[0103] Take the precipitation amount, water storage amount, water level value, and GIS map at the latest time as real-time data, and the data at other times as offline data. Use digital twin technology to construct a three-dimensional scene based on the GIS map at the latest time, and dynamically map the real-time data to the three-dimensional scene.
[0104] Specifically, it includes the following steps:
[0105] S101. Collect the precipitation amount, water storage amount, and water level value at the latest time in real time, and obtain the corresponding GIS map information. These data are updated in real time through sensors. Preprocess the collected real-time data, including data cleaning, format conversion, and standardization, to ensure the consistency and availability of the data.
[0106] S102. Use digital twin technology to construct a three-dimensional scene based on the latest GIS map information. This process involves converting the two-dimensional data in the geographic information system into a three-dimensional model to form a visual virtual environment.
[0107] S103. Dynamically map the processed precipitation, water storage volume, and water level values into a three-dimensional scene. Through a programming interface or visualization tool, bind the data to the corresponding parts of the three-dimensional model, so that the water bodies and terrain in the scene can reflect the data changes in real time.
[0108] S104. Display the dynamically mapped three-dimensional scene through a visualization platform, enabling users to intuitively observe the impact of real-time data on the environment and supporting decision-making and management.
[0109] Through the above steps, the effective combination of real-time data and the three-dimensional scene can be achieved, enhancing the monitoring and management capabilities of environmental changes.
[0110] S200. Analyze the water storage relationship and precipitation relationship of the reservoir based on offline data analysis. Specifically, S200 includes the following steps:
[0111] S201. Analyze the upstream and downstream relationships of the reservoir based on the water flow direction in the three-dimensional scene. All upstream reservoirs directly connected to reservoir Q1 are regarded as its influencing objects. Establish a water storage set and a precipitation set, obtain the precipitation records when the precipitation of reservoir Q1 and all its influencing objects is zero, and put the time of these precipitation records into the water storage set. Obtain the precipitation records when the precipitation of reservoir Q1 is not zero and the precipitation of all its influencing objects is zero, and put the time of these precipitation records into the precipitation set.
[0112] S202. Obtain the hydrological records whose time interval from any element in the water storage set is less than the time length threshold w. Take the water storage volume of reservoir Q1 in each hydrological record as the dependent variable, and the water storage volumes of all influencing objects of reservoir Q1 as independent variables and package them as samples. All samples are input into a polynomial regression model for training to obtain the water storage relationship expression X Q 。
[0113] S203. Obtain the hydrological records whose time interval from any element in the precipitation set is less than w. Substitute the water storage volumes of all influencing objects of reservoir Q1 in each hydrological record into the expression X Q to calculate e, and subtract e from the water storage volume of reservoir Q1 to obtain the water increase amount of the corresponding hydrological record. Each hydrological record is matched with a precipitation record whose time interval is less than w. Take the water increase amount of each hydrological record as the dependent variable, and the precipitation of reservoir Q1 in the matched precipitation record as the independent variable and package them as samples. All samples are input into a polynomial regression model for training to obtain the precipitation relationship expression. Analyze the influencing objects and relationship expressions of each reservoir.
[0114] S300. Analyze the water level relationship between reservoirs, and predict the development trend by integrating real-time data and meteorological data. Specifically, S300 includes the following steps:
[0115] S301. Take the downstream reservoir Q2 as a child node, and the upstream reservoir of Q2 as the parent node of Q2. Establish a tree - like relationship diagram according to the parent - child node relationship. Obtain the predicted precipitation of each reservoir at time t1 in the meteorological data, and substitute the predicted precipitation of the root node Q g into the precipitation relationship expression to calculate u. Add u to the current water storage volume of Q g to obtain the predicted water storage volume L g of Q g .
[0116] There may be multiple tree - like relationships in the tree - like relationship diagram, that is, there are multiple root nodes at the same time. Each root node only has child nodes and no parent nodes. When the predicted precipitation of the root node is zero, the predicted water storage volume takes the current water storage volume.
[0117] S302. Substitute L g continuously into the water storage relationship expression of the child node Q g of Q z to calculate p. Then substitute the predicted precipitation of the child node Q z into the precipitation relationship expression to calculate the result f. Add f to p to obtain the predicted water storage volume of the child node Q z . Calculate the predicted water storage volume of each reservoir progressively in the calculation direction from the parent node to the child node.
[0118] The predicted water storage volume of each child node is the sum of the results of the precipitation relationship expression and the water storage relationship expression, and the calculation is carried out in sequence according to the parent - child node order.
[0119] S303. Analyze the affiliated reservoir according to the location of the monitoring device. Take the water level value of the monitoring device d1 in each hydrological record in the water regime log as the dependent variable, and the water storage volume of the reservoir to which the monitoring device d1 belongs as the independent variable and package them as samples. Input all the samples into the polynomial regression model for training to obtain the water level relationship expression of the monitoring device d1, and analyze the water level relationship expressions of each monitoring device. Substitute the predicted water storage volume into the water level relationship expressions of each monitoring device in the reservoir to calculate the predicted water level of each monitoring device.
[0120] S400. Divide the observation area according to the development trend and allocate computing power, and execute in parallel using the Spark engine. Specifically, S400 includes the following steps:
[0121] S401. Analyze the position distance between the monitoring device n and other monitoring devices in the same reservoir and calculate the average value to obtain the average distance H n , sum the average distances of all monitoring devices in the same reservoir and then calculate the average value H ave , so as to calculate the trend index TR of each reservoir:
[0122] The formula for calculating the trend index is:
[0123] ;
[0124] Wherein, is the number of all monitoring devices in the reservoir, is the predicted water level of the th monitoring device, is the current water level of the th monitoring device, is the warning water level of the th monitoring device.
[0125] Monitoring devices are usually installed near the reservoir dam. Due to the inconsistent heights of the dams under natural environmental conditions, the warning water level values of different monitoring devices are different. The specific values are set in advance by the management personnel according to the actual situation, or are self-adjusted by the system according to the virtual dam mapped in real time in the three-dimensional scene.
[0126] The position of the monitoring device is set in advance by the management personnel. Usually, the more concentrated the monitoring devices are, the more important the location is. By calculating the average distance of the monitoring devices, the importance of the installation location to the reservoir can be analyzed. The shorter the average distance, the more monitoring devices are set near the location, and the more important it is to the reservoir.
[0127] S402. Take the area where the reservoir in the three-dimensional scene has a trend index greater than the threshold c as the observation area. Obtain the remaining available computing power of the data center and the trend index of the reservoir corresponding to each observation area , and substitute them into the formula to calculate the maximum allocated computing power for each observation area respectively:
[0128] ;
[0129] Wherein, is the sum of the trend indices of the reservoirs corresponding to all observation areas.
[0130] Adjust the calculation frequency and the complexity of the model according to the trend index, thereby indirectly affecting the use of resources. The reservoirs corresponding to the observation areas run models and tasks that are more complex and require more computing resources, while the reservoirs in non-observation areas run simple models and low-frequency tasks.
[0131] In such a scenario, the reservoirs in the same batch processing job can be divided into different groups according to the trend index, and each group applies models with different complexities and different running times. In Spark, for the data prediction task of each reservoir, different prediction models are selected and called according to its trend index, and at the same time, the calculation frequency is adjusted as much as possible. Each reservoir in the high-risk group needs to be processed multiple times in the same batch processing, or scheduled at shorter intervals.
[0132] S403, please refer to Figure 2 , the Spark engine is used to dynamically update the observation area and the maximum allocated computing power, and the total computing power used in the observation area is monitored in real time; without exceeding the respective maximum allocated computing power, more complex prediction models are called in parallel for each observation area and a faster calculation frequency is adjusted, thereby improving the prediction frequency and accuracy of each observation area.
[0133] Calling more complex prediction models and adjusting faster calculation frequencies can preview the dynamic evolution of dam breaches under extreme conditions, and display the corresponding dam breach flow process and corresponding dam breach parameter data. On the one hand, it can effectively guide emergency rescue work in reservoirs under extreme conditions, and on the other hand, it can reversely guide the preparation of reservoir flood control plans and improve the accuracy of the plans.
[0134] The dynamic mechanism based on the Spark engine can monitor and update the list of observation areas in real time, and automatically adjust the resource supply according to the preset maximum allocated computing power threshold. When the reservoir trend index of a certain observation area exceeds the threshold, the system will trigger an emergency expansion strategy: temporarily call idle nodes in the cluster or enable elastic expansion of cloud services, increase the number of GPU nodes and increase the prediction frequency to ensure that the prediction tasks of the reservoirs in the observation area can obtain up to 100% of the maximum allocated computing power again. At the same time, dynamically reduce the resource quota of non-observation areas to maintain global load balancing.
[0135] See also Figure 3 Based on the same invention, the present invention provides a real-time and offline reservoir data analysis system based on the Spark engine, including a data acquisition module, a fusion analysis module, a trend prediction module and a resource allocation module.
[0136] The data acquisition module is used to collect monitoring data and meteorological data, and distinguish between real-time data and offline data. The data acquisition module includes a monitoring data acquisition unit and a meteorological data acquisition unit.
[0137] The monitoring data collection unit is used to collect rainfall logs, water logs and work logs.
[0138] The rainfall log includes precipitation records at different times, and each precipitation record includes the precipitation of each reservoir. The water log includes hydrological records at different times, and each hydrological record includes the water storage capacity of each reservoir and the water level values at different locations. The construction log includes GIS maps at different times.
[0139] The meteorological data collection unit is used to collect the predicted precipitation of each reservoir in the future period of time.
[0140] The precipitation, water storage volume, water level value at the latest time and the GIS map are used as real-time data, and the data at other times are used as offline data. The digital twin technology is adopted to construct a three-dimensional scene according to the GIS map at the latest time, and the real-time data is dynamically mapped into the three-dimensional scene.
[0141] The fusion analysis module is used to analyze the water storage relationship and precipitation relationship of the reservoir according to the offline data, and obtain the relationship expression. The fusion analysis module includes a time analysis unit and a relationship construction unit.
[0142] The time analysis unit is used to establish a water storage set and a precipitation set for each reservoir.
[0143] First, analyze the upstream and downstream relationships of the reservoirs according to the water flow direction in the three-dimensional scene. All the upstream reservoirs directly connected to reservoir Q1 are taken as its influencing objects.
[0144] Secondly, establish a water storage set and a precipitation set, obtain the precipitation records with zero precipitation for reservoir Q1 and all its influencing objects, and put the time of these precipitation records into the water storage set.
[0145] Finally, obtain the precipitation records with non-zero precipitation for reservoir Q1 and zero precipitation for all its influencing objects, and put the time of these precipitation records into the precipitation set.
[0146] The relationship construction unit is used to analyze the relationship expressions of each reservoir.
[0147] First, obtain the hydrological records with a time interval less than the duration threshold w from any element in the water storage set. Take the water storage volume of reservoir Q1 in each hydrological record as the dependent variable, and the water storage volumes of all influencing objects of reservoir Q1 as independent variables and package them as samples. All samples are input into the polynomial regression model for training to obtain the water storage relationship expression X Q .
[0148] Secondly, obtain the hydrological records with a time interval less than w from any element in the precipitation set, and substitute the water storage volumes of all influencing objects of reservoir Q1 in each hydrological record into the expression X Q to calculate e, and subtract e from the water storage volume of reservoir Q1 to obtain the water increase volume of the corresponding hydrological record.
[0149] Finally, each hydrological record is matched with the precipitation record with a time interval less than w. Take the water increase volume of each hydrological record as the dependent variable, and the precipitation of reservoir Q1 in the matched precipitation record as the independent variable and package them as samples. All samples are input into the polynomial regression model for training to obtain the precipitation relationship expression. Analyze the influencing objects and relationship expressions of each reservoir.
[0150] The trend prediction module is used to analyze the water level relationship of the reservoir, and predict the development trend of the water storage volume and water level of each reservoir through real-time data and meteorological data. The trend prediction module includes a water storage prediction unit and a water level prediction unit.
[0151] The water storage prediction unit is used to calculate the predicted water storage volume of each reservoir.
[0152] First, take the downstream reservoir Q2 as the child node, and the upstream reservoir of Q2 as the parent node of Q2, and establish a tree-like relationship diagram according to the parent-child node relationship.
[0153] Secondly, obtain the predicted precipitation of each reservoir at time t1 in the meteorological data, and substitute the predicted precipitation of the root node Q g into the precipitation relationship expression to calculate u, and add u to the current water storage volume of Q g to obtain the predicted water storage volume L g of Q. g .
[0154] Then, continue to substitute L g into the water storage relationship expression of the child node Q g of Q to calculate p, and then substitute the predicted precipitation of the child node Q z into the precipitation relationship expression to calculate the result f, and add f to p to obtain the predicted water storage volume of the child node Q z . z Finally, calculate the predicted water storage volume of each reservoir step by step in the calculation direction from the parent node to the child node.
[0155] Finally, calculate the predicted water storage volume of each reservoir step by step in the calculation direction from the parent node to the child node.
[0156] The water level prediction unit is used to calculate the predicted water level of each monitoring device.
[0157] First, analyze the reservoir to which the monitoring device belongs according to the location of the monitoring device, and take the water level value of the monitoring device d1 in each hydrological record in the water regime log as the dependent variable, and the water storage volume of the reservoir to which the monitoring device d1 belongs as the independent variable and package it as a sample.
[0158] Secondly, input all samples into the polynomial regression model for training to obtain the water level relationship expression of the monitoring device d1, and analyze the water level relationship expressions of each monitoring device.
[0159] Finally, substitute the predicted water storage volume into the water level relationship expressions of each monitoring device in the reservoir to calculate the predicted water level of each monitoring device.
[0160] The resource allocation module is used to divide the observation area and allocate computing power, and is executed in parallel scheduling using the Spark engine. The resource allocation module includes a computing power allocation unit and a parallel scheduling unit.
[0161] The computing power allocation unit is used to calculate the maximum allocated computing power of each observation area.
[0162] First, analyze the positional distances between the monitoring device n and other monitoring devices in the same reservoir, calculate the average value to obtain the average distance. Sum up the average distances of all monitoring devices in the same reservoir and then calculate the average value. .
[0163] Secondly, according to the formula: Calculate the trend index of each reservoir. Take the area where the reservoir with a trend index greater than the threshold c is located in the three-dimensional scene as the observation area. Among them, is the total number of all monitoring devices in the reservoir, is the predicted water level of the th monitoring device, is the current water level of the th monitoring device, is the warning water level of the th monitoring device.
[0164] Finally, obtain the remaining available computing power of the data center and the trend index of the reservoir corresponding to each observation area , according to the formula: Calculate the maximum allocated computing power for each observation area respectively. Among them, is the total sum of the trend indices of the reservoirs corresponding to all observation areas.
[0165] The parallel scheduling unit uses the Spark engine to dynamically update the observation area and the maximum allocated computing power, and real-time monitors the total computing power utilized by the observation area. On the premise of not exceeding their respective maximum allocated computing powers, more complex prediction models are called in parallel for each observation area and the calculation frequency with faster adjustment is adopted, so as to improve the prediction frequency and accuracy of each observation area.
[0166] Please refer to Figure 4 , the Spark engine execution includes the following steps:
[0167] Based on the preset trend index formula, use Spark SQL to calculate the real-time trend index of each reservoir in parallel on the full amount of data. Dynamically update the system-level risk threshold through sliding window statistics, and adaptively delimit the high, medium, and low risk level intervals. After the results are marked, a list of reservoirs with priority labels is generated, providing a decision-making basis for computing power scheduling.
[0168] Build a quantitative mapping model from risk levels to computing resources: Set high computing power density (such as exclusive GPU nodes per task) and high prediction frequency (minute level) for high-risk reservoirs. For medium-risk tasks, use a shared resource pool and allocate CPU cores on demand. Merge and execute low-risk batch jobs and use time slice rotation to reduce resource occupancy. Through the Spark configuration parameterization template, realize the linkage adjustment of resource specifications (such as the number of Executors, memory quota) and risk levels.
[0169] Intelligently slice the reservoir dataset according to risk labels: The high-risk group uses fine-grained partitioning (such as single reservoir data per partition) to maximize parallelism, and the medium- and low-risk groups aggregate partitions by geographical area to improve batch processing efficiency. Through a custom Spark scheduler plugin, preempt an independent resource pool for high-risk tasks in the YARN queue and set task priority weights (such as multi-pool configuration of Fair Scheduler) to ensure that high-risk and time-sensitive tasks obtain computing resources first.
[0170] Integrate heterogeneous hardware resources at the Spark computing layer: Schedule high-risk reservoir tasks to the GPU cluster to accelerate LSTM / CNN model inference, use CPU multi-threading to execute gradient boosting tree models for medium-risk tasks, and call pre-trained lightweight regression models for low-risk batches. Reduce data Shuffle overhead through memory optimization strategies (such as off-heap memory allocation, RDD persistence strategy), and combine checkpoint mechanisms to ensure the fault tolerance of long-term prediction tasks.
[0171] Write the prediction results to distributed storage (such as HBase + HDFS) in real time and visualize them through the Dashboard. Establish a closed-loop feedback mechanism: Use Spark ML to analyze the execution effects of historical tasks (prediction error, response delay), adjust the weight coefficients in the maximum allocated computing power formula by setting, and optimize parameter sensitivity combined with actual flood event data to improve the adaptability of dynamic scheduling strategies.
[0172] Deploy a cluster health monitoring subsystem to track the resource utilization rate, queue backlog, and completion rate of tasks at each risk level in real time. Preset a fuse threshold (such as automatically triggering an alarm when the high-risk task delay exceeds 5 minutes), and support manually increasing the computing power quota of specified reservoirs. Compare system metrics under different scheduling strategies through A / B testing, and continuously optimize the resource allocation algorithm to ensure the prediction accuracy and timeliness of key targets during flood peaks.
[0173] Example 1: Suppose there are 4 monitoring devices A1, A2, A3, and A4 in reservoir Q8, and their average distances are: 26m, 24m, 36m, and 18m respectively. Their water level information is as follows:
[0174] A1: Predicted water level: 180cm; Current water level: 150cm; Warning water level: 250cm;
[0175] A2: Predicted water level: 160 cm; Current water level: 140 cm; Warning water level: 210 cm;
[0176] A3: Predicted water level: 150 cm; Current water level: 140 cm; Warning water level: 180 cm;
[0177] A4: Predicted water level: 170 cm; Current water level: 160 cm; Warning water level: 240 cm;
[0178] Substitute into the formula to calculate the trend index of reservoir Q8:
[0179] ;
[0180] Then the trend index of reservoir Q8 is 0.32.
[0181] In summary, the present technical solution realizes precise computing power flow allocation with risk as the weight through a multi-level algorithm, meeting the core requirements of key target key guarantee in intelligent water conservancy. It significantly improves the real-time performance, accuracy and emergency response ability of reservoir monitoring and management, and has obvious advantages compared with the prior art.
[0182] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device.
[0183] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacement on some of the technical features. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
[0184] The above are only the preferred embodiments of the present invention. It should be pointed out that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and deformations can still be made, and these improvements and deformations should also be regarded as the protection scope of the present invention.
Claims
1. A real-time and offline data analysis method for reservoirs based on the Spark engine, characterized in that: The method includes the following steps: S100. Collect monitoring data and meteorological data, and distinguish real-time data and offline data; S200. Analyze the water storage relationship and precipitation relationship of the reservoir according to the offline data; S300. Analyze the water level relationship between reservoirs, and comprehensively predict the development trend based on real-time data and meteorological data; S400. Divide the observation area and allocate computing power according to the development trend, and execute in parallel using the Spark engine; In S100, the monitoring data includes rainfall logs, water regime logs, and engineering situation logs; The rainfall log includes precipitation records at different times, and each precipitation record includes the precipitation of each reservoir; The water regime log includes hydrological records at different times, and each hydrological record includes the water storage volume of each reservoir and the water level values at different positions. The water level values are collected by monitoring devices installed near the reservoir dam; The engineering situation log includes GIS maps at different times; The GIS map includes the topographic structure, relative position, and water flow direction of each reservoir; the meteorological data refers to the predicted precipitation of each reservoir in a future period of time; Take the precipitation, water storage volume, water level value, and GIS map at the latest time as real-time data, and the data at other times as offline data; use digital twin technology to construct a three-dimensional scene based on the GIS map at the latest time, and dynamically map the real-time data into the three-dimensional scene; S200 includes the following steps: S201. Analyze the upstream and downstream relationship of the reservoir according to the water flow direction in the three-dimensional scene. All upstream reservoirs directly connected to reservoir Q1 are regarded as its influencing objects; S202. Establish a water storage set and a precipitation set, obtain precipitation records where the precipitation of reservoir Q1 and all its influencing objects is zero, and put the times of these precipitation records into the water storage set; S203. Obtain precipitation records where the precipitation of reservoir Q1 is not zero and the precipitation of all its influencing objects is zero, and put the times of these precipitation records into the precipitation set; S204. Analyze the water storage set and the precipitation set, establish a water storage relationship expression and a precipitation relationship expression for reservoir Q1 respectively, and analyze the influencing objects and relationship expressions of other reservoirs; specifically including: S2041. Obtain hydrological records with a time interval less than the time length threshold w from any element in the water storage set; Taking the water storage volume of reservoir Q1 in each hydrological record as the dependent variable, and the water storage volumes of all influencing objects of reservoir Q1 as independent variables and packing them into samples, all samples are input into the polynomial regression model for training to obtain the water storage relationship expression X Q ; S2042. Obtain the hydrological records with a time interval less than w from any element in the precipitation ensemble, and substitute the water storage volumes of all the influencing objects of reservoir Q1 in each hydrological record into the expression X Q Calculate e by substituting into the formula, and subtract e from the water storage volume of reservoir Q1 to obtain the increased water volume of the corresponding hydrological record; S2043. Each hydrological record matches a precipitation record with a time interval less than w. The increased water volume of each hydrological record is used as the dependent variable, and the precipitation of reservoir Q1 in the matched precipitation record is used as the independent variable and packaged as a sample. All samples are input into the polynomial regression model for training to obtain the precipitation relationship expression.
2. The real-time and offline data analysis method for reservoir based on Spark engine according to claim 1, characterized in that: S300 includes the following steps: S301. Take the downstream reservoir Q2 as the child node, and the upstream reservoir of Q2 as the parent node of Q2, and establish a tree-like relationship diagram according to the parent-child node relationship; S302. Obtain the predicted precipitation of each reservoir at time t1 in the meteorological data, and substitute the predicted precipitation of the root node Q g into the precipitation relationship expression to calculate u. Add u to the current water storage volume of Q g to obtain the predicted water storage volume L g of Q g ; S303. Substitute L g continuously into the child node Q g of z to calculate p in the water storage relationship expression, and then substitute the predicted precipitation of the child node Q z into the precipitation relationship expression to calculate the result f; S304. Then add f and p to obtain child node Q z For the predicted water storage volume, calculate the predicted water storage volume of each reservoir progressively in sequence according to the calculation direction from the parent node to the child node; S305. Analyze the water level relationship expression of each monitoring device according to the water regime log, substitute the predicted water storage volume into the water level relationship expression of each monitoring device in the reservoir respectively, and calculate the predicted water level of each monitoring device.
3. The real-time and offline data analysis method for reservoir based on Spark engine according to claim 2, wherein: The establishment of the water level relationship expression includes the following steps: Analyze the reservoir to which each monitoring device location belongs. Take the water level value of monitoring device d1 in each hydrological record in the water regime log as the dependent variable, and the water storage volume of the reservoir to which monitoring device d1 belongs as the independent variable and package it as a sample. Input all samples into the polynomial regression model for training to obtain the water level relationship expression of monitoring device d1. And so on, analyze the water level relationship expressions of each monitoring device.
4. A real-time and offline data analysis method for a reservoir based on the Spark engine according to claim 2, characterized in that: S400 includes the following steps: S401. Analyze the position distances between the monitoring device n and other monitoring devices in the same reservoir, and calculate the average value to obtain the average distance H n , sum up the average distances of all monitoring devices in the same reservoir and calculate the average value H ave , thereby calculating the trend index TR of each reservoir; S402. Take the area where the reservoir with a trend index greater than the threshold c in the three-dimensional scene as the observation area; obtain the remaining available computing power CP of the data center m and the trend index TR of the reservoir corresponding to each observation area v , and calculate the maximum allocated computing power CP for each observation area respectively max ; S403. Use the Spark engine to dynamically update the observation area and the maximum allocated computing power, and monitor the total computing power utilized by the observation area in real time; On the premise of not exceeding their respective maximum allocated computing powers, parallelly call more complex prediction models and adjust faster calculation frequencies for each observation area, so as to improve the prediction frequency and accuracy of each observation area.
5. A real-time and offline data analysis method for a reservoir based on the Spark engine according to claim 4, characterized in that: In S401, the calculation formula for the trend index is: ; where k is the number of all monitoring devices in the reservoir, is the predicted water level of the nth monitoring device, is the current water level of the nth monitoring device, is the warning water level of the nth monitoring device.
6. The real-time and offline data analysis method for reservoir based on Spark engine according to claim 4, wherein: In S402, the calculation formula for the maximum allocated computing power is: ; where, TR sum is the sum of the trend indices of the reservoirs corresponding to all the observation areas.
7. A real-time and offline data analysis system for reservoirs based on the Spark engine, which is applied to a real-time and offline data analysis method for reservoirs based on the Spark engine as described in claim 1, and is characterized in that: It includes a data acquisition module, a fusion analysis module, a trend prediction module, and a resource allocation module; The data acquisition module is used to collect monitoring data and meteorological data, and distinguish between real-time data and offline data; The fusion analysis module is used to analyze the water storage relationship and precipitation relationship of the reservoir based on offline data, and obtain the relationship expression; The trend prediction module is used to analyze the water level relationship of the reservoir, and predict the development trends of the water storage volume and water level of each reservoir through real-time data and meteorological data; The resource allocation module is used to divide the observation area and allocate computing power, and use the Spark engine for parallel scheduling and execution.
Citation Information
Patent Citations
Water resource optimization scheduling management method and system based on digital twinning
CN118536773A
KR20240039858A