Water conservancy project monitoring data processing method and system based on cloud computing
By extracting the temporal and spatial correlation characteristics of water conservancy project monitoring data on a cloud computing platform, setting constraints and dynamically allocating weights, the problem of lack of global analysis in water conservancy project monitoring data processing is solved, the accuracy and reliability of data completion are achieved, and scientific decision-making for water conservancy projects is supported.
Patent Information
- Application Number
- CN202511269561.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-08
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2045-09-08
AI Technical Summary
Existing technologies for processing water conservancy project monitoring data lack a global analysis of time and space, cannot effectively complete data, and lack constraints and dynamic allocation weights of contribution, resulting in insufficient data accuracy and reliability.
By using cloud computing-based methods, the temporal and spatial correlation features of monitoring data are extracted to form correlation relationships. Timeliness, correlation and consistency constraints are set, weights are dynamically allocated, and data completion is performed.
It enables comprehensive analysis of monitoring data, improves the accuracy and rationality of data completion, ensures that the output monitoring data can accurately reflect the actual status of key indicators, supports real-time operating condition assessment and safety risk early warning of water conservancy projects, and improves the level of management refinement and intelligence.
Smart Images

Figure CN120804602B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application relates to the technical field of data processing, and in particular relates to a water conservancy project monitoring data processing method and system based on cloud computing. BACKGROUND
[0002] In the process of water conservancy project monitoring, a large amount of monitoring data is generated, which is of great significance for mastering the operation state of the water conservancy project and ensuring the safety of the project. However, due to sensor failure, transmission interruption, adverse environmental interference and other reasons, the monitoring data often appears missing, which affects the accuracy and reliability of the monitoring results. With the continuous expansion of the scale of water conservancy projects and the development of monitoring technology, the amount of monitoring data increases dramatically, and the traditional data processing method has low processing efficiency and cannot realize real-time processing when facing massive monitoring data. Cloud computing has powerful computing power, massive storage capacity and efficient distributed processing capacity, which can provide strong support for the processing of massive monitoring data. Applying cloud computing technology to water conservancy project monitoring data processing is expected to solve the problems existing in the traditional method.
[0003] At present, in the Chinese invention with the publication number CN119647783A, a data analysis and processing method and system based on water conservancy project are disclosed. The application analyzes the to-be-processed water conservancy project data in the database system through the data management terminal, locates the abnormal data position, analyzes the abnormal type, focuses on locating and type judgment of the missing data, determines whether it can be supplemented, and if it cannot be supplemented, it is deleted to avoid the missing data interfering with the subsequent water conservancy project power generation equipment operation state analysis. Although the water conservancy project data that is supplemented is sorted through the continuity of time, the analysis of the water conservancy project data is more convenient and fast, but the related technology does not fuse and analyze the water conservancy project data from different dimensions of time and space, lacks the globality of data analysis, does not realize data completion by setting constraint conditions and dynamically allocating weights according to contribution degrees, and has certain limitations. SUMMARY
[0004] The technical problem solved by the application is that the related technology does not fuse and analyze the water conservancy project data from different dimensions of time and space, lacks the globality of data analysis, does not realize data completion by setting constraint conditions and dynamically allocating weights according to contribution degrees, and has certain limitations.
[0005] To solve the above technical problems, the application provides the following technical scheme, a water conservancy project monitoring data processing method based on cloud computing, comprising the following steps:
[0006] Step S1, pre-processing the received monitoring data to obtain a target data segment;
[0007] Step S2: Analyze the target data segment to obtain temporal correlation features and spatial correlation features, and form a correlation relationship based on the temporal correlation features and spatial correlation features;
[0008] Step S3: Set constraints, calculate the dependency weight of the association relationship based on the constraints, and complete the target data segment using the dependency weight to obtain complete monitoring data.
[0009] As a preferred embodiment of the cloud computing-based water conservancy project monitoring data processing method of the present invention, the preprocessing of the monitoring data includes removing outliers and removing duplicates. The monitoring data includes dam displacement data, reservoir water level data, rainfall data, and monitoring equipment operating voltage.
[0010] The outlier removal is used to remove monitoring data that exceeds a preset reasonable range;
[0011] The deduplication feature is used to remove multiple identical monitoring data points collected by the same monitoring device at the same time point, while retaining only one identical monitoring data point.
[0012] The preprocessed monitoring data is marked for validity to obtain valid monitoring data. The valid monitoring data is compared with the theoretical monitoring sequence, and the proportion of valid monitoring data in each time period is calculated. Time periods whose proportions do not reach the preset complete threshold are marked as target data segments. The target data segments include missing time nodes, missing time intervals, and missing monitoring points.
[0013] The theoretical monitoring sequence is theoretical monitoring data collected according to a preset acquisition frequency.
[0014] As a preferred embodiment of the cloud computing-based water conservancy engineering monitoring data processing method of the present invention, the time correlation features include the numerical change rate and periodic fluctuation characteristics of the M consecutive valid monitoring data before the target data segment and the initial change trend of the K consecutive valid monitoring data after the target data segment, wherein M and K are preset positive integers;
[0015] The numerical change rate is used to extract the numerical values {V1, V2, ..., V} corresponding to the M consecutive valid monitoring data points preceding the target data segment. M} and timestamp {T M1 ,T M2 ,...,T MMThe time interval is calculated by taking the numerical difference and timestamp difference between two adjacent valid monitoring data in the M consecutive valid monitoring data before the target data segment. The ratio of the numerical difference to the time interval is taken as the instantaneous change rate between the two adjacent data. The average value of the M-1 instantaneous change rates calculated from the M consecutive valid monitoring data is taken as the numerical change rate of the data before the target data segment.
[0016] The periodic fluctuation feature is obtained by retrieving monitoring data from historical periods, generating trend curves for the monitoring data in a time series, calculating the overlap of monitoring data fluctuations within historical time intervals, locating historical time intervals where the monitoring data shows periodic increases and decreases, and if there are multiple historical time intervals with a high overlap, selecting the historical time interval with the highest impact on data fluctuations as the periodic pattern, determining the duration of the fluctuation cycle, and based on the fluctuation cycle, extracting the peak, trough, and fluctuation amplitude of the monitoring data in historical periods within each fluctuation cycle, and using the cycle duration, peak, trough, and fluctuation amplitude as the periodic fluctuation feature of the data before the target data segment;
[0017] The time range of the monitoring data for the historical period must cover at least one complete fluctuation cycle;
[0018] The initial trend is used to extract the numerical values {U1, U2, ..., U} corresponding to K consecutive valid monitoring data after the target data segment is extracted. K} and timestamp {T K1 ,T K2 ,...,T KK A Cartesian coordinate system is established with timestamps on the horizontal axis and numerical values on the vertical axis. A regression line is constructed by effectively monitoring the distribution patterns of data points. The initial trend of change is determined based on the slope of the regression line.
[0019] If the slope is positive, then the initial trend of change is upward;
[0020] If the slope is negative, the initial trend of change is downward;
[0021] If the slope is close to 0, the initial trend of change is stable.
[0022] As a preferred embodiment of the cloud computing-based water conservancy engineering monitoring data processing method of the present invention, the spatial correlation features include the range of numerical differences between adjacent monitoring points of other monitoring points in the same area as the target data segment and the data distribution gradient features.
[0023] The range of numerical differences between adjacent monitoring points is obtained by calculating the difference between the monitoring data values of each pair of adjacent monitoring points and taking the absolute value. The absolute differences calculated from the numerical differences of all adjacent monitoring points are statistically analyzed, and the minimum and maximum values are selected to form the range.
[0024] The data distribution gradient feature is obtained by acquiring the three-dimensional structural coordinates of each monitoring point in the same area, calculating the spatial straight-line distance between two adjacent monitoring points as the coordinate difference, and dividing the coordinate difference by the corresponding spatial straight-line distance to obtain the numerical change per unit distance, which is the data distribution gradient feature in the same area.
[0025] As a preferred embodiment of the cloud computing-based water conservancy project monitoring data processing method of the present invention, the formation of a correlation relationship based on the temporal correlation features and the spatial correlation features specifically includes:
[0026] Based on the timestamps of the M consecutive valid monitoring data preceding the target data segment {T M1 ,T M2 ,...,T MM} and the timestamps of the K consecutive valid monitoring data points {T} K1 ,T K2 ,...,T KK}, generate a continuous time series, divide the continuous time series into time periods according to a preset time granularity, so that each time period corresponds to a unique timestamp;
[0027] The three-dimensional structural coordinates of each monitoring point involved in the spatial correlation feature are obtained, and the monitoring points are sorted according to the three-dimensional structural coordinates to form a spatial grid. Each monitoring point in the spatial grid is identified with a unique three-dimensional structural coordinate position.
[0028] By mapping the unique timestamp to the unique three-dimensional structural coordinates, the continuous time series after the time period is divided is dimensionally aligned with the spatial grid to form a one-to-one association with the target data segment. The association includes the monitored values, timestamp information, monitoring points, and value change trends.
[0029] As a preferred embodiment of the cloud computing-based water conservancy project monitoring data processing method of the present invention, the set constraints include timeliness constraints, correlation constraints and consistency constraints.
[0030] The timeliness constraint is used to determine the time interval threshold between the association relationship and the target data segment in the time dimension. If the difference between the timestamp of the association relationship and the timestamp of the target data segment exceeds the time interval threshold, the timeliness weight of the association relationship is reduced.
[0031] The time interval threshold is the maximum effective time difference determined based on the missing time interval and the period duration of the periodic fluctuation characteristics;
[0032] The correlation constraint is used to determine the degree of matching between the correlation relationship on the parameter attributes and the target data segment. The correlation weight is positively correlated with the degree of matching. The parameter attributes include dam displacement, reservoir water level, rainfall and monitoring equipment operating voltage.
[0033] The consistency constraint is used to determine whether the correlation in terms of numerical change trend is consistent with the historical data change trend of the monitoring area where the target data segment is located. If there is a trend conflict, it means that the consistency weight of the correlation is reduced.
[0034] In a preferred embodiment of the cloud computing-based water conservancy project monitoring data processing method of the present invention, the dependency weight of the correlation relationship is calculated according to the constraints, specifically as follows:
[0035] Weighting coefficients are assigned to the constraints, and the formula for calculating the weighting coefficients is as follows:
[0036] W i =Cont i / (Sum cont123 (i=1, 2, 3);
[0037] Where Wᵢ represents the weight coefficient of the i-th constraint, Contᵢ represents the contribution of the i-th constraint, and Sum cont123 This represents the total contribution of the constraints, and the sum of the weight coefficients of all constraints is 1.
[0038] Calculate the score of the association under each constraint, multiply the score under each constraint by the corresponding weight coefficient, and sum the results to obtain the dependency weight of the association.
[0039] As a preferred embodiment of the cloud computing-based water conservancy project monitoring data processing method of the present invention, the calculation of the score of the correlation under various constraints includes timeliness score calculation, correlation score calculation and consistency score calculation.
[0040] The timeliness score calculation sets a threshold for the time interval between the association relationship and the target data segment, and calculates the actual time interval between the timestamp of the association relationship and the timestamp of the target data segment. The smaller the actual time interval, the higher the timeliness score.
[0041] The correlation score is calculated and analyzed to determine the correlation relationship and the parameter attributes of the target data segment. The matching degree between the two is compared through a preset parameter attribute matching table. The higher the matching degree, the higher the correlation score.
[0042] The consistency score calculation compares the trend of the numerical change with the trend of periodic fluctuation characteristics, and calculates the degree of trend consistency between the two. The degree of consistency is the consistency score. The more consistent the trend, the higher the consistency score.
[0043] The scores range from 0 to 1.
[0044] As a preferred embodiment of the cloud computing-based water conservancy project monitoring data processing method of the present invention, the data completion operation of the target data segment according to the dependency weight is specifically as follows:
[0045] The associations are sorted from high to low according to their dependency weights. The temporal and spatial association features corresponding to the top N associations by dependency weight are extracted, and missing value judgment is performed. The missing value judgment calculation includes:
[0046] If the missing value of the target data segment is a periodic physical quantity, then the average value of the corresponding positions in the same period before and after the missing time interval is taken.
[0047] If the missing values of the target data segment are physical quantities that show a trend, then the increment value of the missing time interval is calculated according to the rate of change of the value.
[0048] The missing values of the target data segment are generated and filled into the missing positions of the target data segment to obtain the complete monitoring data after completion, where N is a positive integer set according to the missing length of the target data segment.
[0049] A cloud-based water conservancy project monitoring data processing system, including a preprocessing module, an analysis module, and a completion module;
[0050] The preprocessing module is used to preprocess the received monitoring data to obtain the target data segment;
[0051] The analysis module is used to analyze the target data segment, obtain temporal correlation features and spatial correlation features, and form a correlation relationship based on the temporal correlation features and the spatial correlation features;
[0052] The completion module is used to set constraints, calculate the dependency weight of the association based on the constraints, and perform data completion operation on the target data segment based on the dependency weight to obtain the complete monitoring data after completion.
[0053] The beneficial effects of this invention are as follows: By extracting the temporal and spatial correlation features of monitoring data and performing spatiotemporal fusion analysis based on both, a global control over the analysis of monitoring data is achieved, enabling a more comprehensive capture of the inherent patterns in the monitoring data. By setting timeliness, correlation, and consistency constraints, and dynamically allocating the weight coefficients of each constraint based on contribution to calculate the dependency weight of the correlation relationship, and then performing differentiated data completion based on high-weight correlation features, the accuracy and rationality of data completion are effectively improved. The output complete monitoring data can clearly reflect the actual status and changing trends of key indicators such as dam displacement and reservoir water level, providing reliable support for scientific decision-making in real-time operational condition assessment, safety risk early warning, project scheduling, and maintenance plan formulation for water conservancy projects. This helps to improve the refinement and intelligence level of water conservancy project management, ensuring the safe operation of the project and maximizing its benefits. Attached Figure Description
[0054] Figure 1 A schematic diagram of the basic process of a cloud computing-based water conservancy project monitoring data processing method provided in one embodiment of the present invention. Detailed Implementation
[0055] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.
[0056] Example 1, referring to Figure 1 As an embodiment of the present invention, a cloud computing-based method for processing water conservancy project monitoring data is provided, comprising the following steps:
[0057] Step S1: Preprocess the received monitoring data to obtain the target data segment;
[0058] Step S2: Analyze the target data segment to obtain temporal correlation features and spatial correlation features, and form a correlation relationship based on the temporal correlation features and spatial correlation features;
[0059] Step S3: Set constraints, calculate the dependency weights of the association based on the constraints, and complete the target data segment using the dependency weights to obtain complete monitoring data.
[0060] Preprocessing of monitoring data includes removing outliers and duplicates. The monitoring data includes dam displacement data, reservoir water level data, rainfall data, and the operating voltage of the monitoring equipment.
[0061] Outlier removal is used to remove monitoring data that exceeds a preset reasonable range;
[0062] The function to remove duplicate values is used to remove multiple identical monitoring data points collected by the same monitoring device at the same time point, while retaining only one of the identical monitoring data points.
[0063] The preprocessed monitoring data is marked for validity to obtain valid monitoring data. The valid monitoring data is compared with the theoretical monitoring sequence, and the proportion of valid monitoring data in each time period is calculated. Time periods whose proportions do not reach the preset complete threshold are marked as target data segments. Target data segments include missing time nodes, missing time intervals, and missing monitoring points.
[0064] The theoretical monitoring sequence consists of theoretical monitoring data collected according to a preset acquisition frequency.
[0065] In a specific embodiment, outliers and duplicates are removed. For example, outliers of dam displacement exceeding the equipment's range are eliminated, and only one of the three identical water level data collected by a monitoring device at the same time point is retained. This ensures that the monitoring data has no fundamental problems. Then, by comparing with the theoretical monitoring sequence, target data segments are marked. For example, if the effective data ratio of a certain period is only 80% and does not reach the 90% complete threshold, it is marked as a target data segment containing missing data. This accurately locates the missing data area, clarifies the processing objects for subsequent analysis and completion, and ensures data quality.
[0066] The preset integrity threshold refers to the standard for the proportion of valid data set in advance to determine the integrity of monitoring data for a certain period. It is used to define whether the monitoring data for a certain period needs to be handled in a key manner due to excessive data loss.
[0067] The time-related features include the rate of change of the values of the M consecutive valid monitoring data before the target data segment, the periodic fluctuation characteristics, and the initial trend of the K consecutive valid monitoring data after the target data segment, where M and K are preset positive integers;
[0068] The numerical rate of change is used to extract the numerical values {V1, V2, ..., V} corresponding to the M consecutive valid monitoring data points preceding the target data segment. M} and timestamp {T M1 ,T M2 ,...,T MM The time interval is calculated by taking the difference in values and the difference in timestamps between two adjacent valid monitoring data in the M consecutive valid monitoring data before the target data segment. The ratio of the value difference to the time interval is taken as the instantaneous change rate between the two adjacent data. The average of the M-1 instantaneous change rates calculated from the M consecutive valid monitoring data is taken as the value change rate of the data before the target data segment.
[0069] The periodic fluctuation characteristic is determined by retrieving monitoring data from historical periods and generating trend curves for the monitoring data in a time series. By calculating the overlap of the monitoring data fluctuations within historical time intervals, the historical time intervals in which the monitoring data shows periodic increases and decreases are located. If there are multiple historical time intervals with a high overlap, the historical time interval with the highest impact on the data fluctuation is selected as the periodic pattern, and the duration of the fluctuation cycle is determined. Based on the fluctuation cycle, the peak value, trough value, and fluctuation amplitude of the monitoring data in the historical periods within each fluctuation cycle are extracted. The cycle duration, peak value, trough value, and fluctuation amplitude are used as the periodic fluctuation characteristics of the data before the target data segment.
[0070] The historical monitoring data must cover at least one complete fluctuation cycle.
[0071] The initial trend is used to extract the values {U1, U2, ..., U...} of the K consecutive valid monitoring data after the target data segment. K} and timestamp {T K1 ,T K2 ,...,T KK A Cartesian coordinate system is established with timestamps on the horizontal axis and numerical values on the vertical axis. A regression line is constructed by effectively monitoring the distribution patterns of data points. The initial trend of change is determined based on the slope of the regression line.
[0072] If the slope is positive, the initial trend of change is upward;
[0073] If the slope is negative, the initial trend of change is downward;
[0074] If the slope is close to 0, the initial trend of change is stable.
[0075] In specific embodiments, the core role of temporal correlation features is to provide crucial temporal support for the entire data processing flow. The numerical rate of change, by quantifying the overall rate of change before the target data segment, provides a dynamic reference for judging the timeliness of correlations in the invention, ensuring that the temporal continuity of data changes is accurately reflected when calculating dependency weights. Periodic fluctuation features, by revealing historical periodic patterns, provide a benchmark for extrapolating the missing values of periodic physical quantities in subsequent data completion, ensuring that the completed missing values align with historical fluctuations. Initial change trends, by predicting the short-term direction after the target data segment, provide a trend basis for the dynamic adjustment of correlations in the invention, ensuring the adaptability of correlations over time. These three features work together to comprehensively characterize the temporal dynamics of the target data segment, directly supporting the core aspects of correlation construction, dependency weight calculation, and data completion.
[0076] Spatial correlation features include the range of numerical differences between adjacent monitoring points and the gradient characteristics of data distribution of other monitoring points located in the same region as the target data segment;
[0077] The range of numerical differences between adjacent monitoring points is determined by calculating the difference between the monitoring data values of each pair of adjacent monitoring points and taking the absolute value. The absolute difference is then statistically analyzed for all the absolute differences between adjacent monitoring points, and the minimum and maximum values are selected to form the range.
[0078] The data distribution gradient feature is obtained by acquiring the three-dimensional structural coordinates of each monitoring point in the same area, calculating the spatial straight-line distance between two adjacent monitoring points as the coordinate difference, and dividing the coordinate difference by the corresponding spatial straight-line distance to obtain the numerical change per unit distance, which is the data distribution gradient feature in the same area.
[0079] In a specific embodiment, taking a dam body monitoring area as an example, there are three adjacent monitoring points A, B, and C in this area, with three-dimensional structural coordinates A(0,0,0), B(10,0,0), and C(20,0,0) (unit: m), respectively. The dam body displacement data monitored during a certain period are A=2.0mm, B=2.2mm, and C=2.5mm, respectively. When calculating the range of numerical differences between adjacent monitoring points, the absolute difference between A and B is first calculated as |2.2-2.0|=0.2mm, and the absolute difference between B and C is calculated as |2.5-2.2|=0.3mm. After statistics, the minimum value is 0.2mm and the maximum value is 0.3mm, resulting in a numerical difference range of [0.2mm, 0.3mm]. This range can reflect the spatial consistency of monitoring data within the same area. When calculating the gradient characteristics of the data distribution, the straight-line distance between A and B is 10m, and the change in unit distance is 0.2mm / 10m = 0.02mm / m. The change in unit distance between B and C is 0.3mm / 10m = 0.03mm / m. This gradient characteristic clearly shows that the dam displacement increases from A to C in a gradient trend of 0.02-0.03mm / m. When completing the data, if the data at point B is missing, the reasonable value of point B can be calculated based on the gradient characteristics of points A and C. This provides spatial dimension support for constructing data correlation and data completion, and enhances the scientificity and reliability of monitoring data processing.
[0080] The formation of association relationships based on temporal and spatial correlation features specifically includes:
[0081] Based on the timestamps of the M consecutive valid monitoring data preceding the target data segment {T M1 ,T M2 ,...,T MM} and the timestamps of the K consecutive valid monitoring data points {T} K1 ,T K2 ,...,T KK}, generate a continuous time series, divide the continuous time series into time periods according to a preset time granularity, so that each time period corresponds to a unique timestamp;
[0082] Obtain the three-dimensional structural coordinates of each monitoring point involved in the spatial correlation features, sort the monitoring points according to the three-dimensional structural coordinates to form a spatial grid, and identify the unique three-dimensional structural coordinate position of each monitoring point in the spatial grid.
[0083] By mapping a unique timestamp to a unique three-dimensional structural coordinate position, the continuous time series after the time period is divided is dimensionally aligned with the spatial grid, forming a one-to-one association with the target data segment. The association includes the monitored values, timestamp information, monitoring points, and value change trends.
[0084] In a specific embodiment, a correlation is formed based on temporal and spatial correlation features. A continuous time series is generated by creating timestamps before and after the target data segment and dividing it into time periods according to a preset granularity. For example, if the target data segment is reservoir water level data from 10:00 to 10:30 on August 1, 2024, a continuous sequence is generated by taking the first M=3 timestamps {9:30, 9:45, 9:59} and the last K=2 timestamps {10:35, 10:50}. This sequence is then divided into time periods such as 9:30-9:45 and 9:45-10:00, with each time period corresponding to a unique timestamp. Simultaneously, the three-dimensional coordinates of the monitoring points, such as monitoring points A(10,20,5) and B(15,25,5) on the dam body, are sorted. A spatial grid is formed, giving each point a unique coordinate position. Then, by mapping a unique timestamp to the coordinate position, for example, the coordinates of point A (10, 20, 5) for the 9:30-9:45 time period, the dimensional alignment between the time series and the spatial grid is achieved. This results in a correlation that includes monitoring values, timestamp information, monitoring points, and value change trends. For example, the water level at point A is 10.2m from 9:30 to 9:45, and the water level at point A rises by 0.1m every 15 minutes. This process can accurately establish the correspondence between data in the spatiotemporal dimensions, providing a structured basis for subsequent weighted calculations and data completion. It ensures that the data correlation conforms to both the logic of the time series and the spatial distribution pattern, improving the accuracy and interpretability of the correlation.
[0085] The constraints include timeliness constraints, relevance constraints, and consistency constraints;
[0086] The timeliness constraint is used to determine the time interval threshold between the association relationship and the target data segment in the time dimension. If the difference between the timestamp of the association relationship and the timestamp of the target data segment exceeds the time interval threshold, the timeliness weight of the association relationship is reduced.
[0087] The time interval threshold is the maximum effective time difference determined based on the period duration of the missing time interval and the periodic fluctuation characteristics;
[0088] The correlation constraint is used to determine the degree of matching between the correlation relationship on the parameter attributes and the target data segment. The correlation weight is positively correlated with the degree of matching. The parameter attributes include dam displacement, reservoir water level, rainfall and monitoring equipment operating voltage.
[0089] Consistency constraints are used to determine whether the correlation in numerical change trends is consistent with the historical data change trends of the monitoring area where the target data segment is located. If there is a trend conflict, it means that the consistency weight of the correlation is reduced.
[0090] In a specific embodiment, a time interval threshold is determined based on the missing time interval and the periodic fluctuation characteristics of the periodic fluctuations through a timeliness constraint. For correlations exceeding the threshold, the timeliness weight is reduced to ensure the effectiveness of the correlation in the time dimension. A correlation constraint dynamically adjusts the correlation weight based on the matching degree of parameters such as dam displacement and reservoir water level; the higher the matching degree, the greater the weight, thus strengthening the correlation adaptability of the parameter attributes. A consistency constraint verifies the consistency between the correlation and the historical data change trend of the target data segment's region. For correlations with conflicting trends, the consistency weight is reduced to ensure the coordination of numerical change patterns. The synergistic effect of these three types of constraints can accurately screen out correlations that are superior to the target data segment in terms of time validity, parameter adaptability, and trend consistency. This provides a scientific constraint basis for weighted calculations, thereby improving the accuracy and reliability of subsequent data completion and ensuring that the cloud computing platform's processing logic for water conservancy project monitoring data is more in line with the patterns and needs of actual monitoring scenarios.
[0091] The dependency weights of the association relationships are calculated based on the constraints, specifically as follows:
[0092] Weighting coefficients are assigned to the constraints, and the formula for calculating the weighting coefficients is as follows:
[0093] W i =Cont i / (Sum cont123 (i=1, 2, 3);
[0094] Where Wᵢ represents the weight coefficient of the i-th constraint, Contᵢ represents the contribution of the i-th constraint, and Sum cont123 This represents the total contribution of the constraints, and the sum of the weight coefficients of all constraints is 1.
[0095] Calculate the score of the association under each constraint, multiply the score under each constraint by the corresponding weight coefficient, and sum them to obtain the dependency weight of the association.
[0096] The calculation of the score of the association under each constraint includes the calculation of the timeliness score, the association score, and the consistency score;
[0097] The timeliness score is calculated by setting a threshold for the time interval between the association relationship and the target data segment, and then calculating the actual time interval between the timestamp of the association relationship and the timestamp of the target data segment. The smaller the actual time interval, the higher the timeliness score.
[0098] The correlation score is calculated and analyzed to determine the correlation relationship with the parameter attributes of the target data segment. The matching degree between the two is compared through a preset parameter attribute matching table. The higher the matching degree, the higher the correlation score.
[0099] The consistency score is calculated by comparing the trend of numerical change with the trend of periodic fluctuation characteristics, and calculating the degree of trend agreement between the two. The degree of agreement is the consistency score. The more consistent the trend, the higher the consistency score.
[0100] The scores range from 0 to 1.
[0101] In a specific embodiment, by setting a weighting coefficient formula and combining it with the three types of scores—timeliness, relevance, and consistency—the degree of dependence of different relationships on the target data segment can be scientifically quantified. The weighting coefficient is determined by the contribution ratio of each constraint condition. For example, if the contribution ratios of timeliness, relevance, and consistency are 4, 3, and 3 respectively, and the total contribution ratio is 10, then the corresponding weighting coefficients are 0.4, 0.3, and 0.3 respectively. The dependence weight is obtained by weighting and summing the three types of scores with the corresponding weighting coefficients. This makes the dependence weight calculation reflect both the relative importance of each constraint condition and the degree of matching between the relationship and the target data segment through specific scores. This calculation method can accurately filter out the relationships that have the greatest impact on the target data segment. For example, if a relationship has a timeliness score of 0.9, a correlation score of 0.8, and a consistency score of 0.7, the dependency weight calculated according to the above weight coefficients is 0.9×0.4+0.8×0.3+0.7×0.3=0.81. This provides a reliable basis for subsequent data completion, anomaly detection, and other operations, ensuring that when processing massive amounts of monitoring data, the relationship can be used to improve the efficiency and accuracy of cloud computing processing. At the same time, through standardized scoring ranges and clear calculation logic, the dependency weight has interpretability and comparability, enhancing the scientific nature and stability of the entire data processing method.
[0102] The target data segment is filled in based on dependency weights, specifically as follows:
[0103] The associations are sorted by dependency weight from high to low. The temporal and spatial association features corresponding to the top N associations by dependency weight are extracted, and missing value detection is performed. The missing value detection calculation includes:
[0104] If the missing value of the target data segment is a periodic physical quantity, then take the average value of the corresponding positions in the same period before and after the missing time interval.
[0105] If the missing values of the target data segment are physical quantities that show a trend, then the increment value of the missing time interval is calculated based on the rate of change of the value.
[0106] The missing values of the target data segment are generated and filled into the missing positions of the target data segment to obtain the complete monitoring data after completion, where N is a positive integer set according to the missing length of the target data segment.
[0107] In a specific embodiment, the data completion operation on the target data segment based on dependency weights involves sorting the data by dependency weights and extracting the temporal and spatial correlation features of the top N relationships. This is combined with targeted calculations based on the physical quantity type of the missing values. For example, for periodic reservoir water level missing values, the average water level of the corresponding time period before and after the missing interval is used. For dam displacement missing values showing a trend, the cumulative displacement increment for the missing period is calculated based on the rate of change. This accurately generates missing values that conform to the data patterns and fills them into the target data segment. This approach leverages the strong correlation of relationships to ensure the rationality of missing values while improving completion accuracy by differentiating between periodic and trend-based physical quantities. It effectively solves the problem of incomplete monitoring data caused by equipment failures, transmission interruptions, etc., providing a complete and reliable monitoring data foundation for subsequent feature extraction, relationship construction, and water conservancy project condition analysis. Furthermore, the N value is dynamically set according to the missing length to balance completion efficiency and accuracy, ensuring that the cloud computing platform's processing results for water conservancy project monitoring data better meet actual project needs.
[0108] Example 2: A cloud-based water conservancy project monitoring data processing system, including a preprocessing module, an analysis module, and a completion module;
[0109] The preprocessing module is used to preprocess the received monitoring data to obtain the target data segment;
[0110] The analysis module is used to analyze the target data segment, obtain temporal correlation features and spatial correlation features, and form correlation relationships based on the temporal correlation features and spatial correlation features;
[0111] The completion module is used to set constraints, calculate the dependency weights of the relationships based on the constraints, and complete the target data segment based on the dependency weights to obtain the complete monitoring data.
[0112] This method preprocesses the received monitoring data to obtain target data segments, providing well-organized basic data for subsequent analysis. By analyzing the target data segments, temporal and spatial correlation characteristics are obtained and correlations are formed, constructing the intrinsic connections between data. By setting timeliness, correlation, and consistency constraints, the dependency weights of the correlations are calculated based on the constraints. Then, the target data segments are completed according to the dependency weights, generating and filling in missing values to obtain complete monitoring data. Ultimately, this method significantly improves the completeness, accuracy, and reliability of monitoring data, enabling the output complete monitoring data to more accurately support the assessment of water conservancy project conditions, safety early warning, and scientific decision-making.
[0113] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program code. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Red-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0114] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A cloud computing-based method for processing water conservancy project monitoring data, characterized in that: Includes the following steps: Step S1: Preprocess the received monitoring data to obtain the target data segment; Step S2: Analyze the target data segment to obtain temporal correlation features and spatial correlation features, and form a correlation relationship based on the temporal correlation features and spatial correlation features; Step S3: Set constraints, calculate the dependency weights of the association based on the constraints, and complete the target data segment using the dependency weights to obtain complete monitoring data; The spatial correlation features include the range of numerical differences and data distribution gradient features of adjacent monitoring points of other monitoring points located in the same region as the target data segment; The range of numerical differences between adjacent monitoring points is obtained by calculating the difference between the monitoring data values of each pair of adjacent monitoring points and taking the absolute value. The absolute differences calculated from the numerical differences of all adjacent monitoring points are statistically analyzed, and the minimum and maximum values are selected to form the range. The data distribution gradient feature is obtained by acquiring the three-dimensional structural coordinates of each monitoring point in the same area, calculating the spatial straight-line distance between two adjacent monitoring points as the coordinate difference, and dividing the coordinate difference by the corresponding spatial straight-line distance to obtain the numerical change within a unit distance, which is the data distribution gradient feature in the same area. The formation of association relationships based on the aforementioned temporal and spatial correlation features specifically includes: Based on the timestamps of the M consecutive valid monitoring data preceding the target data segment {T M1 ,T M2 ,...,T MM } and the timestamps of the K consecutive valid monitoring data points {T} K1 ,T K2 ,...,T KK }, generate a continuous time series, divide the continuous time series into time periods according to a preset time granularity, so that each time period corresponds to a unique timestamp; The three-dimensional structural coordinates of each monitoring point involved in the spatial correlation feature are obtained, and the monitoring points are sorted according to the three-dimensional structural coordinates to form a spatial grid. Each monitoring point in the spatial grid is identified with a unique three-dimensional structural coordinate position. By mapping the unique timestamp to the unique three-dimensional structural coordinates, the continuous time series after the time period is divided is dimensionally aligned with the spatial grid to form a one-to-one association with the target data segment. The association includes the monitored values, timestamp information, monitoring points, and value change trends.
2. The cloud computing-based water conservancy project monitoring data processing method as described in claim 1, characterized in that: Preprocessing of the monitoring data includes removing outliers and removing duplicates. The monitoring data includes dam displacement data, reservoir water level data, rainfall data, and the operating voltage of the monitoring equipment. The outlier removal is used to remove monitoring data that exceeds a preset reasonable range; The deduplication feature is used to remove multiple identical monitoring data points collected by the same monitoring device at the same time point, while retaining only one identical monitoring data point. The preprocessed monitoring data is marked for validity to obtain valid monitoring data. The valid monitoring data is compared with the theoretical monitoring sequence, and the proportion of valid monitoring data in each time period is calculated. Time periods whose proportions do not reach the preset complete threshold are marked as target data segments. The target data segments include missing time nodes, missing time intervals, and missing monitoring points. The theoretical monitoring sequence is theoretical monitoring data collected according to a preset acquisition frequency.
3. The cloud computing-based water conservancy project monitoring data processing method as described in claim 2, characterized in that: The time-related features include the rate of change of the values of the M consecutive valid monitoring data before the target data segment, the periodic fluctuation characteristics, and the initial change trend of the K consecutive valid monitoring data after the target data segment, where M and K are preset positive integers; The numerical change rate is used to extract the numerical values {V1, V2, ..., V} corresponding to the M consecutive valid monitoring data points preceding the target data segment. M } and timestamp {T M1 ,T M2 ,...,T MM The time interval is calculated by taking the numerical difference and timestamp difference between two adjacent valid monitoring data in the M consecutive valid monitoring data before the target data segment. The ratio of the numerical difference to the time interval is taken as the instantaneous change rate between the two adjacent data. The average value of the M-1 instantaneous change rates calculated from the M consecutive valid monitoring data is taken as the numerical change rate of the data before the target data segment. The periodic fluctuation feature is obtained by retrieving monitoring data from historical periods, generating trend curves for the monitoring data in a time series, calculating the overlap of monitoring data fluctuations within historical time intervals, locating historical time intervals where the monitoring data shows periodic increases and decreases, and if there are multiple historical time intervals with a high overlap, selecting the historical time interval with the highest impact on data fluctuations as the periodic pattern, determining the duration of the fluctuation cycle, and based on the fluctuation cycle, extracting the peak, trough, and fluctuation amplitude of the monitoring data in historical periods within each fluctuation cycle, and using the cycle duration, peak, trough, and fluctuation amplitude as the periodic fluctuation feature of the data before the target data segment; The time range of the monitoring data for the historical period must cover at least one complete fluctuation cycle; The initial trend is used to extract the numerical values {U1, U2, ..., U} corresponding to K consecutive valid monitoring data after the target data segment is extracted. K } and timestamp {T K1 ,T K2 ,...,T KK A Cartesian coordinate system is established with timestamps on the horizontal axis and numerical values on the vertical axis. A regression line is constructed by effectively monitoring the distribution patterns of data points. The initial trend of change is determined based on the slope of the regression line. If the slope is positive, then the initial trend of change is upward; If the slope is negative, the initial trend of change is downward; If the slope is close to 0, the initial trend of change is stable.
4. The cloud computing-based water conservancy project monitoring data processing method as described in claim 3, characterized in that, The set constraints include timeliness constraints, relevance constraints, and consistency constraints; The timeliness constraint is used to determine the time interval threshold between the association relationship and the target data segment in the time dimension. If the difference between the timestamp of the association relationship and the timestamp of the target data segment exceeds the time interval threshold, the timeliness weight of the association relationship is reduced. The time interval threshold is the maximum effective time difference determined based on the missing time interval and the period duration of the periodic fluctuation characteristics; The correlation constraint is used to determine the matching degree between the correlation relationship on the parameter attributes and the target data segment. The correlation weight is positively correlated with the matching degree. The parameter attributes include dam displacement, reservoir water level, rainfall and monitoring equipment operating voltage. The consistency constraint is used to determine whether the correlation in terms of numerical change trend is consistent with the historical data change trend of the monitoring area where the target data segment is located. If there is a trend conflict, it means that the consistency weight of the correlation is reduced.
5. The cloud computing-based water conservancy project monitoring data processing method as described in claim 4, characterized in that, The dependency weights of the association relationships are calculated based on the constraints, specifically as follows: Weighting coefficients are assigned to the constraints, and the formula for calculating the weighting coefficients is as follows: W i =Cont i / (Sum cont123 )(i=1,2,3); Where Wᵢ represents the weight coefficient of the i-th constraint, Contᵢ represents the contribution of the i-th constraint, and Sum cont123 This represents the total contribution of the constraints, and the sum of the weight coefficients of all constraints is 1. Calculate the score of the association under each constraint, multiply the score under each constraint by the corresponding weight coefficient, and sum the results to obtain the dependency weight of the association.
6. The cloud computing-based water conservancy project monitoring data processing method as described in claim 5, characterized in that, The calculation of the score of the association under each constraint includes the calculation of timeliness score, association score and consistency score; The timeliness score calculation sets a threshold for the time interval between the association relationship and the target data segment, and calculates the actual time interval between the timestamp of the association relationship and the timestamp of the target data segment. The smaller the actual time interval, the higher the timeliness score. The correlation score is calculated and analyzed to determine the correlation relationship and the parameter attributes of the target data segment. The matching degree between the two is compared through a preset parameter attribute matching table. The higher the matching degree, the higher the correlation score. The consistency score calculation compares the trend of the numerical change with the trend of periodic fluctuation characteristics, and calculates the degree of trend consistency between the two. The degree of consistency is the consistency score. The more consistent the trends are, the higher the consistency score. The scores for calculating the association under each constraint are all between 0 and 1.
7. The cloud computing-based water conservancy project monitoring data processing method as described in claim 6, characterized in that, The target data segment is augmented according to the dependency weight, specifically as follows: The associations are sorted from high to low according to their dependency weights. The temporal and spatial association features corresponding to the top N associations by dependency weight are extracted, and missing value judgment is performed. The missing value judgment calculation includes: If the missing value of the target data segment is a periodic physical quantity, then the average value of the corresponding positions in the same period before and after the missing time interval is taken. If the missing values of the target data segment are physical quantities that show a trend, then the increment value of the missing time interval is calculated according to the rate of change of the value. The missing values of the target data segment are generated and filled into the missing positions of the target data segment to obtain the complete monitoring data after completion, where N is a positive integer set according to the missing length of the target data segment.
8. A cloud computing-based water conservancy project monitoring data processing system, characterized in that, It includes a preprocessing module, an analysis module, and a completion module; The preprocessing module is used to preprocess the received monitoring data to obtain the target data segment; The analysis module is used to analyze the target data segment, obtain temporal correlation features and spatial correlation features, and form a correlation relationship based on the temporal correlation features and the spatial correlation features; The completion module is used to set constraints, calculate the dependency weight of the association based on the constraints, and perform data completion operation on the target data segment based on the dependency weight to obtain the complete monitoring data after completion. The spatial correlation features include the range of numerical differences and data distribution gradient features of adjacent monitoring points of other monitoring points located in the same region as the target data segment; The range of numerical differences between adjacent monitoring points is obtained by calculating the difference between the monitoring data values of each pair of adjacent monitoring points and taking the absolute value. The absolute differences calculated from the numerical differences of all adjacent monitoring points are statistically analyzed, and the minimum and maximum values are selected to form the range. The data distribution gradient feature is obtained by acquiring the three-dimensional structural coordinates of each monitoring point in the same area, calculating the spatial straight-line distance between two adjacent monitoring points as the coordinate difference, and dividing the coordinate difference by the corresponding spatial straight-line distance to obtain the numerical change within a unit distance, which is the data distribution gradient feature in the same area. The formation of association relationships based on the aforementioned temporal and spatial correlation features specifically includes: Based on the timestamps of the M consecutive valid monitoring data preceding the target data segment {T M1 ,T M2 ,...,T MM } and the timestamps of the K consecutive valid monitoring data points {T} K1 ,T K2 ,...,T KK }, generate a continuous time series, divide the continuous time series into time periods according to a preset time granularity, so that each time period corresponds to a unique timestamp; The three-dimensional structural coordinates of each monitoring point involved in the spatial correlation feature are obtained, and the monitoring points are sorted according to the three-dimensional structural coordinates to form a spatial grid. Each monitoring point in the spatial grid is identified with a unique three-dimensional structural coordinate position. By mapping the unique timestamp to the unique three-dimensional structural coordinates, the continuous time series after the time period is divided is dimensionally aligned with the spatial grid to form a one-to-one association with the target data segment. The association includes the monitored values, timestamp information, monitoring points, and value change trends.
Citation Information
Patent Citations
Data analysis processing method and system based on water conservancy project
CN119647783A
Geological disaster monitoring, prediction and early warning method based on artificial intelligence
CN119785535A
GIS-based geological disaster monitoring point data expression method and system
CN119961345A