Water conservancy project monitoring data pseudo-correlation identification method and system
By employing a multi-dimensional pseudo-correlation identification method, combined with frequency domain multi-scale decomposition and local geometric feature analysis, the pseudo-correlation problem in the correlation analysis of water conservancy project monitoring data was solved, thereby improving the accuracy and reliability of the correlation analysis of monitoring data.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NANJING HYDRAULIC RES INST
- Filing Date
- 2026-04-13
- Publication Date
- 2026-05-15
AI Technical Summary
Existing technologies struggle to effectively isolate environmental co-driving interference, ignore spatially distant weak coupling, and address cross-structural segmentation effects in the correlation analysis of water conservancy project monitoring data. This leads to the generation of spurious correlations and affects the reliability of the monitoring data correlation analysis.
By constructing a multi-dimensional pseudo-correlation identification method, and combining time-statistical, spatial, structural, and correlation propagation consistency constraints, pseudo-correlation relationships are identified, including frequency domain multi-scale decomposition, local geometric feature analysis, and physical spatial distance correction, thereby improving the accuracy of monitoring data correlation analysis.
It improves the ability to identify spurious correlations driven by environmental factors, accurately restores the physical transmission differences of hydraulic structures, blocks false transmission chains, and enhances the authenticity and reliability of topological inference of monitoring and correlation networks.
Smart Images

Figure CN122046281A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of water conservancy project safety monitoring and data analysis technology, and in particular to a method and system for identifying pseudo-correlation in water conservancy project monitoring data. Background Technology
[0002] Modern water conservancy projects deploy a massive number of sensors. Analyzing the correlation of time series data from multiple sources and constructing a monitoring correlation network is the technical foundation for anomaly diagnosis and structural risk tracking. Accurately extracting data correlations with real mechanical relationships determines the reliability of safety monitoring models and physical inversion early warning systems.
[0003] Current monitoring data correlation analysis largely relies on time-domain statistical indicators. Typical existing techniques directly calculate the global Pearson correlation coefficient between data points throughout the entire observation period, supplemented by conventional isotropic Euclidean distance for correlation threshold screening. In terms of correlation network inference, existing methods generally follow the direct transmission assumption of statistical correlation, reconstructing the overall monitoring topology of the engineering structure through the direct splicing and expansion of strongly correlated edges between nodes.
[0004] Existing technologies face three main challenges when dealing with complex coupled environments: environmental co-driving interference, anisotropic attenuation distortion, and statistical pseudo-transmission chains. Specifically: First, existing global time-domain statistical methods struggle to effectively isolate the high correlations generated by periodic environmental fluctuations such as temperature and water level, easily misclassifying measurement points without physical structural connections as genuine couplings. Second, conventional isotropic spatial models neglect the stiffness differences of hydraulic entities along the principal axis, normal, and vertical direction, causing theoretical spatial attenuation expectations to deviate significantly from actual mechanical transmission laws. Third, statistical correlation coefficients lack direct transferability mathematically; existing technologies, when constructing indirect correlations, lack the natural attenuation penalty for long-span physical distances, easily generating pseudo-correlation transmission chains that violate spatial barrier laws. Summary of the Invention
[0005] Purpose of the invention: To provide a method and system for identifying spurious correlations in water conservancy project monitoring data, in order to solve the problem that spurious correlations are prone to occur in the correlation analysis of monitoring data due to environmental co-driving, spatial long-distance weak coupling, cross-structure segmentation influence, and indirect propagation chain effects in the existing technology.
[0006] Technical solution: A method for identifying pseudo-correlation in water conservancy project monitoring data, comprising:
[0007] Acquire monitoring data and spatial attribute information of monitoring points in water conservancy projects. Spatial attribute information includes the spatial coordinates of the monitoring points and the segmentation information of the engineering structure.
[0008] The monitoring data is preprocessed to obtain a standardized time series, and a spatial topology of the monitoring points is constructed based on the spatial coordinates of the monitoring points.
[0009] Correlation analysis was performed on standardized time series to identify candidate correlations and construct an initial monitoring association network, and time-statistical correlation features of candidate correlations were extracted.
[0010] Based on the spatial distance calculated from the spatial topology and spatial coordinates of the monitoring points, the spatial inconsistency characteristics of candidate correlations deviating from the spatial propagation law are calculated.
[0011] Based on the segmentation information of the engineering structure, structural inconsistency features that cross structural segments are extracted from candidate correlations;
[0012] Based on the spatial topology of the monitoring points, search for indirect association paths between candidate related monitoring point pairs, and calculate the inconsistency features of association transmission on the indirect association paths;
[0013] By fusing time-statistical correlation features, spatial inconsistency features, structural inconsistency features, and association transmission inconsistency features from multiple dimensions, a pseudo-correlation risk index is calculated, and pseudo-correlation relationships in candidate correlation relationships are identified based on this index, thereby updating the monitoring correlation network.
[0014] Among them, when the statistical correlation between monitoring points deviates significantly from the spatial propagation law, structural segmentation constraints, or association transmission logic, the corresponding candidate correlation is determined to be a pseudo-correlation.
[0015] Compared with methods based solely on statistical correlation analysis, this invention achieves pseudo-correlation identification through multi-dimensional consistency constraints such as time-statistics, space, structure, and correlation transitivity.
[0016] Beneficial Effects: To address the issue of spurious coupling caused by environmental co-driving interference, this solution introduces a frequency domain multi-scale decomposition mechanism. By performing multi-scale decomposition on the time series of monitoring data corresponding to candidate correlation edges, a frequency band correlation contribution spectrum is constructed. Combined with a preset set of environmental characteristic frequency bands, the concentration and dispersion of environmental frequency band correlations are extracted. This identifies statistically high correlations induced by fluctuations in external environmental factors such as temperature and water level, thereby improving the ability to identify environmentally driven spurious correlations. This solves the technical problem of distinguishing between real physical coupling and environmentally driven spurious correlations in global time-domain statistical analysis, improving the reliability of monitoring data correlation analysis results.
[0017] To address the attenuation distortion problem caused by isotropic models, this scheme constructs a structural coordinate system based on local engineering geometric features. By projecting the position difference vector of monitoring points onto the principal axis, normal direction, and vertical direction, and combining it with anisotropic attenuation parameters for normalization and synthesis, the directional physical transmission differences caused by the non-uniform stiffness distribution of hydraulic structures are accurately reproduced. This solves the problem of conventional Euclidean distance predictions deviating significantly from actual mechanical transmission laws, improving the mechanical rationality of spatial attenuation predictions.
[0018] To address the issue of spurious transmission chains caused by the lack of direct transitivity in statistical correlation coefficients, this scheme forcibly introduces a natural attenuation correction factor based on the physical spatial distance between the first and last nodes in the deviation calculation of indirect association paths. This approach, starting from physical distance constraints, attenuates and corrects the transmission expectation constructed based on statistical correlation, thus blocking the transmission of statistical spurious correlations that violate spatial attenuation laws under long-span physical barriers. This solves the problem of existing technologies easily generating spurious transmission chains, improving the authenticity and reliability of topological inference in monitored association networks. Attached Figure Description
[0019] Figure 1 This is a diagram illustrating the overall framework for multi-dimensional pseudo-correlation identification in the embodiments of this application.
[0020] Figure 2 This is a flowchart illustrating the steps of the frequency domain multi-scale decomposition and extraction method for time-statistical correlation features in the embodiments of this application.
[0021] Figure 3 This is a flowchart illustrating the steps of the anisotropic attenuation expectation calculation method based on the local structural coordinate system in the embodiments of this application.
[0022] Figure 4 This is a flowchart illustrating the steps of the penalty constraint mechanism across engineering structure segments in this application embodiment. Detailed Implementation
[0023] Example 1, such as Figure 1 As shown in the figure, this embodiment elaborates on the overall framework for multi-dimensional pseudo-correlation identification and the monitoring of the correlation network update process.
[0024] Step 101: Obtain monitoring data and spatial attribute information of the monitoring points of the water conservancy project. The spatial attribute information includes the spatial coordinates of the monitoring points and the segmentation information of the engineering structure.
[0025] Specifically, water conservancy projects continuously monitor the structural status of their structures through a safety monitoring system during long-term operation. The system continuously receives data from monitoring points on structures such as dams, water conveyance channels, or tunnels. The monitoring data specifically covers various types of physical monitoring time series, including structural deformation, seepage pressure, temperature, and strain. Simultaneously, it acquires the spatial attribute information corresponding to each monitoring point. The spatial coordinates of the monitoring points are represented using a three-dimensional coordinate system to indicate the physical location of each sensor. The structural segmentation information records the engineering management or structural unit to which the monitoring points belong, specifically as dam segment numbers, channel segment divisions, or tunnel segment identifiers. This type of basic data provides objective physical boundary conditions for subsequent spatial and structural consistency verification.
[0026] Step 102: Preprocess the monitoring data to obtain a standardized time series, and construct the spatial topology of the monitoring points based on their spatial coordinates.
[0027] In this embodiment, the raw monitoring data collected on-site has discontinuous time series or inconsistent dimensions, requiring cleaning and alignment. The preprocessing specifically includes filling in missing values using linear interpolation or spline interpolation, and resampling into an equally spaced time series. Subsequently, the Z-score standardization method is used to standardize the series, eliminating differences in different physical dimensions, thereby obtaining a standardized time series with a mean of zero and a variance of one.
[0028] Furthermore, regarding the construction of the spatial topology of the monitoring points, the system calculates the Euclidean spatial distance between monitoring points based on their spatial coordinates. A preset spatial neighborhood threshold is obtained. It is then determined whether the spatial distance is less than or equal to the preset threshold; if so, the corresponding two monitoring points are considered to have a spatial neighborhood relationship. Monitoring points with spatial neighborhood relationships are connected as edges to generate a spatial topology reflecting the spatial proximity attributes of the project. The spatial topology of the monitoring points is represented as an undirected graph containing a set of nodes and a set of edges, used to characterize the local physical connectivity of the project entity in space. In some implementations, the preset spatial neighborhood threshold can be selected according to the project scale; for large-scale water conservancy projects, it can be set to 20 to 50 meters, and for small and medium-sized projects, it can be set to 5 to 20 meters.
[0029] Step 103: Perform correlation analysis on the standardized time series, identify candidate correlations and construct an initial monitoring association network, and extract the time-statistical correlation features of the candidate correlations.
[0030] In some embodiments, the system calculates the global Pearson correlation coefficient between the standardized time series corresponding to each pair of monitoring points. A preset correlation threshold is set; when the absolute value of the global Pearson correlation coefficient between two monitoring points is greater than or equal to the preset threshold, a preliminary correlation is determined between the two monitoring points, and they are extracted as candidate correlations. The preset correlation threshold can be set between 0.3 and 0.6, with the specific value depending on the signal-to-noise ratio of the monitoring data and the sensitivity requirements of the project. Subsequently, the absolute value of the global Pearson correlation coefficient is used as a time-statistical correlation feature. This step uses basic mathematical statistics methods to screen node pairs with synchronous fluctuation characteristics in data performance, forming the basic sample pool for multidimensional diagnosis. Simultaneously, the system records the global Pearson correlation coefficient corresponding to each candidate correlation as the actual correlation strength of each monitoring point pair, for subsequent spatial deviation measurement, structural consistency verification, and association propagation analysis.
[0031] Step 104: Based on the spatial distance calculated from the spatial topology and spatial coordinates of the monitoring points, calculate the spatial inconsistency features of the candidate correlations that deviate from the spatial propagation law.
[0032] Specifically, the actual structural response exhibits a decay trend with increasing physical distance. The system uses spatial distance to establish a theoretical expectation of spatial decay and compares it with the actual correlation of candidate relationships. By calculating the deviation of the actual correlation from the theoretical expectation of distance decay, a spatial inconsistency feature is calculated. This feature is specifically used to identify anomalous association pairs that exhibit high correlation despite being geographically distant.
[0033] Step 105: Based on the segmentation information of the engineering structure, extract the structural inconsistency features of candidate correlations that cross structural segments.
[0034] In this embodiment, the system verifies the engineering units to which the nodes at both ends of the candidate correlation belong. For node pairs that cross different physical defense lines or independent operating units, the system applies a specific penalty coefficient. After cross-segment constraint correction, structural inconsistency features are extracted, thereby eliminating cross-domain correlation phenomena caused by coincidence.
[0035] Step 106: Search for indirect association paths between candidate related monitoring point pairs based on the spatial topology of the monitoring points, and calculate the inconsistency features of association transmission on the indirect association paths.
[0036] For example, the system utilizes the previously generated undirected graph structure to search for connected paths between nodes that include intermediate nodes. By comparing the theoretical expectation of segmented propagation with the actual performance of the first and last nodes, inconsistencies in association propagation are extracted. This mechanism is used to identify chain-like spurious correlation errors caused by indirect bridges.
[0037] Step 107: Multi-dimensionally integrate time-statistical correlation features, spatial inconsistency features, structural inconsistency features, and association transmission inconsistency features to calculate the pseudo-correlation risk index, and identify pseudo-correlation relationships among candidate correlation relationships accordingly, and update the monitoring correlation network.
[0038] In this embodiment, pre-configured first, second, third, and fourth weighting coefficients are obtained. The time-statistical correlation feature is multiplied by the first weighting coefficient, the spatial inconsistency feature by the second weighting coefficient, the structural inconsistency feature by the third weighting coefficient, and the association transmission inconsistency feature by the fourth weighting coefficient. The weighted sum of these products yields the spurious correlation risk index.
[0039] Specifically, the formula for calculating the pseudo-correlation risk index satisfies:
[0040] F pseudo =w1*V time +w2*G space +w3*R struct +w4*T trans ;
[0041] Among them, F pseudo V is a pseudo-correlation risk index. time G space R struct and T trans These correspond to time-statistical correlation characteristics, spatial inconsistency characteristics, structural inconsistency characteristics, and correlation transmission inconsistency characteristics, respectively. w1, w2, w3, and w4 are the first, second, third, and fourth weighting coefficients, respectively, and satisfy the constraint that the sum of all weighting coefficients is one. Specifically, the preset risk threshold can be set between 0.5 and 0.7. In engineering practice, the typical configuration of the first to fourth weighting coefficients is w1=0.25, w2=0.30, w3=0.20, and w4=0.25, which can be adjusted according to the data quality and structural complexity at the engineering site.
[0042] Furthermore, when the spurious correlation risk index exceeds a preset risk threshold, the corresponding candidate correlation is determined to be a spurious correlation. Based on the determination result, the system cleans up the connection topology within the project, specifically by removing edges determined to be spuriously correlated from the original monitoring and correlation network, or proportionally reducing the connection weight of such edges. The updated monitoring and correlation network is represented as a cleaned adjacency matrix, which is output to downstream modules for tracking structural anomalies and locating risks.
[0043] Example 2: This example details the basic scheme for global extraction of time-statistical correlation features.
[0044] Step 201: Calculate the global Pearson correlation coefficient between the standardized time series corresponding to the candidate correlations.
[0045] In this embodiment, for each extracted candidate correlation, the system traverses the entire monitoring timeline and calculates the global Pearson correlation coefficient between the time series corresponding to the two monitoring points. This calculation process calculates the degree of linear correlation between the two variables over the entire observation period by statistically analyzing the ratio of the product of their covariance and standard deviation over the global time range.
[0046] Step 202: The absolute value of the global Pearson correlation coefficient is used as a time-statistical correlation feature to characterize the overall statistical correlation of candidate correlations.
[0047] Furthermore, the system maps the numerical outputs calculated at the global level to indicators for use by subsequent multi-dimensional fusion modules. In conventional monitoring scenarios where environmental factors change gradually and structural responses are simple, this basic solution features low computational overhead and a simple implementation process.
[0048] Example 3, such as Figure 2 As shown in the figure, this embodiment elaborates on the frequency domain multi-scale decomposition extraction method of time-statistical correlation features, as well as the complete process of environmental feature diagnosis using frequency band concentration and dispersion.
[0049] Step 301: Use orthogonal wavelets to perform multi-scale decomposition on the standardized time series corresponding to the candidate correlation to obtain the approximate components of each monitoring point and the detail components of multiple frequency bands.
[0050] Specifically, after acquiring the standardized time series, the system performs discrete wavelet transform processing using Daubechies orthogonal wavelets. The Daubechies orthogonal wavelet can specifically be a Daubechies 4th order db4 basis function. Based on fundamental signal processing principles, multi-scale decomposition separates the original broadband time-domain signal into different frequency bands. After decomposition, two types of components are output: one is an approximate component, used to characterize the ultra-low frequency slowly varying trend components in the original signal with frequencies below a set threshold; the other is multiple detail components, used to characterize the high-frequency and mid-frequency fluctuation components in the original signal that fall within specific frequency ranges.
[0051] Step 302: Calculate the correlation contribution of candidate correlations at the approximate component and each detail component level, and construct the band correlation contribution spectrum.
[0052] In this embodiment, since the standardized time series has a variance of one and a mean of zero after preprocessing, the global Pearson correlation coefficient of the time series of the two monitoring points is numerically equal to their covariance. Utilizing the strictly orthogonal property of orthogonal wavelets, the sum of the cross products between components at different scales is zero. Therefore, the global correlation coefficient can be expanded without error into the sum of independently calculated inner products within each frequency band. The system calculates the inner product of the time series of the approximate components of the two monitoring points to obtain the correlation contribution of the 0th frequency band; simultaneously, it calculates the inner product of the time series of each detail component at the same scale to obtain the detail correlation contribution of each frequency band.
[0053] The principle of additive decomposition of the global correlation coefficient in terms of frequency bands is calculated using the following formula:
[0054] ρ ik =C ik 0 +∑ j=1 J C ik j ;
[0055] Where, ρ ik Let C be the global Pearson correlation coefficient between monitoring point i and monitoring point k. ik 0 C represents the correlation contribution of the approximate component. ik j denoted as the correlation contribution of the detail component at layer j, where J is the total number of layers in the multi-scale decomposition.
[0056] Subsequently, the system arranges and combines the calculated correlation contributions of each frequency band into a vector form to construct a frequency band correlation contribution spectrum that characterizes the physical structure of the correlation distribution in the frequency domain.
[0057] Step 303: Based on the distribution characteristics of the frequency band correlation contribution in the frequency band correlation contribution spectrum, extract the quantitative index that characterizes the degree of concentration of frequency domain correlation, and use the quantitative index as a time-statistical correlation feature to identify pseudo-correlation patterns driven by environmental factors in candidate correlation relationships.
[0058] After obtaining the frequency band correlation contribution spectrum, the system extracts the energy proportion of each frequency band component in the global correlation, and then maps it to generate computational feature parameters for multidimensional consistency fusion.
[0059] Furthermore, an implementation method for pre-constructing an environmental feature mask is provided, specifically including the following steps.
[0060] Step 304: Obtain the characteristic period of known environmental factors and the sampling interval of the standardized time series.
[0061] Specifically, the monitoring data of the engineering structure is affected by the periodic changes in the external environment. The system pre-acquires the characteristic cycles of known environmental factors in the area where the project is located, including annual temperature cycles and seasonal reservoir water level adjustments. Simultaneously, the system acquires the recorded sampling interval parameters.
[0062] Step 305: Calculate the wavelet decomposition layer number corresponding to each known environmental factor based on the ratio of the characteristic period to the sampling interval.
[0063] The system uses logarithmic operations to map the characteristic period of the physical time domain to the decomposition level of the wavelet transform domain. By calculating the dimensionless ratio of the characteristic period to the sampling interval, it locates the frequency band where the environmental fluctuation energy is mainly concentrated.
[0064] Calculate the corresponding wavelet decomposition layer number j m The formula is:
[0065] j m =floor(log2(P env,m / Δt));
[0066] Among them, P env,m Let be the characteristic period of the m-th known environmental factor, Δt be the sampling interval, floor(·) be the floor function, and j m This is the calculated corresponding frequency band layer number.
[0067] For example, assuming the sampling interval is set to 1 day and the characteristic period of a certain environmental factor is measured to be 365 days, substituting into the above formula, the base-2 logarithm of 365 is approximately 8.51. After rounding down, the output wavelet decomposition layer number is 8. This value indicates that the energy of a specific physical annual cycle phenomenon is concentrated in the 8th layer detail component.
[0068] Step 306: Merge the frequency bands corresponding to the extracted wavelet decomposition layer numbers to construct a pre-configured set of environmental feature frequency bands.
[0069] The system iterates through all known environmental factors and aggregates the wavelet decomposition layer numbers obtained through mapping calculation into a discrete set, forming a pre-configured set of environmental feature frequency bands.
[0070] In an alternative solution, in response to the long-term gradual climate change trend of large-scale water conservancy projects, the system synchronously incorporates the 0th frequency band corresponding to the approximate component representing the ultra-low frequency trend into the above set to form an expanded feature mask set, thereby preventing the omission of ultra-long-term gradual change driving components.
[0071] Step 307: Obtain the pre-configured set of environmental feature frequency bands.
[0072] After completing the preliminary construction, the online computing module reads the pre-configured set of environmental feature frequency bands and uses it as a data index scale for extracting strongly correlated environmental components.
[0073] Step 308: Calculate the ratio of the sum of the absolute values of the correlation contribution of the frequency band correlation contribution spectrum within the set of environmental characteristic frequency bands to the absolute value of the total correlation contribution, and obtain the environmental frequency band correlation concentration.
[0074] The system extracts the contribution values of all frequency bands that hit the pre-configured environmental feature frequency band set index from the frequency band correlation contribution spectrum. To eliminate the influence of positive and negative phases, the system takes the absolute value of each extracted frequency band contribution and sums them up. The ratio of the sum of absolute contributions within the environmental frequency band to the sum of the absolute values of the total contributions including all components is the environmental frequency band correlation concentration.
[0075] Specifically, the environmental frequency band correlation concentration η ik The calculation formula is:
[0076] η ik =(∑ j∈E |C ik j |) / (∑ j=0 J |C ik j |);
[0077] Where, η ik Let E be the concentration of environmental frequency band correlation, and C be the pre-configured set of environmental characteristic frequency bands. ik j Let J be the correlation contribution of the j-th frequency band, where J is the highest frequency band number and 0 is the frequency band number corresponding to the approximate component.
[0078] Step 309: Calculate the frequency band dispersion of the correlation contribution based on the proportion of the absolute value of the correlation contribution of each frequency band in the frequency band correlation contribution spectrum.
[0079] The system utilizes the information entropy mechanism to calculate the distribution characteristics of relevant components at different decomposition scales. The system calculates the percentage probability of the absolute value of the contribution of each frequency band within the total absolute value, and substitutes this probability value into the Shannon entropy formula to calculate the dispersion of the relevant contribution frequency bands.
[0080] Correspondingly, the relevant contribution frequency band dispersion H ik The calculation formula is:
[0081] H ik =-∑ j=0 J [p ikj *ln(p ikj )];
[0082] Among them, H ik To contribute to the frequency band dispersion, p ikj p represents the proportion of the absolute value of the correlation contribution of the j-th frequency band to the absolute value of the total correlation contribution. ikj =|C ikj | / (∑ l=0 J |C ik l |), C ikj This represents the correlation contribution of the j-th frequency band. Where, it is stipulated that when p... ikj When = 0, the corresponding term p ikj *ln(p ikj The value of p is zero, which conforms to the standard convention of information entropy. Those skilled in the art can first check p in numerical implementation. ikj The conventional method for protecting the logarithm when it is greater than 0.
[0083] Step 310: Obtain the preset main index weight coefficients, and perform weighted fusion of the environmental frequency band correlation concentration and the normalized correlation contribution frequency band dispersion based on the main index weight coefficients to calculate the frequency domain pseudo-correlation diagnostic score.
[0084] After obtaining the concentration and dispersion indicators, the system extracts the preset weight parameters and performs linear weighted fusion calculation on the feature parameters of the two independent dimensions.
[0085] Frequency domain pseudo-correlation diagnostic score F ik The calculation formula is:
[0086] F ik =γ*η ik +(1-γ)*(1-H ik / ln(J+1));
[0087] Among them, F ik For frequency domain pseudo-correlation diagnostic scores, γ is the main index weighting coefficient, and η is the frequency domain pseudo-correlation diagnostic score. ik H represents the concentration of environmental frequency band correlation. ik The relevant contribution frequency band dispersion is represented by ln(J+1), which is the normalized extreme value parameter of the information entropy theory.
[0088] The main index weight coefficient γ∈(0,1) is used to adjust the relative contributions of the two types of indicators: environmental frequency band concentration and related contribution frequency band dispersion. When the periodic excitation characteristics of the engineering site environment are significant (such as large annual temperature fluctuations and strong reservoir water level scheduling regularity), the value of γ can be appropriately increased to strengthen the weight of the environmental concentration index; when the multi-band response characteristics of the engineering structure are obvious, the value of γ can be appropriately decreased. In an optional implementation, the value of γ is 0.6. Those skilled in the art can evaluate F based on historical monitoring data according to the environmental excitation characteristics and structural response patterns of the actual engineering project. ik To improve the accuracy of identifying known spurious correlations, a suitable value for γ is determined using conventional methods such as cross-validation.
[0089] Step 311: Use the frequency domain pseudo-correlation diagnostic score as a time-statistical correlation feature.
[0090] The system outputs the fused score, which is then used to replace the global Pearson correlation coefficient in the subsequent diagnostic framework. This method does not rely on additional external meteorological instruments; it extracts the internal structural features of the signal through mathematical transformations, thus achieving the separation of correlation components.
[0091] Example 4: This example details the method for measuring spatial expected deviation and the complete operation process for calculating spatial attenuation expectations using an isotropic benchmark model.
[0092] Step 401: Based on spatial distance, construct the spatial decay expectation corresponding to the candidate correlation.
[0093] Specifically, in hydraulic engineering structures, the correlation of responses between monitoring points typically exhibits a non-linear, natural decay characteristic as the physical distance increases. Based on pre-calculated Euclidean distances between three-dimensional spatial coordinates, the system establishes a theoretical benchmark characterizing the natural decay of correlation with physical distance. This benchmark is used to subsequently determine whether the actually acquired statistical correlations conform to the true physical propagation laws of the engineering space.
[0094] Step 402: Obtain the preset global spatial attenuation parameters.
[0095] In this embodiment, the system extracts a preset global spatial attenuation parameter from a pre-configured parameter set. This parameter is a fixed scalar reflecting the overall structural attenuation rate of the engineering project, and its value is determined based on the specific type of hydraulic engineering project. For example, for relatively uniform structural entities such as homogeneous earth-rock dams or long-distance water conveyance channels, a single global spatial attenuation parameter can be used to summarize the isotropic propagation attenuation characteristics of the monitoring response signal in the structural medium.
[0096] Step 403: Based on the spatial distance and global spatial attenuation parameters, calculate the expected spatial attenuation using the isotropic exponential attenuation function to characterize the theoretical propagation law of monitoring data in space.
[0097] The system assumes that the monitored response decays at the same rate in any three-dimensional direction within the engineering space. Based on the Euclidean distance between monitoring points obtained from preprocessing, the system employs a negative exponential model for mathematical mapping, converting the numerically continuous physical distance into a theoretical correlation intensity expectation located within a closed interval of zero to one. As the spatial distance gradually increases, the spatial decay expectation calculated using the exponential decay function shows a gradually decreasing trend.
[0098] Specifically, the expected spatial decay S ij The calculation formula is:
[0099] S ij =exp(-d ij / L);
[0100] Among them, S ij For spatial decay expectation, d ij Let L be the spatial distance, L be the preset global spatial decay parameter, and exp be the natural constant exponential function.
[0101] Step 404: Based on the correlation coefficients calculated in the correlation analysis, obtain the actual correlation strength of the candidate correlations.
[0102] In this embodiment, the system reads the absolute value of the correlation coefficient calculated for each candidate correlation node. In the basic implementation, the actual correlation strength is specifically extracted from the absolute value of the global Pearson correlation coefficient; however, in the case of using a time-frequency diagnostic alternative, the actual correlation strength can be characterized by the correlation ratio feature of a specific frequency band after frequency domain decomposition.
[0103] Step 405: Calculate the positive excess of the actual correlation strength relative to the expected spatial attenuation to obtain the spatial inconsistency feature, which is used to characterize the degree of deviation of the candidate correlation from the spatial propagation law.
[0104] In this embodiment, the system uses a difference operation instead of the traditional product operation to accurately calculate the physical deviation risk in the spatial dimension. Conventional algorithms, which use the product of actual correlation and expected spatial decay, cannot accurately characterize the abnormal logical conflict features of distant and highly correlated relationships. This solution subtracts the actual correlation strength from the theoretical expected spatial decay and extracts the positive value from the result. When the actual correlation strength exceeds the theoretical expected baseline corresponding to its spatial distance, the difference result takes a large positive value, indicating that the correlation relationship deviates significantly from the natural decay law of physical space and has a high risk of spurious correlation. Conversely, when the actual correlation strength is less than or equal to the theoretical expected distance, the difference result is forcibly set to zero, indicating that the correlation relationship is within the normal spatial physical conduction decay range and does not possess spatial anomaly characteristics.
[0105] Computational space inconsistency feature G ij The specific formula is:
[0106] G ij =max(0,|r ij |-S ij );
[0107] Among them, G ij For spatial inconsistency features, |r ij | represents the actual correlation strength between monitoring point i and monitoring point j, S ij Let represent the expected spatial decay, and max be the function to maximize the value.
[0108] Step 406: Based on the spatial topology of the monitoring points, extract the shortest topological hop count between two monitoring points corresponding to the candidate correlation.
[0109] Furthermore, in addition to relying on continuous Euclidean distance for judgment, the system introduces topological connectivity properties from discrete graph theory as a supplementary defense. Within a pre-constructed undirected graph model reflecting entity connectivity, the system uses a graph traversal algorithm to search for the shortest connected path between target node pairs, and outputs the number of network edges contained in this path as the shortest topological hop count. This hop count physically reflects the degree of obstruction in the indirect transmission of influence between nodes in an engineering structure.
[0110] Step 407: When the shortest topology hop count exceeds the preset maximum hop count threshold, a preset topology amplification factor is applied to the spatial inconsistency feature to obtain the corrected spatial inconsistency feature.
[0111] The system determines whether the calculated shortest topological hop count exceeds the preset graph connectivity boundary condition. If two monitoring points have neither a spatial physical adjacency nor a short-chain topological path reachable through a few intermediate nodes, the system infers strong physical isolation in the physical engineering structure. In this case, for the spatial inconsistency features that have already been exposed by the difference operation, the system further multiplies them by an additional penalty coefficient to amplify the penalty for the non-connectivity features in the topological dimension.
[0112] In some optional implementations, the preset maximum hop count threshold can be fixed at 2, and the preset topology amplification factor can be set to 0.5. For example, if a monitoring point pair exhibits data characteristics of being far apart but highly correlated, and the shortest topology hop count calculated by graph traversal reaches 3, the penalty mechanism of this embodiment is triggered, and its originally calculated deviation feature value will be amplified to 1.5 times the original value.
[0113] Accordingly, the corrected spatial inconsistency feature G ij * The calculation formula is:
[0114] G ij * =G ij *(1+β*1[h ij h max ]);
[0115] Among them, G ij * For the corrected spatial inconsistency feature, G ij The spatial inconsistency characteristic, β is the topological magnification factor, h ij h is the shortest topological hop count. max 1 represents the maximum number of hops threshold, and 1[·] represents the indicator function. When the condition within the parentheses is true, the indicator function outputs one; otherwise, it outputs zero.
[0116] Step 408: Use the corrected spatial inconsistency features as spatial inconsistency features for multi-dimensional fusion.
[0117] The system ultimately outputs a deviation index that has been double-cored and corrected based on both spatial physical distance and structural topology conditions. This corrected output index fully encompasses the isotropic natural attenuation characteristics of the space medium and penalizes the discretely distributed barrier fracture surfaces of the structure. Subsequently, this characteristic value is transmitted to the fusion scoring module of the data bus as the input basis for multi-dimensional analysis.
[0118] Example 5: This example is a preferred alternative to the above-described isotropic spatial attenuation model, such as... Figure 3As shown, the method for calculating the anisotropic attenuation expectation based on the local structural coordinate system and the offline data-driven estimation process of the characteristic attenuation length parameter are described in detail.
[0119] Step 501: Obtain the pre-configured engineering geometric parameters and construct a local structural coordinate system accordingly. The local structural coordinate system includes the principal axis direction, the normal direction, and the vertical direction.
[0120] Specifically, in actual hydraulic engineering structures, due to uneven distribution of structural stiffness and differences in physical boundary conditions, the propagation of monitoring responses often exhibits directional dependence. The system pre-acquires pre-configured engineering geometric parameters, including geometric feature information such as the orientation of the main structure's axis and the normal vectors of physical surfaces. Based on these parameters, the system constructs a local structural coordinate system that fits the engineering entity. This local structural coordinate system is established in the form of a three-dimensional Cartesian coordinate system, specifically defining the principal axis direction along the extension direction of the main structure, the normal direction perpendicular to the main water-retaining or water-passing surface, and the vertical direction along the gravity direction. By establishing this coordinate system, the system can map relative positions in global space to a physical coordinate space reflecting differences in structural stiffness.
[0121] Step 502: Based on the spatial coordinates of the two monitoring points corresponding to the candidate correlation, calculate the position difference vector and project the position difference vector onto the local structure coordinate system to obtain the distance components in the principal axis direction, normal direction and vertical direction.
[0122] After obtaining the spatial coordinates of the two monitoring points corresponding to the candidate correlation, the system performs a vector subtraction operation on the three-dimensional coordinates to generate the position difference vector between the two points in the global three-dimensional space. Subsequently, based on the direction cosine of the previously constructed local structural coordinate system, the system calculates the rotation matrix from the global coordinate system to the local structural coordinate system. The system multiplies the position difference vector with the rotation matrix and performs a coordinate rotation projection operation. For linear engineering structures, the rotation matrix is taken as a constant matrix; for curved structures, the system updates the rotation matrix using the local geometric tangent direction at the midpoint of the line connecting the two monitoring points. After projection mapping, the system extracts an index reflecting the degree of separation between the two monitoring points in the three principal inertial directions of the structure itself, namely the distance components in the principal axis direction, the normal direction, and the vertical direction.
[0123] Step 503: Obtain the preset anisotropic feature attenuation length parameter, and normalize the distance components based on the anisotropic feature attenuation length parameter to synthesize the anisotropic normalized distance.
[0124] In this embodiment, the system extracts preset anisotropic characteristic attenuation length parameters from an external parameter library or an offline calibration module. This parameter set contains three independent scalar values, each corresponding to a characteristic physical length reflecting the attenuation rate of the monitoring response signal propagating along three orthogonal directions. The system divides the distance components in each of the three directions by their respective characteristic physical lengths and performs dimensionless normalization. After completing the normalization process for each direction, the system calculates the square root of the sum of the squares of the three dimensionless quantities, synthesizes an equivalent spatial metric that comprehensively considers the differences in directional attenuation, and obtains the anisotropic normalized distance.
[0125] Specifically, the formula for calculating the anisotropic normalized distance is:
[0126] d ij aniso =sqrt((Δu ij / L u ) 2 +(Δv ij / L v ) 2 +(Δw ij / L w ) 2 );
[0127] Where, d ij aniso For the anisotropic normalized distance, Δu ij Δv ij and Δw ij Let L represent the projection values of the distance component in the principal axis direction, normal direction, and vertical direction, respectively. u L v and L w The preset anisotropic characteristic attenuation length parameters correspond to the attenuation lengths in the principal axis direction, normal direction, and vertical direction, respectively, and sqrt is the square root function.
[0128] Step 504: Calculate the expected spatial decay based on the anisotropic normalized distance.
[0129] The system uses the previously synthesized anisotropic normalized distance as an independent variable and inputs it into the negative exponential decay model. Since the directional parameter has already been absorbed, there is no need to introduce an additional global decay constant. The system calculates the natural exponential function of the negative anisotropic normalized distance and outputs the theoretically relevant intensity prediction value that reflects the directional physical conduction law.
[0130] Spatial decay expectation S ij The calculation formula is:
[0131] S ij =exp(-d ij aniso );
[0132] Among them, S ij For spatial decay expectation, d ij aniso Let be the anisotropic normalized distance, and exp be the natural constant exponential function.
[0133] Based on the aforementioned anisotropic attenuation model, to avoid systematic biases caused by subjective parameter settings, an offline data-driven estimation scheme for the preset anisotropic characteristic attenuation length parameter is provided. This method is executed before the system's real-time online pseudo-correlation monitoring phase.
[0134] Step 505: In the pre-stored historical monitoring data of water conservancy projects, select monitoring point pairs that are spatially distant from each other and belong to the same engineering structure segment, and construct a high-confidence reference sample set.
[0135] The system loads accumulated historical monitoring data from water conservancy projects. Based on prior engineering knowledge, data responses generated between monitoring points that are spatially close and without physical barriers caused by structural joints are highly likely to reflect real structural internal force relationships. Therefore, the system retrieves nodes whose spatial distance is shorter than a preset neighborhood distance threshold and filters out node pairs belonging to the same structural segment attribute, forming a high-confidence reference sample set for subsequent parameter inversion and fitting.
[0136] Step 506: Calculate the actual correlation coefficients of each monitoring point pair in the high-confidence reference sample set.
[0137] For each monitoring point pair included in the high-confidence reference sample set, the system calculates the global correlation strength within the historical operating cycle and extracts the absolute value as the actual correlation coefficient output under this spatial combination.
[0138] Step 507: Construct a least-squares objective function based on the actual correlation coefficient and the corresponding anisotropic normalized distance.
[0139] The system determines the optimal decay parameters by constructing an optimization mathematical model. The model considers the previously calculated actual correlation coefficient as the target fitted value and the spatial decay expectation calculated from the anisotropic normalized distance as the model prediction. The system calculates the sum of squares of the differences between the target fitted value and the model prediction, using this as the overall penalty cost to measure the bias of the decay model, thereby establishing the least squares objective function.
[0140] Step 508: By solving the least squares objective function, the optimal parameter that minimizes the deviation between the actual correlation coefficient and the expected spatial attenuation is fitted and used as the preset anisotropic characteristic attenuation length parameter.
[0141] The system invokes a multivariate nonlinear numerical optimization algorithm to minimize the objective function. Specifically, it can iteratively update the decay length variables in the three dimensions by executing the Levenberg-Marquardt algorithm until the iterative convergence condition is met, and finally output the set of parameter vectors that minimizes the penalty cost.
[0142] The optimization criterion for the least squares objective function is as follows:
[0143] (L u ,L v ,L w )=argmin Lu,Lv,Lw>0 ∑ (i,j)∈N_ref (|r ij |-exp(-d ij aniso )) 2 ;
[0144] In the formula, L u L v and L w These are the preset anisotropic feature attenuation length parameters that need to be fitted, N. ref Let (i,j) be a set of monitoring point pairs in the high-confidence reference sample set, and |r ij | represents the actual correlation strength between monitoring point i and monitoring point j, d ij aniso For the anisotropic normalized distance corresponding to the monitoring point pair, argmin represents the set of independent variable parameters that minimize the objective function.
[0145] Furthermore, as a fault-tolerant alternative, the system verifies the spatial distribution diversity of the high-confidence reference sample set before executing the solution. If it is found that all nodes in the reference sample set are distributed on the same elevation plane, it indicates that there is a lack of effective separation gradients in the vertical direction, causing the attenuation length in the vertical direction to fail to converge. At this time, the system triggers a parameter pre-set backoff mechanism, replacing the incorrectly estimated attenuation length parameter in a specific dimension with the default global spatial attenuation parameter, ensuring that the parameter estimation algorithm has robustness against the defects of a single spatial distribution sample.
[0146] Example 6, as Figure 4 As shown, this embodiment elaborates in detail the penalty constraint mechanism across engineering structure segments, as well as the spatial physical natural decay correction method to prevent statistical spurious transmission traps.
[0147] Step 601: Based on the two monitoring points corresponding to the candidate correlation, determine whether the two monitoring points belong to the same engineering structure segment in the engineering structure segment information.
[0148] The system reads the partition information configuration table of the engineering entity. It extracts the spatial attribution labels of the two monitoring points that constitute the candidate correlation in the physical structure. By comparing the consistency of the two labels, it determines whether the two monitoring points are located in the same continuous physical segment or belong to different structural segments isolated by physical gaps, expansion joints, or independent casting units.
[0149] Step 602: When the two monitoring points do not belong to the same engineering structure segment, obtain the preset cross-segment penalty coefficient.
[0150] For monitoring point pairs belonging to different physical structural segments, the system extracts a scalar parameter from the pre-configured parameter space, with a value greater than zero and less than one, as a cross-segment penalty coefficient. This coefficient reflects the degree to which the physical boundary weakens the transmission of the structural response.
[0151] In one optional implementation, the span penalty coefficient α can be set according to the joint type and physical isolation degree of the engineering structure: for adjacent dam sections completely isolated by permanent expansion joints, α can be set to 0.3 to 0.5; for sections using temporary construction joints, α can be set to 0.5 to 0.8. Those skilled in the art can determine an appropriate value for α by comparing the correlation distribution differences between span-segment measuring point pairs and same-segment measuring point pairs based on the historical statistical characteristics of the structural design documents and on-site monitoring data, using quantile statistical methods.
[0152] Step 603: Use the cross-segment penalty coefficient to correct the actual correlation strength calculated in the correlation analysis of the candidate correlations, and obtain the structural consistency score.
[0153] The system multiplies this coefficient by the actual correlation strength calculated between the two nodes. After the product operation, the original statistical correlation strength is proportionally reduced, and the reduced value is defined as the structural consistency score. For monitoring point pairs belonging to the same structural segment, this score is equal to the actual correlation strength.
[0154] Step 604: Calculate the structural deviation based on the structural consistency score, and use the structural deviation as a feature of structural inconsistency.
[0155] The system uses a method of subtracting the structural consistency score from the unit score to invert the positive score representing the degree of consistency into a risk metric representing the degree of inconsistency. The output is mapped to the structural inconsistency features required in the multidimensional fusion stage.
[0156] Structural inconsistency characteristics (1-R) ij The calculation process satisfies:
[0157] R ij =|r ij |*δ ij ;
[0158] When two monitoring points belong to the same engineering structure segment, δ ij =1; when they do not belong to the same engineering structure segment, δ ij =α, where α is the cross-segment penalty coefficient; in the formula, R ij For structural consistency score, |r ij | represents the actual correlation strength between monitoring point i and monitoring point j.
[0159] Step 605: In the spatial topology of the monitoring points, traverse and search for neighborhood triples that satisfy the direct spatial neighborhood connectivity condition. A neighborhood triple consists of a first node, an intermediate node, and a tail node.
[0160] The system performs a graph traversal algorithm within a pre-generated undirected graph topology. The system sets the path search condition to a two-hop path connected through a single intermediate node, selecting node groups consisting of three neighboring nodes. This connectivity constraint reduces the computational scale and focuses computation on adjacent regions with potential physical connections within the engineering structure.
[0161] Step 606: Obtain the actual correlation strength between the first node and the intermediate node, and between the intermediate node and the tail node to form the association propagation expectation; perform attenuation correction on the association propagation expectation based on the spatial distance between the first node and the tail node to obtain the corrected association propagation expectation.
[0162] For the node groups obtained from the search, the system extracts the actual correlation strength of the first and second path segments respectively, and calculates the expected correlation propagation of the indirect path by multiplying the two. At the mathematical and statistical level, the Pearson correlation coefficient does not possess inherent propagation properties. When two independent sequences are both driven by a third independent factor, even if the preceding node is highly correlated with the intermediate node, and the intermediate node is highly correlated with the subsequent node, mathematically, it is permissible for the first and last nodes to exhibit low correlation or even no correlation. Therefore, relying on the mathematical propagation properties of correlation strength to judge spurious correlation chains can lead to logical misjudgments.
[0163] Step 607: Compare the corrected expected correlation propagation with the actual correlation strength between the first and last nodes, extract the propagation deviation, and construct correlation propagation inconsistency features to identify pseudo-correlation chain structures that do not meet propagation consistency.
[0164] The system subtracts and normalizes the theoretically calculated expected transmission of the association from the actual extracted association strengths of the first and last nodes. It then calculates the deviation ratio of the difference from the expected transmission, generating an intermediate variable characterizing the risk of the transmission chain breaking.
[0165] Obtain the preset transmission attenuation constant and construct a spatial natural attenuation correction factor based on the spatial distance between the first node and the last node.
[0166] Specifically, the system obtains the preset transmission attenuation constant L natural And based on the spatial distance d between the first node i and the last node j ij The spatial natural decay correction factor S is constructed using an exponential decay function. ij natural The spatial natural attenuation correction factor is used to characterize the natural attenuation of the indirect correlation transmission strength between the first and last nodes as the spatial distance increases under the condition of physical spatial separation.
[0167] In one alternative implementation, the transmission attenuation constant L natural It can be set to the same value as the global spatial attenuation parameter L to maintain physical consistency of spatial attenuation; or it can be calibrated separately using the least squares regression method by statistically analyzing the transmission behavior of neighborhood triples in historical monitoring data, based on the spatial span characteristics of indirect transmission paths in actual engineering projects. Those skilled in the art can make adaptive selections according to the actual engineering situation.
[0168] Step 608: The correlation transit expectation is corrected using the spatial natural decay correction factor to obtain the corrected correlation transit expectation.
[0169] For any neighborhood triple consisting of the first node i, the middle node k, and the tail node j, the system constructs an uncorrected correlation transit expectation E based on the actual correlation strength between the first node and the middle node, and the actual correlation strength between the middle node and the tail node. ikj Then, the spatial natural decay correction factor S is used. ij natural The transitive expectation of association is modified to obtain the modified transitive expectation of association. This ensures that the correlation propagation expectation satisfies the spatial natural decay constraint between the first and last nodes.
[0170] Step 609: Compare the corrected expected correlation propagation with the actual correlation strength between the first and last nodes, and extract the propagation inconsistencies of a single indirect correlation path.
[0171] The system will pass the corrected correlation expectation. The actual correlation strength |r between the first node i and the last node j ij | Compare and calculate the transitive inconsistency T corresponding to a single indirect association path. ikj When the actual correlation strength between the first and last nodes is significantly lower than the expected correlation propagation strength after correction, it indicates that the path has a high risk of propagation inconsistency.
[0172] Step 610: Summarize the transitive inconsistencies corresponding to each neighborhood triplet to obtain the transitive inconsistency features of the candidate correlation.
[0173] For the first and last node pairs ij corresponding to the same candidate correlation, the system traverses all intermediate nodes k that satisfy the direct spatial neighborhood connectivity condition, and calculates the inconsistency T of the single indirect association path transmission for each neighborhood triple. ikj The maximum value is extracted from this and used as the inconsistency feature T of the association between the first and last nodes. ij This characterizes the highest transmission break risk of the candidate correlation across all possible indirect propagation paths.
[0174] Extracted association transitivity inconsistency feature T ij The calculation formula is:
[0175] Correlation transitive expectation:
[0176]
[0177] Among them, E ikj Let r be the uncorrected transitive expectation of the neighborhood triple consisting of the first node i, the middle node k, and the tail node j; ik r is the correlation coefficient between the first node and the intermediate nodes. kj This represents the correlation coefficient between the intermediate node and the tail node.
[0178] Spatial natural decay correction factor:
[0179]
[0180] in, is the spatial natural decay correction factor between the first node i and the last node j; This represents the spatial distance between the first and last nodes. This is the preset transmission attenuation constant.
[0181] Revised expectation of association transitivity:
[0182]
[0183] in, This represents the expected correlation propagation after spatial natural attenuation correction.
[0184] Inconsistency in transmission along a single indirect association path:
[0185]
[0186] Among them, T ikj For neighborhood triples The corresponding single indirect association path transmission is inconsistent; r ij The correlation coefficient between the first and last nodes; ε represents the actual correlation strength between the first and last nodes; ε is a very small positive number to prevent the denominator from being zero.
[0187] Inconsistency characteristics of association transitivity:
[0188]
[0189] Among them, T ij For monitoring points corresponding to candidate correlations The inconsistency characteristic of association transmission represents the maximum value of the inconsistency in the transmission of a single indirect association path among all intermediate nodes k that satisfy the direct spatial neighborhood connectivity condition.
[0190] To verify the effect of introducing spatial physical attenuation correction, a normalized dimensionless parameter example was constructed. The normalized spatial distance between the first node A and the intermediate node B was set to 0.5, the normalized spatial distance between the intermediate node B and the tail node C was set to 0.6, and the normalized long span distance between the first node A and the tail node C was measured to be 2.0. The transmission attenuation constant was set to 1.0.
[0191] The system extracts the actual correlation strength between node A and node B as 0.85, and the actual correlation strength between node B and node C as 0.82. The measured actual correlation strength between node A and node C is only 0.15. The system first calculates the uncorrected correlation transit expectation, obtaining... Furthermore, based on the normalized spatial distance of 2.0 between node A and node C and the propagation attenuation constant of 1.0, the spatial natural attenuation correction factor is calculated. .
[0192] The system corrects the association transit expectation to obtain the corrected association transit expectation. .
[0193] Since the actual correlation strength between node A and node C (0.15) is higher than the corrected expected correlation transitivity (0.094), substituting this into the transitivity inconsistency calculation formula yields... .
[0194] The results show that after introducing spatial natural decay correction, the high transmission expectation between long-span node pairs, which was originally derived solely from statistical paths, is reasonably reduced, thereby avoiding misjudging node relationships that conform to the spatial decay law as pseudo-transmission chains.
[0195] Example 7: In another aspect of this example, a method for identifying pseudo-correlation in water conservancy project monitoring data includes:
[0196] Acquire monitoring data and spatial attribute information of monitoring points in water conservancy projects. Spatial attribute information includes the spatial coordinates of the monitoring points and the segmentation information of the engineering structure.
[0197] The monitoring data is preprocessed to obtain a standardized time series, and a spatial topology of the monitoring points is constructed based on the spatial coordinates of the monitoring points.
[0198] Correlation analysis is performed on standardized time series to identify candidate correlations and construct an initial monitoring correlation network. Time-statistical correlation features of candidate correlations are extracted, including correlation coefficient features based on global statistical correlation and / or frequency band correlation contribution spectrum and its distribution features obtained based on multi-scale decomposition.
[0199] Based on the spatial distance calculated from the spatial topology and spatial coordinates of the monitoring points, a spatial attenuation expectation corresponding to the candidate correlation is constructed, and spatial inconsistency characteristics are calculated based on the degree of deviation between the actual correlation strength and the spatial attenuation expectation.
[0200] Based on the segmentation information of the engineering structure, structural inconsistency features that cross structural segments are extracted from candidate correlations;
[0201] Based on the spatial topology of the monitoring points, indirect association paths between candidate correlation monitoring point pairs are searched, association transmission expectation is constructed, and the association transmission expectation is attenuated and corrected based on the spatial distance between the monitoring point pairs. Then, the association transmission inconsistency characteristics are calculated according to the degree of deviation between the corrected association transmission expectation and the actual correlation strength.
[0202] By fusing time-statistical correlation features, spatial inconsistency features, structural inconsistency features, and association transmission inconsistency features from multiple dimensions, a pseudo-correlation risk index is calculated, and pseudo-correlation relationships in candidate correlation relationships are identified based on this index, thereby updating the monitoring correlation network.
[0203] A pseudo-correlation identification system for water conservancy project monitoring data, comprising:
[0204] The data acquisition module is used to acquire monitoring data collected by the monitoring system.
[0205] The spatial attribute construction module is used to construct the spatial topology of monitoring points;
[0206] The correlation analysis module is used to identify candidate correlations;
[0207] The consistency verification module is used to perform spatial consistency verification and structural consistency verification.
[0208] The association propagation analysis module is used to identify spurious correlation chains;
[0209] The spurious correlation identification module is used to calculate the spurious correlation risk index and identify spurious correlation relationships.
[0210] Example 8: This example uses a safety monitoring system for a water conservancy project. The system is deployed on water conveyance channels, tunnels, or dam structures. Monitoring includes one or more of the following: deformation monitoring, seepage pressure monitoring, temperature monitoring, strain monitoring, and environmental quantity monitoring. The system consists of multiple monitoring points, each with a unique monitoring number. Its spatial coordinates, station number, structural segment information, and monitoring data time series are recorded. The monitoring data time series and spatial attribute information serve as input data for pseudo-correlation identification analysis in this example.
[0211] First, monitoring data collected by the water conservancy project safety monitoring system is acquired, and the raw monitoring data corresponding to each monitoring point is preprocessed. Preprocessing includes missing value completion, outlier correction, time base alignment, and standardization to eliminate the impact of differences in the dimensions of different monitoring items and abnormal sampling disturbances on the subsequent correlation analysis results, obtaining the standardized time series corresponding to each monitoring point. Assume there are a total of [number missing] monitoring points in the project. The monitoring point, the first The time series of monitoring data from each monitoring point is denoted as follows: ,in, Indicates monitoring point At any moment The monitoring value, This indicates the length of the monitoring data sequence. After standardization, a standardized time series is obtained. .
[0212] Subsequently, spatial attribute information for each monitoring point is extracted, including spatial coordinates, station number, and the segment of the engineering structure to which the monitoring point belongs. Based on the spatial coordinates of each monitoring point, a spatial topology is constructed, and the spatial adjacency relationships between monitoring points are determined in conjunction with the engineering layout characteristics, providing a foundation for subsequent extraction of spatial inconsistencies and related feature transfer. For any two monitoring points... and It can calculate the corresponding spatial distance based on its spatial coordinates and determine whether it belongs to a pair of spatially adjacent monitoring points according to preset neighborhood rules.
[0213] After obtaining the standardized time series and spatial topology of the monitoring points, the statistical correlation between the monitoring points is calculated to identify candidate correlations and construct an initial monitoring correlation network. For each candidate correlation edge in the initial monitoring correlation network, the corresponding time-statistical correlation feature is extracted. In one implementation, the absolute value of the global Pearson correlation coefficient can be used as the time-statistical correlation feature; in another preferred implementation, the standardized time series corresponding to the candidate correlations are decomposed into multiple scales to construct a frequency band correlation contribution spectrum, extract the environmental frequency band correlation concentration and correlation contribution frequency band dispersion, and use the fused frequency domain pseudo-correlation diagnostic score as supplementary discriminant information on the time side to enhance the ability to identify pseudo-correlations caused by common environmental excitations.
[0214] For each candidate association edge in the initial monitoring association network, spatial inconsistency features are extracted. In the preferred implementation, the anisotropic normalized distance is calculated in the local structural coordinate system, and the anisotropic spatial inconsistency features are extracted accordingly to improve the identification accuracy of pseudo-correlation relationships in hydraulic structures with obvious structural orientation.
[0215] Simultaneously, for each candidate correlation edge, structural inconsistency features are extracted based on the structural segment information of the monitoring point. These features characterize whether the candidate correlation crosses discontinuous structural segments, deviates from normal structural connection paths, and exhibits abnormally strong cross-segment correlation characteristics. When the two monitoring points corresponding to a candidate correlation belong to different structural segments, and this cross-segment relationship does not conform to the continuity of the engineering structure or the actual load transfer pattern, the structural inconsistency risk corresponding to this candidate correlation increases.
[0216] Furthermore, for each candidate association edge, an indirect association path between the first and last nodes is searched in the spatial topology of the monitoring points. An association propagation expectation is constructed based on this indirect path, and this expectation is attenuated and corrected according to the spatial distance between the first and last nodes. Then, association propagation inconsistency features are extracted. These features characterize whether candidate associations are formed by indirect propagation from intermediate nodes, and whether the actual correlation strength between the first and last nodes deviates significantly from the corrected association propagation expectation.
[0217] After extracting time-statistical correlation features, spatial inconsistency features, structural inconsistency features, and association transmission inconsistency features, the above multidimensional features are fused to calculate the pseudo-correlation risk index corresponding to each candidate association edge. The fusion process can be implemented using a linear weighting method, or using a logistic regression model, tree model, or neural network model. In a preferred embodiment, multidimensional feature fusion is implemented using a linear weighting method; in another embodiment, multidimensional feature fusion is implemented using a nonlinear model. For any candidate association edge, if its corresponding pseudo-correlation risk index is greater than a preset judgment threshold, the candidate association edge is judged as a pseudo-correlation; if its corresponding pseudo-correlation risk index is not greater than the preset judgment threshold, the candidate association edge is retained as a reliable correlation.
[0218] Finally, the initial monitoring correlation network is updated based on the results of pseudo-correlation identification. Candidate correlation edges identified as pseudo-correlation edges are deleted, downweighted, or marked; candidate correlation edges identified as reliable correlation edges are retained, thus forming a revised monitoring correlation network. The revised monitoring correlation network can be further used in applications such as anomaly analysis of monitoring data, identification of engineering operation status, safety risk assessment, construction of monitoring and early warning models, and optimization of monitoring system deployment, thereby improving the reliability and engineering applicability of water conservancy project safety monitoring data analysis results.
[0219] The method described in this embodiment can effectively identify spurious correlations caused by shared environmental excitation, long-distance weak coupling, cross-segmental influence of discontinuous structures, or indirect propagation chains, thereby improving the accuracy of identifying the true correlations of monitoring points. The corrected monitoring correlation network exhibits higher consistency with the actual connection relationships, spatial propagation laws, and physical response mechanisms of the engineering structure, providing a more reliable data correlation foundation for subsequent safety diagnosis and risk warning.
[0220] The preferred embodiments of the present invention have been described in detail above. However, the present invention is not limited to the specific details in the above embodiments. Within the scope of the technical concept of the present invention, various equivalent transformations can be made to the technical solutions of the present invention, and these equivalent transformations all fall within the protection scope of the present invention.
Claims
1. A method for identifying pseudo-correlation in water conservancy project monitoring data, characterized in that, include: Acquire monitoring data and spatial attribute information of monitoring points in water conservancy projects. Spatial attribute information includes the spatial coordinates of the monitoring points and the segmentation information of the engineering structure. The monitoring data is preprocessed to obtain a standardized time series, and a spatial topology of the monitoring points is constructed based on the spatial coordinates of the monitoring points. Correlation analysis was performed on standardized time series to identify candidate correlations and construct an initial monitoring association network, and time-statistical correlation features of candidate correlations were extracted. Based on the spatial distance calculated from the spatial topology and spatial coordinates of the monitoring points, the spatial inconsistency characteristics of candidate correlations deviating from the spatial propagation law are calculated. Based on the segmentation information of the engineering structure, structural inconsistency features that cross structural segments are extracted from candidate correlations; Based on the spatial topology of the monitoring points, search for indirect association paths between candidate related monitoring point pairs, and calculate the inconsistency features of association transmission on the indirect association paths; By fusing time-statistical correlation features, spatial inconsistency features, structural inconsistency features, and association transmission inconsistency features from multiple dimensions, a pseudo-correlation risk index is calculated, and pseudo-correlation relationships in candidate correlation relationships are identified based on this index, thereby updating the monitoring correlation network.
2. The method according to claim 1, characterized in that, Extract time-statistical correlation features of candidate correlations, including: Calculate the global Pearson correlation coefficient between the standardized time series corresponding to the candidate correlations; The absolute value of the global Pearson correlation coefficient is used as a time-statistical correlation feature to characterize the overall statistical correlation of candidate correlations.
3. The method according to claim 1, characterized in that, Extract time-statistical correlation features of candidate correlations, including: Orthogonal wavelets are used to perform multi-scale decomposition on the standardized time series corresponding to the candidate correlation to obtain the approximate components of each monitoring point and the detail components of multiple frequency bands. Calculate the correlation contribution of candidate correlations at the approximate component and each detail component level respectively, and construct the band correlation contribution spectrum; Based on the distribution characteristics of the correlation contribution of each frequency band in the frequency band correlation contribution spectrum, a quantitative index characterizing the degree of concentration of frequency domain correlation is extracted. The quantitative index is used as a time-statistical correlation feature to identify pseudo-correlation patterns driven by environmental factors in candidate correlation relationships.
4. The method according to claim 3, characterized in that, Based on the distribution characteristics of the correlation contributions of each frequency band in the frequency band correlation contribution spectrum, a quantitative index characterizing the concentration of frequency domain correlation is extracted. This quantitative index is used as a time-statistical correlation feature, including: Obtain the pre-configured set of environmental feature frequency bands; The environmental frequency band correlation concentration is obtained by calculating the proportion of the sum of the absolute values of the correlation contributions of the frequency band correlation contribution spectrum within the set of environmental characteristic frequency bands to the absolute value of the total correlation contributions. The correlation contribution frequency band dispersion is calculated based on the proportion of the absolute value of the correlation contribution of each frequency band in the frequency band correlation contribution spectrum. Obtain the preset main index weight coefficients, and perform weighted fusion of the environmental frequency band correlation concentration and the normalized correlation contribution frequency band dispersion based on the main index weight coefficients to calculate the frequency domain pseudo-correlation diagnostic score. The frequency domain pseudo-correlation diagnostic score was used as a time-statistical correlation feature.
5. The method according to claim 1, characterized in that, Calculate the spatial inconsistency features of candidate correlations that deviate from the spatial propagation law, including: Based on spatial distance, construct spatial decay predictions for candidate correlations; Based on the correlation coefficients calculated in the correlation analysis, the actual correlation strength of the candidate correlations is obtained. The positive excess of the actual correlation strength relative to the expected spatial attenuation is calculated to obtain the spatial inconsistency feature, which is used to characterize the degree of deviation of candidate correlations from the spatial propagation law.
6. The method according to claim 5, characterized in that, After obtaining the spatial inconsistency features, the following is also included: Based on the spatial topology of the monitoring points, the shortest topological hop count between two monitoring points corresponding to candidate correlations is extracted. When the shortest topology hop count exceeds the preset maximum hop count threshold, a preset topology amplification factor is applied to the spatial inconsistency feature to obtain the corrected spatial inconsistency feature. The modified spatial inconsistency features are used as spatial inconsistency features for multi-dimensional fusion.
7. The method according to claim 5, characterized in that, Based on spatial distance, construct the spatial decay expectation corresponding to the candidate correlation, including: Obtain the preset global spatial decay parameters; Based on spatial distance and global spatial attenuation parameters, spatial attenuation expectation is calculated using an isotropic exponential attenuation function to characterize the theoretical propagation law of monitoring data in space.
8. The method according to claim 5, characterized in that, Based on spatial distance, construct the spatial decay expectation corresponding to the candidate correlation, including: Obtain pre-configured engineering geometric parameters and construct a local structural coordinate system accordingly. The local structural coordinate system includes principal axis direction, normal direction and vertical direction. Based on the spatial coordinates of the two monitoring points corresponding to the candidate correlation, the position difference vector is calculated, and the position difference vector is projected onto the local structure coordinate system to obtain the distance components in the principal axis direction, normal direction and vertical direction; Obtain the preset anisotropic feature attenuation length parameter, and normalize the distance components based on the anisotropic feature attenuation length parameter to synthesize the anisotropic normalized distance. Spatial decay prediction is calculated based on anisotropic normalized distance.
9. The method according to claim 1, characterized in that, Based on the segmented information of the engineering structure, structural inconsistency features that cross structural segments are extracted from candidate correlations, including: Based on the two monitoring points corresponding to the candidate correlation, determine whether the two monitoring points belong to the same engineering structure segment in the engineering structure segment information; When two monitoring points do not belong to the same engineering structure segment, obtain the preset cross-segment penalty coefficient; The cross-segment penalty coefficient is used to correct the actual correlation strength calculated in the correlation analysis of candidate correlations, and the structural consistency score is obtained. Structural deviation is calculated based on structural consistency score, and structural deviation is used as a feature of structural inconsistency.
10. The method according to claim 1, characterized in that, Based on the spatial topology of monitoring points, indirect association paths between candidate related monitoring point pairs are searched, and inconsistencies in association transmission along these indirect association paths are calculated, including: In the spatial topology of the monitoring points, the neighbor triples that satisfy the direct spatial neighborhood connectivity condition are traversed and searched. The neighbor triples consist of a first node, an intermediate node, and a tail node. The actual correlation strength between the first node and the middle node, and between the middle node and the tail node, are obtained respectively to form the expected correlation propagation. The association propagation expectation is attenuated and corrected based on the spatial distance between the first node and the last node to obtain the corrected association propagation expectation. By comparing the corrected expected correlation propagation with the actual correlation strength between the first and last nodes, the propagation deviation is extracted, and correlation propagation inconsistency features are constructed to identify pseudo-correlation chain structures that do not meet the propagation consistency.
11. A system for identifying pseudo-correlation in water conservancy project monitoring data, characterized in that, include: The data acquisition module is used to acquire monitoring data collected by the monitoring system. The spatial attribute construction module is used to construct the spatial topology of monitoring points; The correlation analysis module is used to identify candidate correlations; The consistency verification module is used to perform spatial consistency verification and structural consistency verification. The association propagation analysis module is used to identify spurious correlation chains; The spurious correlation identification module is used to calculate the spurious correlation risk index and identify spurious correlation relationships.