A water level monitoring equipment fault prediction method based on multi-source data fusion

CN122113006BActive Publication Date: 2026-08-11ZHEJIANG RUILIN INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-04-27
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0002]水位监测设备作为水文监测体系的核心终端,其运行稳定性直接影响水文数据采集的准确性与连续性,现阶段设备故障预测多依赖单一传感数据的阈值判断或简单趋势分析,未能捕捉设备内部各传感单元间的时变交互关系,仅通过独立数据异常判定故障,忽略了传感单元间的耦合关联特性,难以识别故障潜伏期的微弱状态变化,导致故障预测的提前性和准确性不足,同时,现有技术多采用静态特征建模方式,未构建能够动态反映传感单元线性关联与因果关系的交互网络,无法从系统层面刻画设备的整体运行状态,易出现故障漏判或误判问题

Benefits of technology

1.本发明依托融合时变相关性与因果关系的动态耦合关系网络构建方法,打造了精准的故障特征提取机制,相比传统方法实现了多维度的技术突破,该机制先对多源时序数据做平稳化处理,再通过动态条件协方差模型挖掘传感变量间的时变线性关联强度,结合滑动时间窗口下的多元格兰杰因果检验量化变量间的领先-滞后因果关系,融合生成有向加权的动态耦合关系网络,从强度、方向、动态性三个维度精准表征设备内部传感单元间的复杂时变交互状态,这一方式突破了传统单一数据分析、静态特征建模的局限,摒弃了仅关注独立传感数据异常的弊端,能够捕捉到故障潜伏期传感单元交互关系的微弱异常变化,实现了从单一数据特征到系统耦合特征的提取升级,让故障特征的挖掘更贴合设备实际运行的复杂状态,为后续故障预测奠定了精准的特征基础。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122113006B_ABST
    Figure CN122113006B_ABST
Patent Text Reader

Abstract

This invention relates to the field of fault prediction technology and discloses a fault prediction method for water level monitoring equipment based on multi-source data fusion. The method includes: stabilizing multi-source time-series data to obtain a weakly stationary multivariate sequence; constructing a dynamic coupling relationship network through a dynamic conditional covariance model and Granger causality test; extracting topological feature vectors from the network; establishing a deviation reference model based on the topological feature vectors from historical normal periods; calculating the deviation of the topological feature vectors to obtain an equipment coupling health index; and generating a fault prediction and early warning based on the evolution trend of the index and its matching degree with pre-stored fault modes. This invention can improve the accuracy of fault prediction for water level monitoring equipment based on multi-source data fusion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of fault prediction technology, and in particular to a fault prediction method for water level monitoring equipment based on multi-source data fusion. Background Technology

[0002] As the core terminal of the hydrological monitoring system, the operational stability of water level monitoring equipment directly affects the accuracy and continuity of hydrological data acquisition. At present, equipment fault prediction mostly relies on threshold judgment or simple trend analysis of single sensor data, failing to capture the time-varying interaction relationship between various sensor units within the equipment. Faults are determined solely by independent data anomalies, ignoring the coupling and correlation characteristics between sensor units. It is difficult to identify subtle state changes during the fault latency period, resulting in insufficient advance warning and accuracy of fault prediction. At the same time, existing technologies mostly adopt static feature modeling methods and fail to construct an interactive network that can dynamically reflect the linear correlation and causal relationship of sensor units. This makes it impossible to characterize the overall operating status of the equipment at the system level, which easily leads to missed or misjudged faults.

[0003] Existing methods for predicting faults in water level monitoring equipment lack quantitative assessment and trend analysis mechanisms for equipment health status. They judge equipment status solely based on deviations from characteristics at a single moment, without establishing a standardized deviation reference model based on historical normal conditions or conducting pattern matching analysis on the evolution trajectory of health indices. This makes it difficult to distinguish between occasional data fluctuations and systemic deterioration caused by faults, and fails to achieve closed-loop early warning from equipment status perception to fault development trend prediction. Furthermore, traditional methods lack a fault pattern library and a matching mechanism for dynamic time warping, resulting in low accuracy in fault type identification. This makes it difficult to provide targeted fault warnings and maintenance suggestions for operation and maintenance personnel, and fails to meet the actual needs of hydrological monitoring for early prediction and accurate assessment of equipment faults. Therefore, improving the accuracy of hydrological monitoring in predicting equipment faults has become an urgent problem to be solved. Summary of the Invention

[0004] This invention provides a method for predicting the faults of water level monitoring equipment based on multi-source data fusion, in order to solve the problems mentioned in the background art.

[0005] To achieve the above objectives, the present invention provides a method for predicting faults in water level monitoring equipment based on multi-source data fusion, comprising: S1, acquire multi-source time-series data from water level monitoring equipment, and perform stabilization processing on the multi-source time-series data to obtain a weakly stationary multivariate sequence; S2, Based on the weakly stationary multivariate sequence, a dynamic coupling relationship network characterizing the time-varying interaction relationship between sensing units inside the device is constructed through a dynamic conditional covariance model and Granger causality test. S3, extract the topological feature vector representing the stability of the network structure from the dynamic coupling relationship network; S4. Establish a deviation reference model based on the topological feature vectors of historical normal periods, and analyze and calculate the deviation of the topological feature vectors based on the deviation reference model to obtain the equipment coupling health index. S5. Based on the evolution trend of the device coupled health index and its matching degree with the pre-stored fault modes, generate a fault prediction and early warning.

[0006] In a preferred embodiment, the step of acquiring multi-source time-series data from a water level monitoring device and performing stationarization processing on the multi-source time-series data to obtain a weakly stationary multivariate sequence includes: Acquire raw multi-source time-series data generated by the target water level monitoring equipment within a preset sampling period; The original multi-source time series data is subjected to time alignment and outlier removal operations to obtain preprocessed multi-source time series data; The preprocessed multi-source time series data were subjected to white noise self-testing method to obtain the data stationarity test results. Based on the data stationarity test results, the sequences that are determined to be non-white noise are subjected to first-order differencing, and the differencing sequences are iteratively tested for white noise until they pass the white noise test, so as to obtain the weakly stationary multivariate sequence.

[0007] In a preferred embodiment, the step of constructing a dynamic coupling relationship network characterizing the time-varying interaction relationships between sensing units within the device based on the weakly stationary multivariate sequence, through a dynamic conditional covariance model and Granger causality test, includes: Based on the weakly stationary multivariate sequence, the time-varying covariance matrix among the sensing variables is estimated by a dynamic conditional covariance model. Extract the time-varying correlation coefficient matrix representing the change in the strength of the linear association between variables over time from the time-varying covariance matrix; Based on the weakly stationary multivariate sequence, within a preset sliding time window, the time-varying causal strength matrix representing the lead-lag causal relationship between variables is calculated using the multivariate Granger causality test algorithm. By fusing the time-varying correlation coefficient matrix and the time-varying causal strength matrix, a directed weighted dynamic coupling relationship network is constructed.

[0008] In a preferred embodiment, the mathematical expression of the multivariate Granger causality test algorithm is as follows: The multivariate Granger causality test algorithm consists of both restricted and unrestricted models. The mathematical expression for the constrained model is as follows: ; The mathematical expression for the unrestricted model is as follows: ; In the formula X i (t) and X j (tp) represents the observed values ​​of the i-th and j-th sensed variables in the weakly stationary multivariate sequences at times t and tp, respectively, where P is the lag order and A is the lag order. p and B p These are the model coefficients, e1(t) and e2(t) are the residual terms, and p is the order index.

[0009] In a preferred embodiment, extracting the topological feature vector representing the stability of the network structure from the dynamically coupled network includes: Based on the adjacency matrix of the dynamically coupled network, a global clustering coefficient is calculated to measure the tightness of the connections between nodes within the network. Based on the shortest path information between nodes in the dynamically coupled network, the global efficiency, which measures the efficiency of network information transmission, is calculated. Based on the number of connections of each node in the dynamic coupling network, the standard deviation of node degree centrality, which measures the uniformity of network connection distribution, is calculated. The global clustering coefficient, the global efficiency, and the standard deviation of node degree centrality are combined to generate a topological feature vector.

[0010] In a preferred embodiment, calculating the standard deviation of node degree centrality, which measures the uniformity of network connectivity distribution, includes: First, calculate the degree centrality of each node in the network, i.e., the number of connections k of each node. i Then calculate the standard deviation of the degree centrality of all nodes, which is expressed mathematically as follows: ; In the formula, D t Let N be the standard deviation of the degree centrality of nodes at time t, and N be the total number of nodes, k i Let i be the degree of node i. Let be the average degree of all nodes, and i be the node index.

[0011] In a preferred embodiment, establishing the deviation reference model based on topological feature vectors from historical normal periods includes: Collect all topological feature vectors generated by the equipment during its historical normal and fault-free operation phase to form a set of historical normal topological feature vectors; The mean vector reflecting the central position of the set of historical normal topological feature vectors and the covariance matrix representing the dispersion and correlation of the data are calculated by using parameter estimation methods. The deviation reference model is defined by the mean vector and the covariance matrix.

[0012] In a preferred embodiment, the step of analyzing and calculating the deviation of the topological feature vector based on the deviation reference model to obtain the device coupling health index includes: Based on the deviation reference model and the topological feature vector at the current time, calculate the outlier value, which represents the negative logarithm of the probability density of the topological feature vector at the current time under the deviation reference model. Based on a preset scaling function, the anomaly value is mapped and transformed to obtain the device coupling health index.

[0013] In a preferred embodiment, generating a fault prediction and early warning based on the evolution trend of the device coupled health index and its matching degree with pre-stored fault modes includes: Continuously monitor the device's coupled health index to obtain the health index sequence within the most recent time window; Based on the health index sequence, its short-term decline trend is calculated using a trend estimation algorithm; The health index sequence is dynamically time-warped similarity matched with typical fault mode sequences in the pre-stored fault mode library to obtain the matching distance. Based on the short-term decay trend and the matching distance, a fault prediction logic is executed to generate the fault prediction warning.

[0014] In a preferred embodiment, the pre-stored fault mode library is a database that stores the typical evolution trajectory of the equipment coupling health index before and after various typical faults in history. Each typical fault mode is represented by a health index sequence, which captures the characteristic change shape of the health index from the fault latency period to the fault occurrence period.

[0015] Compared with the prior art, the present invention has the following beneficial effects: 1. This invention relies on a dynamic coupling relationship network construction method that integrates time-varying correlation and causal relationships to create a precise fault feature extraction mechanism. Compared with traditional methods, it achieves multi-dimensional technical breakthroughs. This mechanism first performs stabilization processing on multi-source time-series data, then mines the strength of time-varying linear correlations between sensor variables through a dynamic conditional covariance model, and combines multivariate Granger causality tests under a sliding time window to quantify the lead-lag causal relationship between variables. It then fuses and generates a directed weighted dynamic coupling relationship network, which accurately characterizes the complex time-varying interaction state between sensor units inside the equipment from three dimensions: strength, direction, and dynamism. This approach breaks through the limitations of traditional single data analysis and static feature modeling, and abandons the drawback of only focusing on anomalies in independent sensor data. It can capture subtle abnormal changes in the interaction relationship of sensor units during the fault latency period, realizing an upgrade from single data feature extraction to system coupling feature extraction. This makes the mining of fault features more consistent with the complex state of actual equipment operation, laying a precise feature foundation for subsequent fault prediction.

[0016] 2. The fault prediction mechanism constructed in this invention, based on the evolution trend of topological feature deviation and health index, combines statistical modeling and pattern matching to achieve a closed-loop fault early warning from "state perception" to "trend prediction," significantly improving the accuracy and foresight of fault prediction. This mechanism extracts a topological feature vector composed of global clustering coefficient, global efficiency, and node degree centrality standard deviation from a dynamically coupled network. Based on historical normal data, it establishes a multivariate Gaussian distribution deviation reference model. By calculating Mahalanobis distance and negative log-likelihood, it derives anomaly values ​​and maps them to the equipment coupling health index, thus realizing the monitoring of equipment health status. The system quantifies the perception of the state; simultaneously, it analyzes the short-term decay trend of the health index through trend estimation algorithms, combines dynamic time warping algorithms with a pre-stored fault mode library for similarity matching, and performs fault prediction logic judgment based on the overall decay trend and matching distance. This not only distinguishes between occasional data fluctuations and systemic deterioration caused by faults, but also accurately identifies fault types. It achieves a leap from single state judgment to trend-based and precise fault early warning, effectively capturing fault precursors in advance, providing accurate decision-making basis for preventive maintenance of equipment, significantly reducing the probability of missed or false faults, and improving the stability and continuity of water level monitoring equipment operation. Attached Figure Description

[0017] Figure 1 This is a flowchart illustrating a method for predicting faults in water level monitoring equipment based on multi-source data fusion, provided in an embodiment of the present invention. The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0018] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.

[0019] This application provides a method for predicting faults in water level monitoring equipment based on multi-source data fusion. The execution entity of this method includes, but is not limited to, at least one of the following electronic devices that can be configured to execute the method provided in this application: a server, a terminal, etc. In other words, the method for predicting faults in water level monitoring equipment based on multi-source data fusion can be executed by software or hardware installed on a terminal device or a server device. The server includes, but is not limited to, a single server, a server cluster, a cloud server, or a cloud server cluster. The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms.

[0020] Reference Figure 1 The diagram shown is a flowchart illustrating a method for predicting faults in water level monitoring equipment based on multi-source data fusion, according to an embodiment of the present invention. In this embodiment, the method for predicting faults in water level monitoring equipment based on multi-source data fusion includes: S1, acquire multi-source time-series data from water level monitoring equipment, and perform stabilization processing on the multi-source time-series data to obtain a weakly stationary multivariate sequence; In this embodiment of the invention, the step of acquiring multi-source time-series data from a water level monitoring device and performing stationarization processing on the multi-source time-series data to obtain a weakly stationary multivariate sequence includes: Acquire raw multi-source time-series data generated by the target water level monitoring equipment within a preset sampling period; The original multi-source time series data is subjected to time alignment and outlier removal operations to obtain preprocessed multi-source time series data; The preprocessed multi-source time series data were subjected to white noise self-testing method to obtain the data stationarity test results. Based on the data stationarity test results, the sequences that are determined to be non-white noise are subjected to first-order differencing, and the differencing sequences are iteratively tested for white noise until they pass the white noise test, so as to obtain the weakly stationary multivariate sequence.

[0021] It should be noted that the original multi-source time-series data refers to a set of multi-data elements with timestamps, obtained synchronously or asynchronously from the target water level monitoring equipment and its directly associated physical environment through heterogeneous sensors. Specifically, it includes: core monitoring data: water level height time series collected by the water level gauge, which directly reflects the physical quantity of the monitoring target of the equipment; key environmental disturbance data: rainfall intensity or cumulative rainfall time series collected by the rain gauge, which is the main external driving factor causing water level changes; equipment electrical status data: equipment operating voltage or battery voltage time series collected by the voltage sensor, whose abnormal fluctuations may indicate power supply aging, poor contact or load failure; and equipment communication status data: wireless signal reception strength and signal-to-noise ratio time series provided by the equipment communication module.

[0022] It should be noted that the time alignment operation is based on the timestamps of different sensing units. The original data sequences such as water level, rainfall, equipment voltage, and signal strength are resampled and aligned to the same time point through a preset time grid with equal intervals, so as to eliminate the timing misalignment caused by sensor sampling frequency or transmission delay.

[0023] It should be noted that the outlier removal operation is based on a preset statistical range threshold. It identifies and removes extreme values ​​in each data sequence that exceed the set standard deviation in order to eliminate noise caused by instantaneous sensor interference and data acquisition errors.

[0024] Furthermore, the preset statistical range threshold is usually set as the mean of the data series plus or minus three standard deviations.

[0025] It should be noted that by calculating the autocorrelation coefficients of the sequence at multiple time lag points and combining this information to form a single statistic, we can determine whether the entire sequence can be considered white noise. The mathematical expression for the white noise test is as follows: ; In the formula, Here, y is the autocorrelation coefficient, n is the number of samples in the sequence, and y is the number of samples in the sequence. t Let be the observation value at time t. Let be the sample mean of the entire sequence, and Q be the Q statistic. To test the maximum lag order set, l is the time interval between the current observation and the l-th previous observation.

[0026] It should be noted that the calculated Q statistic is compared with the theoretical critical value of the chi-square distribution. If the Q statistic is greater than the critical value at the selected significance level of 5%, the null hypothesis is rejected, and the series is considered not to be white noise, i.e., there is significant autocorrelation. Otherwise, the null hypothesis is accepted, and the series is considered to be white noise. The degrees of freedom are k. For example, when the degrees of freedom are set to k=10 and the significance level is 5%, the corresponding theoretical critical value can be found to be 18.31 by looking up the table.

[0027] It should be noted that the results of the data stationarity test are the conclusions of the above tests. If the conclusion is "the sequence is white noise", then the sequence is considered to have the characteristics of a stationary sequence and no difference processing is required. If the conclusion is "the sequence is not white noise", then the sequence is considered to have autocorrelation or trend and difference stationarization processing is required.

[0028] Furthermore, the differential stabilization process is performed iteratively. That is, after performing a first-order difference on the original non-white noise sequence, a new sequence is obtained, and the new sequence is tested again. This process is repeated until the new sequence passes the white noise test. The resulting sequence is a weakly stationary multivariate sequence that meets the modeling requirements.

[0029] It should be noted that differential stabilization is a process used to eliminate trend or seasonal components in non-stationary sequences by calculating the difference between the current data value and the previous data value. Essentially, it constructs a new stationary sequence by calculating the change in adjacent observations. The number of differential operations is determined based on the sequence characteristics.

[0030] S2, Based on the weakly stationary multivariate sequence, a dynamic coupling relationship network characterizing the time-varying interaction relationship between sensing units inside the device is constructed through a dynamic conditional covariance model and Granger causality test. In this embodiment of the invention, the construction of a dynamic coupling relationship network characterizing the time-varying interaction relationship between sensing units within the device, based on the weakly stationary multivariate sequence and through a dynamic conditional covariance model and Granger causality test, includes: Based on the weakly stationary multivariate sequence, the time-varying covariance matrix among the sensing variables is estimated by a dynamic conditional covariance model. Extract the time-varying correlation coefficient matrix representing the change in the strength of the linear association between variables over time from the time-varying covariance matrix; Based on the weakly stationary multivariate sequence, within a preset sliding time window, the time-varying causal strength matrix representing the lead-lag causal relationship between variables is calculated using the multivariate Granger causality test algorithm. By fusing the time-varying correlation coefficient matrix and the time-varying causal strength matrix, a directed weighted dynamic coupling relationship network is constructed.

[0031] It should be noted that the time-varying covariance matrix is ​​a quantification matrix that describes the degree of coordinated fluctuation between multiple sensing variables at any given time. The magnitude and sign of its off-diagonal elements reflect the instantaneous intensity and direction of the mutual influence between variables. For water level monitoring equipment, these variables include water level, rainfall, equipment voltage, signal strength, etc.

[0032] It should be noted that the implementation process of the dynamic conditional covariance model is as follows: First, a GARCH(1,1) model is established for each sensor variable sequence i in the weakly stationary multivariate sequence to calculate the time-varying conditional variance of the variable. The mathematical expression of the GARCH(1,1) model is as follows: ; In the formula, ω represents the conditional variance of variable i at time t. i It is a constant term, α i Measuring new information items The effect on the current variance, i.e. the squared residual of the previous time step, β i Measure the variance of the previous period The lasting impact, where t is the time index.

[0033] Furthermore, the process of establishing the GARCH(1,1) model is as follows: based on the sequence of each sense variable, the model parameters ω are fitted using the maximum likelihood estimation method. i α i and β i This allows us to calculate the conditional variance of the sequence at each time step.

[0034] Furthermore, the model parameters ω are fitted using the maximum likelihood estimation method. i α i and β i The set of parameter estimates that maximizes the likelihood function value is denoted as the model parameters θ. i .

[0035] It should be noted that the core steps of the dynamic conditional covariance model also include calculating the standardized residuals. And a dynamic conditional correlation matrix Q is constructed based on the standardized residual sequence. t Its mathematical expression is: ; In the formula, Q t This is the dynamic conditional correlation matrix. It is the unconditional covariance matrix of the standardized residuals, θ1 and θ2 are model parameters, and u t-1 It is the standardized residual vector of all variables at the previous moment.

[0036] Furthermore, the expression for the dynamic conditional correlation matrix describes how the correlation coefficient dynamically adjusts over time: Part 1 Represents the long-term average correlation coefficient level; Part Two Represents the impact of recent variable co-shocks; Part 3 θ2Q t-1 This represents the continued influence of previous correlations, and the entire expression realizes the time-varying evolution of the correlation coefficient.

[0037] It should be noted that the time-varying correlation coefficient matrix P t By analyzing the dynamic conditional correlation matrix Q t The elements are obtained through standardization. The calculation formula is: ; In the formula, q ij,t It is Q t The element in the i-th row and j-th column of the matrix, q ii,t and q jj,t These are the diagonal elements corresponding to variables i and j, respectively.

[0038] It should be noted that the elements of the time-varying causality strength matrix are calculated using the following formula when testing for rejection of the null hypothesis: ; In the formula, G ij,t G is the time-varying causal intensity matrix t The element in the i-th row and j-th column, SSE r It is the sum of squared residuals of the constrained model, SSE u It is the sum of squared residuals of the unrestricted model; if the test does not reject the null hypothesis, then G ij,t =0.

[0039] Furthermore, the time-varying causality strength quantifies the time-varying causality strength of variable X within a sliding window [tL,t]. j After incorporating past information into the model, for variable X i The closer the improvement in prediction accuracy is to 1, the stronger the causality. This is because G... ij,t ≠G ji,t This reflects the directionality of causality. L is a pre-set positive integer parameter that represents the length of historical data used for model estimation and calculation, with a value of 24.

[0040] It should be noted that the process of constructing a dynamically coupled relationship network is as follows: each sensed variable is used as a network node; for any two nodes i and j, if their causal strength G ij,t If the value is greater than 0, then a directed edge is constructed from node j to node i; the weight of this directed edge is set to the absolute value of the corresponding time-varying correlation coefficient. .

[0041] Furthermore, in the dynamic coupling network, nodes represent different sensing units of the device, such as water level gauges, rain gauges, voltage sensors, and signal modules; directed edges represent statistically significant, temporally sequential influences of one unit on another; the weight of the edge represents the synchronicity or synergy of the numerical changes of the two when such influences occur; the network comprehensively reflects the complex, time-varying interaction relationships between the sensing units within the device from three dimensions: "intensity," "direction," and "dynamics."

[0042] In this embodiment of the invention, the mathematical expression of the multivariate Granger causality test algorithm is as follows: The multivariate Granger causality test algorithm consists of both restricted and unrestricted models. The mathematical expression for the constrained model is as follows: ; The mathematical expression for the unrestricted model is as follows: ; In the formula X i (t) and X j (tp) represents the observed values ​​of the i-th and j-th sensed variables in the weakly stationary multivariate sequences at times t and tp, respectively, where P is the lag order and A is the lag order. p and B p These are the model coefficients, e1(t) and e2(t) are the residual terms, and p is the order index.

[0043] It should be noted that the multivariate Granger causality test algorithm refers to the algorithm that, within a sliding window [tL,t], performs a causality test on any two variables X. i and X j By establishing two vector autoregressive models and performing an F-test, we can determine the variable X. j Does the past value of X affect the predictor variable X? i The current value has a statistically significant contribution.

[0044] Furthermore, the null hypothesis of the F-test is that variable X... j With variable X i No Granger causality, i.e., all B p If all coefficients are 0, and the test rejects the null hypothesis, then it is considered that there exists a coefficient from X. j To X i Granger causality.

[0045] Furthermore, A p and B pThese are coefficients automatically estimated by the model during the fitting process, used to quantify the influence weight of historical observations on the current value. e1(t) and e2(t) are the fitting residuals of the two models at time t, that is, the errors between the model predictions and the actual observations.

[0046] S3, extract the topological feature vector representing the stability of the network structure from the dynamic coupling relationship network; In this embodiment of the invention, extracting the topological feature vector representing the stability of the network structure from the dynamically coupled network includes: Based on the adjacency matrix of the dynamically coupled network, a global clustering coefficient is calculated to measure the tightness of the connections between nodes within the network. Based on the shortest path information between nodes in the dynamically coupled network, the global efficiency, which measures the efficiency of network information transmission, is calculated. Based on the number of connections of each node in the dynamic coupling network, the standard deviation of node degree centrality, which measures the uniformity of network connection distribution, is calculated. The global clustering coefficient, the global efficiency, and the standard deviation of node degree centrality are combined to generate a topological feature vector.

[0047] It should be noted that the calculation process of the global clustering coefficient is as follows: First, calculate the local clustering coefficient of each node in the network, that is, the actual number of edges between the node's neighboring nodes, divided by the total number of possible edges between these neighboring nodes; then, calculate the arithmetic mean of the local clustering coefficients of all nodes to obtain the global clustering coefficient of the network.

[0048] It should be noted that the calculation process for global efficiency is as follows: First, calculate the shortest path length d between any two distinct nodes i and j in the network. ij This is the minimum number of edges required to get from one node to another; then, the sum of the reciprocals of the shortest path lengths between all pairs of nodes is calculated and normalized. A higher global efficiency value indicates better overall network connectivity, meaning that the states of any two sensing units can influence each other more quickly. Lower efficiency may indicate a "bottleneck" in the network or that the connections have become sparse, which is usually a sign of a loose or abnormal coupling structure.

[0049] It should be noted that the calculation process for the standard deviation of nodal degree centrality is as follows: The larger the standard deviation of node degree centrality, the greater the difference in the number of connections between nodes. It is possible that a few "hub" nodes dominate a large number of connections, while other nodes are sparsely connected. The smaller the value, the more uniform the distribution of network connections. For water level monitoring equipment, an abnormal standard deviation of node degree centrality may mean that a certain sensing unit is strongly coupled with all other units, or that a certain unit has become isolated.

[0050] It should be noted that the topological feature vector quantifies the structural state of a dynamically coupled network from three complementary dimensions: local compactness, global connectivity, and uniform distribution. The stability of the network structure is reflected through these macroscopic topological properties. For example, a stable sensing coupled network may exhibit a moderate clustering coefficient, high global efficiency, and low degree distribution variation. The initiation or evolution of any fault may first cause a systematic deviation of these topological features. The topological feature vector is a key state indicator representing the health of the internal coupling structure of the device.

[0051] In this embodiment of the invention, the calculation of the standard deviation of node degree centrality, which measures the uniformity of network connectivity distribution, includes: First, calculate the degree centrality of each node in the network, i.e., the number of connections k of each node. i Then calculate the standard deviation of the degree centrality of all nodes, which is expressed mathematically as follows: ; In the formula, D t Let N be the standard deviation of the degree centrality of nodes at time t, and N be the total number of nodes, k i Let i be the degree of node i. Let be the average degree of all nodes, and i be the node index.

[0052] S4. Establish a deviation reference model based on the topological feature vectors of historical normal periods, and analyze and calculate the deviation of the topological feature vectors based on the deviation reference model to obtain the equipment coupling health index. In this embodiment of the invention, establishing a deviation reference model based on topological feature vectors from historical normal periods includes: Collect all topological feature vectors generated by the equipment during its historical normal and fault-free operation phase to form a set of historical normal topological feature vectors; The mean vector reflecting the central position of the set of historical normal topological feature vectors and the covariance matrix representing the dispersion and correlation of the data are calculated by using parameter estimation methods. The deviation reference model is defined by the mean vector and the covariance matrix.

[0053] It should be noted that the parameter estimation method uses the maximum likelihood estimation method to calculate the average value of all vectors in the set of historical normal topological feature vectors to obtain the mean vector, and calculates the average value of the outer product of the deviations between these vectors and the mean vector to obtain the covariance matrix.

[0054] Furthermore, the formula for calculating the mean vector is as follows: ; In the formula, μ is the mean vector, M is the total number of historical normal topological feature vectors, and T (m) It is the m-th topological feature vector in the set, where m is the index of the topological feature vector; By default, the historical normal topology feature vector is constructed by taking the normal topology feature vector of the past month.

[0055] It should be noted that the covariance matrix... The calculation formula is: ; In the formula, It is the covariance matrix, T (m) -μ is the deviation vector between the m-th eigenvector and the mean vector, T represents the transpose operation, M is the total number of historical normal topological eigenvectors, and m is the index of the topological eigenvector.

[0056] It should be noted that the deviation reference model M normal The mean vector μ and covariance matrix obtained from the estimation Common definition, namely M normal , representing a matrix with mean μ and covariance matrix of μ. The multivariate Gaussian distribution.

[0057] Furthermore, the mathematical expression for the multivariate Gaussian distribution is as follows: ; In the formula, M normal Let X be an arbitrary topological feature vector, K be the dimension of the topological feature vector, and μ be the mean vector, representing the deviation reference model. Let covariance matrix be the variance matrix. The determinant of the covariance matrix. Let T denote the inverse of the covariance matrix, and let T denote the transpose operation. Let K be a Gaussian distribution function, where the dimension K=3.

[0058] In this embodiment of the invention, the step of analyzing and calculating the deviation of the topological feature vector based on the deviation reference model to obtain the device coupling health index includes: Based on the deviation reference model and the topological feature vector at the current time, calculate the outlier value, which represents the negative logarithm of the probability density of the topological feature vector at the current time under the deviation reference model. Based on a preset scaling function, the anomaly value is mapped and transformed to obtain the device coupling health index.

[0059] It should be noted that the anomaly value is calculated based on the aforementioned deviation reference model and the current topological feature vector, by calculating the squared Mahalanobis distance and the negative log-likelihood value; the mathematical expression for calculating the squared Mahalanobis distance is as follows: ; In the formula, T is the squared Mahalanobis distance. t Let μ be the current topological feature vector, and μ be the mean vector. Let T be the covariance matrix, and let T denote the transpose operation. The inverse matrix of the covariance matrix; The mathematical expression for calculating the negative log-likelihood is as follows: ; In the formula, L t This is the anomaly value. It is the logarithm of the determinant of the covariance matrix, and K is the dimension of the topological eigenvectors. It is a constant.

[0060] Furthermore, the larger the negative log-likelihood value, the stronger the T... t The smaller the probability density under this distribution, the further the current network structure deviates from the normal pattern, and the higher the degree of abnormality.

[0061] It should be noted that the scaling function is a monotonically decreasing exponential function, used to scale the outlier value L. t Mapping to the [0,1] interval to generate an intuitive health index, its mathematical expression is: CHI t =e -λLt ; In the formula, CHI t λ is the device coupling health index, and λ is a preset positive scaling parameter used to control the sensitivity or decay rate of the mapping.

[0062] Furthermore, the scaling parameter defaults to 0.05, and its range is from 0.01 to 0.5, which can be adjusted manually.

[0063] S5. Based on the evolution trend of the device coupled health index and its matching degree with the pre-stored fault modes, generate a fault prediction and early warning.

[0064] In this embodiment of the invention, generating a fault prediction and early warning based on the evolution trend of the device coupled health index and its matching degree with pre-stored fault modes includes: Continuously monitor the device's coupled health index to obtain the health index sequence within the most recent time window; Based on the health index sequence, its short-term decline trend is calculated using a trend estimation algorithm; The health index sequence is dynamically time-warped similarity matched with typical fault mode sequences in the pre-stored fault mode library to obtain the matching distance. Based on the short-term decay trend and the matching distance, a fault prediction logic is executed to generate the fault prediction warning.

[0065] It should be noted that the health index sequence within the recent time window refers to the set of device coupled health indices arranged in chronological order from the current moment back to a predetermined length in the past, denoted as Q = {q1, q2, ..., q}. ω}, where ω is the sequence length, and the preset length is 24.

[0066] It should be noted that the trend estimation algorithm uses the derivative of the exponentially weighted moving average to calculate the short-term decay trend of the health index series. First, the exponentially weighted moving average of the health index series is calculated: EMA t =α·q t + (1-α)·EMA t-1 ; In the formula, EMA t Let q be the exponentially weighted moving average at time t. t Let α be the health index at time t, and EMA be the smoothing factor. t-1 This is the moving average value of the previous time step, where the smoothing factor is 0.1. Then, the derivative of the moving average sequence at the most recent moment is calculated as the short-term decaying trend. t ; When the short-term decay trend is negative and the absolute value is large, it indicates that the health index is declining rapidly, that is, the state of the equipment coupling relationship structure is deteriorating at an accelerated pace. This is an important precursor signal that a failure may be imminent.

[0067] It should be noted that dynamic time warping similarity matching is an algorithm used to measure the shape similarity between two time series that may have different lengths. Its calculation process is as follows: Let there be a real-time health index sequence Q = {q1, ..., q} ω} and failure mode sequence ; Build a The cumulative cost matrix D, where the matrix element D(i,j) represents the minimum cumulative alignment cost between the first i points of sequence Q and the first j points of sequence C; The initialization and recursive formulas for matrix D are as follows: ; In the formula, D(i,j) represents the minimum cumulative alignment cost between the first i points of sequence Q and the first j points of sequence C, (q i -c j ) 2 For point q i With point c j Local costs between them; Matching distance Dist DTW (Q, C) is defined as the square root of the final element value of the cumulative cost matrix, i.e. .

[0068] Furthermore, the algorithm uses dynamic programming to find an optimal twist path that non-linearly aligns the two sequences on the time axis, thereby minimizing their cumulative distance (q). i -c j ) 2 It measures the numerical difference between two sequences at corresponding points; the smaller the matching distance value, the more similar the shapes of the real-time health index sequence Q and the fault mode sequence C are, that is, the closer the current health evolution trajectory is to a certain known fault development mode.

[0069] It should be noted that the fault prediction logic makes a comprehensive judgment based on two conditions: first, whether the short-term decay trend exceeds a preset negative threshold; second, whether it is related to a certain fault mode C. K The matching distance is less than the preset similarity threshold; if and only if both conditions are met, it is determined that the device is evolving toward the fault type corresponding to the fault mode.

[0070] It should be noted that the fault prediction and early warning is an output information containing prediction conclusions. Its content includes at least: determining that the equipment has a fault risk, the predicted fault type, the early warning level, and the recommended maintenance measures. After the early warning is generated, it is sent to the monitoring center or the terminal of the maintenance personnel to realize early intervention and preventive maintenance of equipment faults.

[0071] In this embodiment of the invention, the pre-stored fault mode library is a database that stores the typical evolution trajectory of the equipment coupling health index before and after various typical faults in history. Each typical fault mode is represented by a health index sequence, which captures the characteristic change shape of the health index from the fault latency period to the fault occurrence period.

[0072] In the several embodiments provided by this invention, it should be understood that the disclosed method can be implemented in other ways.

[0073] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.

[0074] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence is the theory, method, and technology that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.

[0075] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A method for predicting faults in water level monitoring equipment based on multi-source data fusion, characterized in that, The method includes: S1, acquire multi-source time-series data from water level monitoring equipment, and perform stabilization processing on the multi-source time-series data to obtain a weakly stationary multivariate sequence; S2, Based on the weakly stationary multivariate sequence, a dynamic coupling relationship network characterizing the time-varying interaction relationship between sensing units within the device is constructed through a dynamic conditional covariance model and Granger causality test. The specific method is as follows: Based on the weakly stationary multivariate sequence, the time-varying covariance matrix among the sensing variables is estimated by a dynamic conditional covariance model. Extract the time-varying correlation coefficient matrix representing the change in the strength of the linear association between variables over time from the time-varying covariance matrix; Based on the weakly stationary multivariate sequence, within a preset sliding time window, the time-varying causal strength matrix representing the lead-lag causal relationship between variables is calculated using the multivariate Granger causality test algorithm. By fusing the time-varying correlation coefficient matrix and the time-varying causality strength matrix, a directed weighted dynamic coupling relationship network is constructed; S3, extracting topological feature vectors representing network structural stability from the dynamically coupled network, specifically including: Based on the adjacency matrix of the dynamically coupled network, a global clustering coefficient is calculated to measure the tightness of the connections between nodes within the network. Based on the shortest path information between nodes in the dynamically coupled network, the global efficiency, which measures the efficiency of network information transmission, is calculated. Based on the number of connections of each node in the dynamic coupling network, the standard deviation of node degree centrality, which measures the uniformity of network connection distribution, is calculated. The global clustering coefficient, the global efficiency, and the standard deviation of node degree centrality are combined to generate a topological feature vector. S4. Establish a deviation reference model based on the topological feature vectors of historical normal periods, and analyze and calculate the deviation of the topological feature vectors based on the deviation reference model to obtain the equipment coupling health index. S5. Based on the evolution trend of the device coupled health index and its matching degree with the pre-stored fault modes, generate a fault prediction and early warning.

2. The method for predicting faults in water level monitoring equipment based on multi-source data fusion as described in claim 1, characterized in that, The process of acquiring multi-source time-series data from water level monitoring equipment and performing stationarization processing on the multi-source time-series data to obtain a weakly stationary multivariate sequence includes: Acquire raw multi-source time-series data generated by the target water level monitoring equipment within a preset sampling period; The original multi-source time series data is subjected to time alignment and outlier removal operations to obtain preprocessed multi-source time series data; The preprocessed multi-source time series data were subjected to white noise self-testing method to obtain the data stationarity test results. Based on the data stationarity test results, the sequences that are determined to be non-white noise are subjected to first-order differencing, and the differencing sequences are iteratively tested for white noise until they pass the white noise test, so as to obtain the weakly stationary multivariate sequence.

3. The method for predicting faults in water level monitoring equipment based on multi-source data fusion as described in claim 1, characterized in that, The mathematical expression of the multivariate Granger causality test algorithm is as follows: The multivariate Granger causality test algorithm consists of both restricted and unrestricted models. The mathematical expression for the constrained model is as follows: ; The mathematical expression for the unrestricted model is as follows: ; In the formula and Let represent the observed values ​​of the i-th and j-th sensed variables in the weakly stationary multivariate sequences at times t and tp, respectively. It is the lag order. and These are model coefficients. and It is the residual term. This is the order index.

4. The method for predicting faults in water level monitoring equipment based on multi-source data fusion as described in claim 1, characterized in that, The calculation of the standard deviation of node degree centrality, which measures the uniformity of network connectivity distribution, includes: First, calculate the degree centrality of each node in the network, that is, the number of connections of each node. Then calculate the standard deviation of the degree centrality of all nodes, which is expressed mathematically as follows: ; In the formula, Let be the standard deviation of the nodal degree centrality at time t. The total number of nodes. Let i be the degree of node i. Let be the average degree of all nodes, and i be the node index.

5. The method for predicting faults in water level monitoring equipment based on multi-source data fusion as described in claim 1, characterized in that, The deviation reference model established based on topological feature vectors from historical normal periods includes: Collect all topological feature vectors generated by the equipment during its historical normal and fault-free operation phase to form a set of historical normal topological feature vectors; The mean vector reflecting the central position of the set of historical normal topological feature vectors and the covariance matrix representing the dispersion and correlation of the data are calculated by using parameter estimation methods. The deviation reference model is defined by the mean vector and the covariance matrix.

6. The method for predicting faults in water level monitoring equipment based on multi-source data fusion as described in claim 1, characterized in that, The process of analyzing and calculating the deviation of the topological feature vector based on the deviation reference model to obtain the device coupling health index includes: Based on the deviation reference model and the topological feature vector at the current time, calculate the outlier value, which represents the negative logarithm of the probability density of the topological feature vector at the current time under the deviation reference model. Based on a preset scaling function, the anomaly value is mapped and transformed to obtain the device coupling health index.

7. The method for predicting faults in water level monitoring equipment based on multi-source data fusion as described in claim 1, characterized in that, The generation of fault prediction and early warning based on the evolution trend of the device coupled health index and its matching degree with pre-stored fault modes includes: Continuously monitor the device's coupled health index to obtain the health index sequence within the most recent time window; Based on the health index sequence, its short-term decline trend is calculated using a trend estimation algorithm; The health index sequence is dynamically time-warped similarity matched with typical fault mode sequences in the pre-stored fault mode library to obtain the matching distance. Based on the short-term decay trend and the matching distance, a fault prediction logic is executed to generate the fault prediction warning.

8. The method for predicting faults in water level monitoring equipment based on multi-source data fusion as described in claim 7, characterized in that, The pre-stored fault mode library is a database that stores the typical evolution trajectory of the equipment coupling health index before and after various typical faults in history. Each typical fault mode is represented by a health index sequence, which captures the characteristic change shape of the health index from the fault latency period to the fault occurrence period.

Citation Information

Patent Citations

  • Remote fault diagnosis method and system for solar device

    CN120449057A

  • AI fusion processing method, system and equipment for state monitoring data of power transmission and transformation equipment and storage medium

    CN121542799A