Monitoring data anomaly diagnosis method and device fusing causal inference
Patent Information
- Application Number
- CN202610655237.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-13
- Publication Date
- 2026-09-04
AI Technical Summary
这类方法能够识别偏离正常模式的数据点并触发报警,但无法判断异常的根本原因,也即无法判断是传感器自身漂移、环境自然波动,还是真正预示泄漏或失稳等风险事件
[0010] The technical solution provided by the embodiments of this disclosure can include the following beneficial effects: By acquiring and preprocessing sensor data from the carbon dioxide storage area, an unsupervised anomaly detection algorithm is first used to automatically filter out suspected anomalies and aggregate them into time segments, quickly identifying the period of interest without manual annotation; then, a directed acyclic causal graph is constructed for each segment, and the causal effect strength of each upstream variable on the downstream anomaly variable is quantified based on a causal inference algorithm, thereby accurately locating the anomaly source node; finally, based on the type of source node, data change characteristics, and propagation path in the causal graph, high-risk events such as instrument failure, natural environmental fluctuations, carbon dioxide leakage, or geomechanical instability are intelligently distinguished. This method elevates anomaly detection from simple alarms to interpretable diagnoses, avoiding the high false alarm rate caused by the inability to distinguish causes in traditional methods, and providing maintenance personnel with clear root causes and propagation paths, significantly reducing the burden of manual investigation and decision-making delays, and achieving highly reliable, automated early warning and diagnosis of the safety status of carbon dioxide storage.
Smart Images

Figure CN122692818A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of carbon dioxide sequestration technology, and in particular to a method and apparatus for diagnosing monitoring data anomalies by integrating causal inference. Background Technology
[0002] In related technologies, sensor networks continuously generate massive amounts of time-series monitoring data in fields such as carbon dioxide geological storage, underground energy storage, and environmental monitoring. Currently, data analysis mainly relies on threshold alarms or unsupervised anomaly detection algorithms. These methods can identify data points that deviate from normal patterns and trigger alarms, but they cannot determine the root cause of the anomaly, i.e., whether it is sensor drift, natural environmental fluctuations, or a genuine indication of a risk event such as leakage or instability. Therefore, in practical applications, the inability to distinguish between faults and real risks often leads to a high false alarm rate. Maintenance personnel still need to conduct extensive on-site investigations and manual diagnoses, increasing decision-making delays and maintenance burdens, and failing to meet the requirements of high-risk scenarios for accurate and operable early warnings. Summary of the Invention
[0003] To overcome the problems existing in related technologies, this disclosure provides a method and apparatus for diagnosing abnormal monitoring data by integrating causal inference.
[0004] According to a first aspect of the present disclosure, a method for diagnosing monitoring data anomalies by incorporating causal inference is provided, comprising:
[0005] Acquire multivariate time-series monitoring data collected by a sensor network deployed in the carbon dioxide storage area, and preprocess the multivariate time-series monitoring data; An unsupervised anomaly detection algorithm is used to model and score the preprocessed data. Sample points with scores higher than a preset threshold are marked as suspected anomalies, and the time interval formed by multiple consecutive suspected anomalies is determined as a suspected anomaly time segment. For each suspected abnormal time segment, all monitoring data within the suspected abnormal time segment are used as associated data, and a causal discovery algorithm is used to construct a directed acyclic causal graph between each monitoring variable based on the associated data. On the directed acyclic causal graph, variables whose data change amplitude exceeds a preset amplitude within the suspected abnormal time segment are marked as downstream abnormal variables. A causal inference algorithm is used to calculate the causal effect strength of each upstream monitoring variable on the downstream abnormal variable. Based on the causal effect strength, the initial disturbance node that leads to the propagation of the abnormality in the directed acyclic causal graph is located to obtain the abnormality source node. Based on the node type, data change characteristics, and causal propagation pattern of the anomaly source node, the root cause of the anomaly event is diagnosed; the root cause includes instrument failure, natural environmental fluctuations, and high-risk events; the high-risk events include carbon dioxide leakage or geomechanical instability; the causal propagation pattern is the causal propagation pattern reflected by the directed paths from the anomaly source node to each downstream anomaly variable in the directed acyclic causal graph.
[0006] According to a second aspect of the present disclosure, a monitoring data anomaly diagnosis device integrating causal inference is provided, comprising: The acquisition unit is used to acquire multivariate time-series monitoring data collected by a sensor network deployed in the carbon dioxide storage area, and to preprocess the multivariate time-series monitoring data. The labeling unit is used to model and score the preprocessed data using an unsupervised anomaly detection algorithm. Sample points with scores higher than a preset threshold are labeled as suspected anomalies, and the time interval formed by multiple consecutive suspected anomalies is defined as a suspected anomaly time segment. The construction unit is used to, for each suspected abnormal time segment, take all monitoring data within the suspected abnormal time segment as associated data, and use a causal discovery algorithm to construct a directed acyclic causal graph between each monitoring variable based on the associated data; The positioning unit is used to mark variables whose data change amplitude exceeds a preset amplitude within the suspected abnormal time segment as downstream abnormal variables on the directed acyclic causal graph, calculate the causal effect strength of each upstream monitoring variable on the downstream abnormal variable using a causal inference algorithm, and locate the initial disturbance node that leads to the propagation of the abnormality in the directed acyclic causal graph based on the causal effect strength to obtain the abnormality source node. The diagnostic unit is used to diagnose the root cause of anomalies based on the node type of the anomaly source node, data change characteristics, and causal propagation pattern. The root cause includes instrument failure, natural environmental fluctuations, and high-risk events. The high-risk events include carbon dioxide leaks or geomechanical instability. The causal propagation pattern is the causal propagation pattern reflected by the directed paths from the anomaly source node to each downstream anomaly variable in the directed acyclic causal graph.
[0007] According to a third aspect of the present disclosure, an electronic device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the method as described in any one of the first aspects.
[0008] According to a fourth aspect of the present disclosure, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the method as described in any one of the first aspects.
[0009] According to a fifth aspect of the present disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the method as described in any one of the first aspects.
[0010] The technical solution provided by the embodiments of this disclosure can include the following beneficial effects: By acquiring and preprocessing sensor data from the carbon dioxide storage area, an unsupervised anomaly detection algorithm is first used to automatically filter out suspected anomalies and aggregate them into time segments, quickly identifying the period of interest without manual annotation; then, a directed acyclic causal graph is constructed for each segment, and the causal effect strength of each upstream variable on the downstream anomaly variable is quantified based on a causal inference algorithm, thereby accurately locating the anomaly source node; finally, based on the type of source node, data change characteristics, and propagation path in the causal graph, high-risk events such as instrument failure, natural environmental fluctuations, carbon dioxide leakage, or geomechanical instability are intelligently distinguished. This method elevates anomaly detection from simple alarms to interpretable diagnoses, avoiding the high false alarm rate caused by the inability to distinguish causes in traditional methods, and providing maintenance personnel with clear root causes and propagation paths, significantly reducing the burden of manual investigation and decision-making delays, and achieving highly reliable, automated early warning and diagnosis of the safety status of carbon dioxide storage.
[0011] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0012] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.
[0013] Figure 1 This is a flowchart illustrating a monitoring data anomaly diagnosis method that integrates causal inference, according to an exemplary embodiment.
[0014] Figure 2 This is a block diagram illustrating a monitoring data anomaly diagnostic device that integrates causal inference, according to an exemplary embodiment.
[0015] Figure 3 This is a block diagram illustrating an apparatus for a monitoring data anomaly diagnosis method that integrates causal inference, according to an exemplary embodiment. Detailed Implementation
[0016] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. Rather, they are merely examples of apparatuses and methods consistent with some aspects of the invention as detailed in the appended claims.
[0017] The terminology used in this disclosure is for the purpose of describing particular embodiments only and is not intended to be limiting of the present disclosure. The singular forms “a” and “the” as used in this disclosure and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise.
[0018] It should be understood that although the terms first, second, third, etc., may be used to describe various information in embodiments of this disclosure, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first information may also be referred to as second information without departing from the scope of embodiments of this disclosure, and similarly, second information may also be referred to as first information. Depending on the context, the words “if” and “suppose” as used herein may be interpreted as “when”, “when”, or “in response to a determination”.
[0019] Furthermore, various forms of processes shown in the embodiments of this disclosure can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and no limitation is imposed herein.
[0020] It should be noted that the collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the technical solution disclosed herein all comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0021] Figure 1 This is a flowchart illustrating a monitoring data anomaly diagnosis method that integrates causal inference, according to an exemplary embodiment, such as... Figure 1 As shown, it should be noted that the monitoring data anomaly diagnosis method based on fused causal inference in this embodiment is applied in the monitoring data anomaly diagnosis device based on fused causal inference. Figure 1 As shown, the method may include the following steps: Step 101: Obtain multivariate time-series monitoring data collected by the sensor network deployed in the carbon dioxide storage area, and preprocess the multivariate time-series monitoring data.
[0022] In some embodiments of this disclosure, missing values in multivariate time-series monitoring data are filled by interpolation or forward padding, detection data that exceeds the sensor's range are removed, and the timestamps of data from different sensors are aligned.
[0023] Specifically, sensor networks may include pressure gauges, thermometers, strain gauges, carbon dioxide concentration sensors, etc., with various sensors generating multivariate time-series data at different sampling frequencies. The raw data often contains missing values due to communication interruptions, gross errors exceeding reasonable ranges caused by momentary sensor failures, and inconsistencies in time references between different devices.
[0024] By using interpolation methods (such as linear interpolation and spline interpolation) or forward imputation to handle missing values, the continuity of the time series can be maintained; removing data points that are significantly beyond the physical range of the sensor can avoid the interference of spurious extreme values on subsequent modeling; aligning the timestamps of multi-source data ensures that variables are comparable at the same time cross-section.
[0025] The above steps establish a high-quality data input pipeline, filter out obvious errors and noise in the raw data, improve the data signal-to-noise ratio, and provide clean and well-organized input data for all subsequent analysis steps, thereby ensuring the accuracy of anomaly detection and causal inference.
[0026] Step 102: An unsupervised anomaly detection algorithm is used to model and score the preprocessed data. Sample points with scores higher than a preset threshold are marked as suspected anomalies, and the time interval formed by multiple consecutive suspected anomalies is determined as a suspected anomaly time segment.
[0027] In some embodiments of this disclosure, the unsupervised anomaly detection algorithm in step 102 may employ an isolated forest or a deep autoencoder.
[0028] As an example, when using an isolated forest, the anomaly score is calculated based on the average path length of the sample in the isolated tree; the shorter the average path length, the higher the anomaly score. When using a deep autoencoder, the anomaly score is calculated based on the reconstruction error of the sample; the larger the reconstruction error, the higher the anomaly score.
[0029] Specifically, the isolation forest constructs multiple isolation trees by randomly and recursively segmenting the feature space. Normal data needs to be segmented multiple times to be isolated, while abnormal data is more easily isolated quickly. Therefore, the shorter the average path length, the higher the abnormality score of the sample.
[0030] A deep autoencoder trains a model using normal historical data from the carbon dioxide storage region, enabling the model to learn the reconstruction mapping under normal patterns. For samples deviating from the normal pattern, the reconstruction error increases significantly, and this error is directly used as the anomaly score. The preset threshold can be dynamically determined based on a specified quantile or multiple of the mean of the anomaly score distribution of normal samples. Samples with scores higher than the threshold are marked as suspected anomalies, and the time interval formed by multiple temporally consecutive suspected anomalies is defined as a suspected anomaly time segment.
[0031] It should be noted that the above steps enable automated cleaning and preliminary screening of massive amounts of monitoring data. Without the need for a large number of manually labeled abnormal samples, key event signals that deviate significantly from the normal pattern can be quickly and effectively identified, greatly reducing the workload of manual data review and identifying key targets for subsequent in-depth diagnosis.
[0032] Step 103: For each suspected abnormal time segment, all monitoring data within the suspected abnormal time segment are used as associated data, and a causal discovery algorithm is used to construct a directed acyclic causal graph between each monitoring variable based on the associated data.
[0033] Specifically, for each marked suspected anomalous time segment, data from all sensors (pressure, temperature, strain, carbon dioxide concentration, etc.) within that time interval are extracted to form a correlation data matrix. A causal discovery algorithm, such as the LINGAM algorithm for linear non-Gaussian acyclic models, is used to analyze the statistical dependencies between variables, resulting in a directed acyclic graph (DAG). Each node in the graph represents a monitored variable, and directed edges indicate the direction of the causal relationship between variables (e.g., changes in formation pressure lead to changes in strain gauge readings).
[0034] As an example, the direction of causality can be constrained and verified by combining physical knowledge from the field of carbon dioxide sequestration, such as pressure changes should lead to flow rate changes rather than the other way around.
[0035] It should be noted that by constructing a causal graph among the monitored variables, the mutual influence and transmission paths between various physical quantities during the occurrence of abnormal events are clearly revealed. Isolated anomalies are connected into a logically related network, providing a structured causal graph for subsequent anomaly source tracing.
[0036] Step 104: On the directed acyclic causal graph, variables whose data changes exceeding a preset range within a suspected abnormal time segment are marked as downstream abnormal variables. A causal inference algorithm is used to calculate the causal effect strength of each upstream monitoring variable on the downstream abnormal variable. Based on the causal effect strength, the initial disturbance node that leads to the propagation of the abnormality in the directed acyclic causal graph is located, and the abnormality source node is obtained.
[0037] In some embodiments of this disclosure, the causal inference algorithm in step 104 is counterfactual reasoning. Specifically, step 104 may include the following steps: For each upstream monitoring variable in the directed acyclic causal graph, counterfactual data is generated when the upstream monitoring variable takes different values through counterfactual reasoning. The strength of the causal effect of the upstream monitoring variable on the downstream abnormal variable when it changes is calculated. All upstream monitoring variables are traversed, and the upstream monitoring variable with the strongest causal effect is determined as the initial disturbance node, thus obtaining the abnormal source node.
[0038] It should be noted that variables whose data changes exceed a preset range refer to those whose numerical variation within a suspected abnormal time segment exceeds a predetermined multiple of the variable's historical fluctuation range during normal periods. These variables are marked as downstream anomalous variables and used as targets for causal effect analysis.
[0039] The basic idea of counterfactual reasoning is to keep other variables in the causal graph constant, change the value of a certain upstream variable, generate counterfactual data for different values of that variable, and then observe the changes in the response of downstream anomalous variables to quantify the strength of the causal effect of the upstream variable on the downstream anomalous variable. After traversing all upstream nodes, the node with the strongest effect is the most likely initial source of disturbance.
[0040] It should be noted that the above has achieved a leap from correlation to causation, making anomaly localization no longer based on empirical guesses, but on data-driven quantitative analysis, which has higher credibility and accuracy, and can accurately trace the source of anomalies and distinguish between single sensor failures and multi-physical field coupling changes.
[0041] Step 105: Based on the node type, data change characteristics, and causal propagation pattern of the anomaly source node, diagnose the root cause of the anomaly event.
[0042] The root causes include instrument malfunction, natural environmental fluctuations, and high-risk events; high-risk events include carbon dioxide leaks or geomechanical instability; the causal propagation pattern is the causal propagation pattern reflected by the directed paths from the source node of the anomaly to each downstream anomaly variable in the directed acyclic causal graph.
[0043] In this embodiment, the root cause is determined by comprehensively considering the sensor type corresponding to the anomaly source node, the data change pattern of that node within the abnormal time segment (e.g., constant step change, slow recoverable change, or drastic deviation), and the propagation range and direction reflected by the directed path from that node to the downstream anomaly variable in the causal graph. If the source node is singular, the change is step-like, and the downstream impact range is minimal, it is diagnosed as an instrument malfunction; if the change is slow and recoverable, and the propagation path conforms to a geological or hydrological cycle model, it is diagnosed as a natural environmental fluctuation; if the source is a safety-critical variable, the change is drastic, and the causal path points to multiple physical quantities and conforms to the causal chain template of leakage or instability, it is diagnosed as a high-risk event. The technical significance of this step lies in fusing the causal graph, node attributes, and dynamic characteristics to output an operable diagnostic conclusion, avoiding the limitation of only providing alarms without being able to explain the cause.
[0044] In some embodiments of this disclosure, step 105 may specifically include the following sub-steps: Step e1: If the sensor type corresponding to the anomaly source node is a single type, and the data change pattern of the node within the abnormal time segment is a step change followed by a constant value, and the out-degree of the node in the directed acyclic causal graph does not exceed the preset out-degree threshold, and the proportion of the number of affected downstream anomaly variables to the total number of sensors does not exceed the preset proportion threshold, then it is diagnosed as an instrument malfunction; if the data change pattern of the anomaly source node is that the absolute value of the rate of change is lower than the normal fluctuation threshold, there is no abrupt inflection point, and it can recover to the baseline range after the abnormal time segment ends, and the distribution of downstream anomaly variables of the node in the directed acyclic causal graph is consistent with the propagation path predicted by the known geological or hydrological cycle model of the carbon dioxide sequestration area, then it is diagnosed as a natural environmental fluctuation.
[0045] Step e2: If the sensor corresponding to the anomaly source node belongs to any of the safety-critical variables among wellhead pressure, caprock strain, or carbon dioxide concentration, and the data change pattern of the node is a continuous deviation from the baseline exceeding a predetermined multiple of the normal fluctuation range, and the change rate exceeds a predetermined multiple of the historical maximum change rate, and there is a directed path in the directed acyclic causal graph pointing from the node to multiple downstream anomaly variables of different physical quantities, and the path conforms to the predefined causal chain template of carbon dioxide leakage or geomechanical instability, then it is diagnosed as a high-risk event. The risk level of the high-risk event is calculated by weighting the three indicators of anomaly amplitude, affected area, and change rate, and is divided into three levels: low, medium, and high according to the preset level classification rules. Among them, the anomaly amplitude is defined as the ratio of the difference between the node peak value and the baseline to the baseline standard deviation, the affected area is defined as the proportion of the number of downstream anomaly variables to the total number of sensors, and the change rate is defined as the maximum change of the node per unit time.
[0046] As an example, the criteria for judging instrument failure are: the source node is isolated (low degree), there are few affected variables (low proportion), and the data shows a step constant characteristic, which is consistent with the typical behavior of sensor drift or jamming.
[0047] The criteria for judging natural environmental fluctuations are: slow changes, no sudden changes, reversibility, and a propagation path consistent with known geomechanical or hydrological periodic patterns, such as the periodic pressure fluctuations caused by tides.
[0048] The criteria for judging high-risk events are: the source points to safety-critical variables, the degree and rate of deviation are significant, and the causal propagation path conforms to the physical model of leakage or instability (such as a drop in wellhead pressure accompanied by an increase in shallow carbon dioxide concentration).
[0049] For high-risk events, the abnormal magnitude, scope of impact, and rate of change are further weighted and calculated, and low, medium, and high risk labels are output according to preset classification rules (such as equal-width binning or quantile binning).
[0050] It should be noted that transforming complex causal analysis results into intuitive and easy-to-understand diagnostic conclusions and risk warnings significantly improves the interpretability and accuracy of the early warning system. By intelligently distinguishing between faults, interference, and real risks, the false alarm rate is greatly reduced, enabling maintenance personnel to quickly focus on genuine risk events and take targeted countermeasures based on the risk level (such as on-site instrument calibration, enhanced monitoring, or activation of emergency plans). This greatly reduces the burden of manual diagnosis and improves the efficiency and reliability of carbon dioxide storage safety management.
[0051] According to the monitoring data anomaly diagnosis method based on causal inference proposed in this disclosure, sensor data from the carbon dioxide storage area is acquired and preprocessed. First, an unsupervised anomaly detection algorithm automatically filters out suspected anomalies and aggregates them into time segments, quickly identifying the relevant time periods without manual annotation. Then, a directed acyclic causal graph is constructed for each segment, and the causal effect strength of each upstream variable on the downstream anomaly variable is quantified based on a causal inference algorithm, thereby accurately locating the anomaly source node. Finally, based on the type of the source node, data change characteristics, and propagation paths in the causal graph, high-risk events such as instrument malfunction, natural environmental fluctuations, carbon dioxide leakage, or geomechanical instability are intelligently distinguished. This method elevates anomaly detection from simple alarms to interpretable diagnosis, avoiding the high false alarm rate caused by the inability to distinguish causes in traditional methods, and providing maintenance personnel with clear root causes and propagation paths. This significantly reduces the burden of manual investigation and decision-making delays, achieving highly reliable and automated early warning and diagnosis of the safety status of carbon dioxide storage.
[0052] Figure 2This is a block diagram illustrating a monitoring data anomaly diagnostic device that integrates causal inference, according to an exemplary embodiment. (Refer to...) Figure 2 The device includes an acquisition unit 201, a marking unit 202, a construction unit 203, a positioning unit 204, and a diagnostic unit 205.
[0053] The acquisition unit 201 is used to acquire multivariate time-series monitoring data collected by the sensor network deployed in the carbon dioxide storage area and to preprocess the multivariate time-series monitoring data. The labeling unit 202 is used to model and score the preprocessed data using an unsupervised anomaly detection algorithm, mark sample points with scores higher than a preset threshold as suspected anomalies, and determine the time interval formed by multiple consecutive suspected anomalies as a suspected anomaly time segment. Construction unit 203 is used to take all monitoring data within each suspected abnormal time segment as associated data and use a causal discovery algorithm to construct a directed acyclic causal graph between each monitoring variable based on the associated data for each suspected abnormal time segment. The positioning unit 204 is used to mark variables whose data change amplitude exceeds a preset amplitude in a directed acyclic causal graph as downstream anomalous variables, and to use a causal inference algorithm to calculate the causal effect strength of each upstream monitoring variable on the downstream anomalous variable. Based on the causal effect strength, the initial disturbance node that causes the propagation of the anomaly in the directed acyclic causal graph is located to obtain the anomaly source node. The diagnostic unit 205 is used to diagnose the root cause of anomalies based on the node type of the anomaly source node, data change characteristics, and causal propagation pattern. Root causes include instrument failure, natural environmental fluctuations, and high-risk events. High-risk events include carbon dioxide leaks or geomechanical instability. The causal propagation pattern is the causal propagation pattern reflected by the directed paths from the anomaly source node to each downstream anomaly variable in the directed acyclic causal graph.
[0054] In some embodiments of this disclosure, the acquisition unit 201 may specifically be used for: Missing values in multivariate time-series monitoring data are filled using interpolation or forward imputation, detection data exceeding the sensor's range are removed, and data timestamps from different sensors are aligned.
[0055] In some embodiments of this disclosure, the unsupervised anomaly detection algorithm employs an isolated forest or a deep autoencoder.
[0056] In some embodiments of this disclosure, the diagnostic unit 205 may specifically be used for: If the sensor type corresponding to the abnormal source node is a single type, and the data change pattern of the node in the abnormal time segment is a step change followed by a constant value, and the out-degree of the node in the directed acyclic causal graph does not exceed the preset out-degree threshold, and the proportion of the number of affected downstream abnormal variables to the total number of sensors does not exceed the preset proportion threshold, then it is diagnosed as an instrument failure. If the data change pattern of the abnormal source node is that the absolute value of the rate of change is lower than the normal fluctuation threshold, there is no abrupt inflection point, and it can recover to the baseline range after the abnormal time segment ends, and the distribution of the downstream abnormal variables of the node in the directed acyclic causal graph is consistent with the propagation path predicted by the known geological or hydrological cycle model of the carbon dioxide sequestration area, then it is diagnosed as natural environmental fluctuation.
[0057] In some embodiments of this disclosure, the diagnostic unit 205 may specifically be used for: If the sensor corresponding to the abnormal source node belongs to any of the safety-critical variables such as wellhead pressure, caprock strain, or carbon dioxide concentration, and the data change pattern of the node is a continuous deviation from the baseline by a predetermined multiple of the normal fluctuation range, and the rate of change exceeds a predetermined multiple of the historical maximum rate of change, and there is a directed path in the directed acyclic causal graph pointing from the node to multiple downstream abnormal variables of different physical quantities, and the path conforms to the predefined causal chain template of carbon dioxide leakage or geomechanical instability, then it is diagnosed as a high-risk event. The risk level of high-risk events is calculated by weighting three indicators: abnormal amplitude, affected area, and rate of change, and then classified into three levels: low, medium, and high according to the preset level classification rules. Among them, abnormal amplitude is defined as the ratio of the difference between the node peak value and the baseline to the baseline standard deviation; affected area is defined as the proportion of the number of downstream abnormal variables to the total number of sensors; and rate of change is defined as the maximum change of the node per unit time.
[0058] In some embodiments of this disclosure, the causal inference algorithm is counterfactual reasoning, and the positioning unit 204 can specifically be used for: For each upstream monitoring variable in the directed acyclic causal graph, counterfactual data is generated when the upstream monitoring variable takes different values through counterfactual reasoning, and the strength of the causal effect of the upstream monitoring variable on the downstream abnormal variable when it changes is calculated. By iterating through all upstream monitoring variables, the upstream monitoring variable with the strongest causal effect is identified as the initial disturbance node, thus obtaining the anomaly source node.
[0059] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.
[0060] According to the monitoring data anomaly diagnosis device integrating causal inference proposed in this disclosure, by acquiring and preprocessing sensor data from the carbon dioxide storage area, an unsupervised anomaly detection algorithm is first used to automatically filter out suspected anomalies and aggregate them into time segments, quickly identifying the period of interest without manual annotation. Then, a directed acyclic causal graph is constructed between monitoring variables for each segment, and the causal effect strength of each upstream variable on the downstream anomaly variable is quantified based on the causal inference algorithm, thereby accurately locating the anomaly source node. Finally, based on the type of source node, data change characteristics, and propagation path in the causal graph, the device intelligently distinguishes between instrument malfunction, natural environmental fluctuations, and high-risk events such as carbon dioxide leakage or geomechanical instability. This method elevates anomaly detection from simple alarms to interpretable diagnosis, avoiding the high false alarm rate caused by the inability to distinguish causes in traditional methods, and providing maintenance personnel with clear root causes and propagation paths, significantly reducing the burden of manual investigation and decision-making delays, and achieving highly reliable, automated early warning and diagnosis of the safety status of carbon dioxide storage.
[0061] Figure 3 This is a block diagram illustrating an apparatus for a monitoring data anomaly diagnosis method that integrates causal inference, according to an exemplary embodiment. For example, apparatus 300 may be an electronic device, such as a mobile phone, computer, digital broadcasting terminal, messaging device, tablet device, personal digital assistant, etc.
[0062] Reference Figure 3 The device 300 may include one or more of the following components: processing component 302, memory 304, power component 306, multimedia component 308, audio component 310, input / output (I / O) interface 312, sensor component 314, and communication component 316.
[0063] Processing component 302 typically controls the overall operation of device 300, such as operations associated with display, telephone calls, data communication, camera operation, and recording. Processing component 302 may include one or more processors 320 to execute instructions to perform all or part of the steps of the methods described above. Furthermore, processing component 302 may include one or more modules to facilitate interaction between processing component 302 and other components. For example, processing component 302 may include a multimedia module to facilitate interaction between multimedia component 308 and processing component 302.
[0064] Memory 304 is configured to store various types of data to support the operation of device 300. Examples of such data include instructions for any application or method operating on device 300, contact data, phonebook data, messages, pictures, videos, etc. Memory 304 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0065] The power supply component 306 provides power to the various components of the device 300. The power supply component 306 may include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power to the device 300.
[0066] Multimedia component 308 includes a screen that provides an output interface between the device 300 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may sense not only the boundaries of the touch or swipe action but also the duration and pressure associated with the touch or swipe operation. In some embodiments, multimedia component 308 includes a front-facing camera and / or a rear-facing camera. When the device 300 is in an operating mode, such as a shooting mode or a video mode, the front-facing camera and / or the rear-facing camera may receive external multimedia data. Each front-facing camera and rear-facing camera may be a fixed optical lens system or have focal length and optical zoom capabilities.
[0067] Audio component 310 is configured to output and / or input audio signals. For example, audio component 310 includes a microphone (MIC) configured to receive external audio signals when device 300 is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 304 or transmitted via communication component 316. In some embodiments, audio component 310 also includes a speaker for outputting audio signals.
[0068] I / O interface 312 provides an interface between processing component 302 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, home buttons, volume buttons, start buttons, and lock buttons.
[0069] Sensor assembly 314 includes one or more sensors for providing status assessments of various aspects of device 300. For example, sensor assembly 314 may detect the on / off state of device 300, the relative positioning of components such as the display and keypad of device 300, changes in the position of device 300 or a component of device 300, the presence or absence of user contact with device 300, the orientation or acceleration / deceleration of device 300, and temperature changes of device 300. Sensor assembly 314 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 314 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, sensor assembly 314 may also include an accelerometer, a gyroscope, a magnetometer, a pressure sensor, or a temperature sensor.
[0070] Communication component 316 is configured to facilitate wired or wireless communication between device 300 and other devices. Device 300 can access wireless networks based on communication standards, such as WiFi, 2G, or 3G, or combinations thereof. In one exemplary embodiment, communication component 316 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, communication component 316 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0071] In an exemplary embodiment, the apparatus 300 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the methods described above.
[0072] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 304 including instructions, which can be executed by a processor 320 of the device 300 to perform the above-described method. For example, the non-transitory computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.
[0073] In an exemplary embodiment, a computer program product is also provided, including a computer program that implements the above-described method when executed by the processor 320 of the device 300.
[0074] Other embodiments of the invention will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of the invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of the invention are indicated by the following claims.
[0075] It should be understood that the present invention is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of the invention is limited only by the appended claims.
Claims
1. A method for diagnosing anomalies in monitoring data by integrating causal inference, characterized in that, include: Acquire multivariate time-series monitoring data collected by a sensor network deployed in the carbon dioxide storage area, and preprocess the multivariate time-series monitoring data; An unsupervised anomaly detection algorithm is used to model and score the preprocessed data. Sample points with scores higher than a preset threshold are marked as suspected anomalies, and the time interval formed by multiple consecutive suspected anomalies is determined as a suspected anomaly time segment. For each suspected abnormal time segment, all monitoring data within the suspected abnormal time segment are used as associated data, and a causal discovery algorithm is used to construct a directed acyclic causal graph between each monitoring variable based on the associated data. On the directed acyclic causal graph, variables whose data change amplitude exceeds a preset amplitude within the suspected abnormal time segment are marked as downstream abnormal variables. A causal inference algorithm is used to calculate the causal effect strength of each upstream monitoring variable on the downstream abnormal variable. Based on the causal effect strength, the initial disturbance node that leads to the propagation of the abnormality in the directed acyclic causal graph is located to obtain the abnormality source node. Based on the node type, data change characteristics, and causal propagation pattern of the anomaly source node, the root cause of the anomaly event is diagnosed; the root cause includes instrument failure, natural environmental fluctuations, and high-risk events; the high-risk events include carbon dioxide leakage or geomechanical instability; the causal propagation pattern is the causal propagation pattern reflected by the directed paths from the anomaly source node to each downstream anomaly variable in the directed acyclic causal graph.
2. The monitoring data anomaly diagnosis method based on causal inference according to claim 1, characterized in that, The preprocessing of multivariate time-series monitoring data includes: Missing values in multivariate time-series monitoring data are filled using interpolation or forward imputation, detection data exceeding the sensor's range are removed, and data timestamps from different sensors are aligned.
3. The monitoring data anomaly diagnosis method based on causal inference according to claim 1, characterized in that, The unsupervised anomaly detection algorithm employs either an isolated forest or a deep autoencoder.
4. The monitoring data anomaly diagnosis method based on causal inference according to claim 1, characterized in that, The method of diagnosing the root cause of abnormal events based on the node type of the anomaly source node, data change characteristics, and causal propagation patterns includes: If the sensor type corresponding to the abnormal source node is a single type, and the data change pattern of the node in the abnormal time segment is a step change followed by a constant value, and the out-degree of the node in the directed acyclic causal graph does not exceed the preset out-degree threshold, and the proportion of the number of affected downstream abnormal variables to the total number of sensors does not exceed the preset proportion threshold, then it is diagnosed as an instrument failure. If the data change pattern of the abnormal source node is that the absolute value of the rate of change is lower than the normal fluctuation threshold, there is no abrupt inflection point, and it can recover to the baseline range after the abnormal time segment ends, and the distribution of the abnormal variables downstream of the node in the directed acyclic causal graph is consistent with the propagation path predicted by the known geological or hydrological cycle model of the carbon dioxide sequestration area, then it is diagnosed as natural environmental fluctuation.
5. The monitoring data anomaly diagnosis method based on causal inference according to claim 1, characterized in that, The method of diagnosing the root cause of abnormal events based on the node type of the anomaly source node, data change characteristics, and causal propagation patterns includes: If the sensor corresponding to the abnormal source node belongs to any of the safety-critical variables such as wellhead pressure, caprock strain, or carbon dioxide concentration, and the data change pattern of the node is a continuous deviation from the baseline by a predetermined multiple of the normal fluctuation range, and the rate of change exceeds a predetermined multiple of the historical maximum rate of change, and there is a directed path in the directed acyclic causal graph pointing from the node to multiple downstream abnormal variables of different physical quantities, and the path conforms to the predefined causal chain template of carbon dioxide leakage or geomechanical instability, then it is diagnosed as a high-risk event. The risk level of the high-risk event is calculated by weighting three indicators: abnormal amplitude, affected area, and rate of change, and is divided into three levels: low, medium, and high according to the preset level classification rules. Among them, abnormal amplitude is defined as the ratio of the difference between the node peak value and the baseline to the baseline standard deviation, affected area is defined as the proportion of the number of downstream abnormal variables to the total number of sensors, and rate of change is defined as the maximum change of the node per unit time.
6. The monitoring data anomaly diagnosis method based on causal inference according to claim 1, characterized in that, The causal inference algorithm is counterfactual reasoning; the algorithm is used to calculate the causal effect strength of each upstream monitored variable on the downstream abnormal variable, and based on the causal effect strength, the initial disturbance node leading to the propagation of the anomaly in the directed acyclic causal graph is located to obtain the anomaly source node, including: For each upstream monitoring variable in the directed acyclic causal graph, counterfactual data is generated when the upstream monitoring variable takes different values through counterfactual reasoning, and the causal effect strength of the upstream monitoring variable on the downstream abnormal variable when it changes is calculated. By traversing all upstream monitoring variables, the upstream monitoring variable with the strongest causal effect is determined as the initial disturbance node, thus obtaining the anomaly source node.
7. A monitoring data anomaly diagnosis device integrating causal inference, characterized in that, include: The acquisition unit is used to acquire multivariate time-series monitoring data collected by a sensor network deployed in the carbon dioxide storage area, and to preprocess the multivariate time-series monitoring data. The labeling unit is used to model and score the preprocessed data using an unsupervised anomaly detection algorithm. Sample points with scores higher than a preset threshold are labeled as suspected anomalies, and the time interval formed by multiple consecutive suspected anomalies is defined as a suspected anomaly time segment. The construction unit is used to, for each suspected abnormal time segment, take all monitoring data within the suspected abnormal time segment as associated data, and use a causal discovery algorithm to construct a directed acyclic causal graph between each monitoring variable based on the associated data; The positioning unit is used to mark variables whose data change amplitude exceeds a preset amplitude within the suspected abnormal time segment as downstream abnormal variables on the directed acyclic causal graph, calculate the causal effect strength of each upstream monitoring variable on the downstream abnormal variable using a causal inference algorithm, and locate the initial disturbance node that leads to the propagation of the abnormality in the directed acyclic causal graph based on the causal effect strength to obtain the abnormality source node. The diagnostic unit is used to diagnose the root cause of anomalies based on the node type of the anomaly source node, data change characteristics, and causal propagation pattern. The root cause includes instrument failure, natural environmental fluctuations, and high-risk events. The high-risk events include carbon dioxide leaks or geomechanical instability. The causal propagation pattern is the causal propagation pattern reflected by the directed paths from the anomaly source node to each downstream anomaly variable in the directed acyclic causal graph.
8. An electronic device, characterized in that, include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the method as described in any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 6.
10. A computer program product, comprising a computer program, characterized in that, The computer program, when executed by a processor, implements the method as described in any one of claims 1 to 6.