Method for monitoring and evaluating operation condition of equipment in carbon dioxide sequestration site
By combining multimodal feature fusion and machine learning models with data from multiple sensors for preprocessing and anomaly detection, the problem of difficulty in detecting hidden equipment faults in traditional monitoring systems has been solved. This has enabled intelligent and automated monitoring and evaluation of carbon dioxide storage equipment, improving the safety and reliability of the system.
Patent Information
- Application Number
- CN202510888300.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-10-28
AI Technical Summary
Traditional monitoring systems rely on data from a single sensor, making it difficult to capture the continuous degradation trajectory of equipment performance. This makes it difficult to detect hidden equipment faults, affecting the safety and reliability of the carbon dioxide storage system.
By employing multimodal feature fusion and machine learning models, combined with data from multiple sensors for preprocessing and anomaly detection, an equipment health index is calculated to achieve intelligent assessment of equipment operation.
It enables real-time monitoring and fault prediction of carbon dioxide storage equipment, improving the safety and reliability of the system and making it suitable for large-scale geological storage projects.
Smart Images

Figure CN120850199A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of carbon dioxide sequestration, and more particularly to a method for monitoring and evaluating the operation of equipment in a carbon dioxide sequestration site. Background Technology
[0002] As a key strategy for reducing greenhouse gas emissions, carbon dioxide geological storage faces the challenge of long-term exposure of equipment in storage sites to high-pressure and highly corrosive environments, which may lead to problems such as leaks, blockages, or mechanical failures.
[0003] In the operation and maintenance of carbon dioxide geological storage systems, traditional monitoring systems rely on independent threshold alarm mechanisms from discrete sensors, making it difficult to capture the continuous degradation trajectory of equipment performance. Current monitoring technologies mainly depend on single sensor data (such as pressure or temperature threshold alarms), but single-modal data acquisition leads to a lack of spatiotemporal correlation between key parameters and a lack of comprehensive analysis capabilities for multi-source data. This makes it difficult to detect equipment performance degradation trends or potential faults in a timely manner. In addition, traditional manual inspection methods are inefficient and cannot adapt to real-time dynamic changes during equipment operation.
[0004] For example, the lack of a collaborative analysis framework between pressure and vibration sensor data makes it impossible to identify complex fault characteristics caused by early corrosion. The mismatch between manual inspection cycles and the dynamic operating status of the equipment, coupled with data acquisition frequency limited by the number of physical visits, results in a lag in identifying performance degradation trends.
[0005] For example, in the operation of the gas injection pump unit in a deep saline aquifer sequestration project, pressure sensors, temperature sensors, and vibration sensors are each set with independent alarm thresholds. When slow corrosion occurs in the plunger seal, the outlet temperature recorded by the temperature sensor rises at a rate of 0.3°C per hour, but does not reach the preset ±5°C alarm threshold. The high-frequency harmonic components detected by the vibration sensor are ignored because they are not analyzed in conjunction with the temperature change trend. Such potential failure modes continue to accumulate under the existing monitoring system until the seal fails, triggering a CO2 leakage event. At this point, a sudden pressure drop triggers an emergency shutdown, but by then, sequestration efficiency has been lost and environmental monitoring costs have increased.
[0006] If the above issues are not addressed, the hidden degradation process of equipment health will continue to deplete system safety redundancy, leading to an increase in unplanned downtime. Progressive failures such as bearing wear and valve jamming in critical rotating machinery are difficult to detect in a timely manner, significantly increasing the probability of well integrity failure. The fragmented processing of multi-source heterogeneous data reduces the generalization ability of condition assessment models, making it difficult to establish unified monitoring standards for equipment in different geological regions, ultimately affecting the feasibility of large-scale carbon sequestration project deployment.
[0007] Therefore, there is an urgent need to develop a more intelligent and automated monitoring and evaluation system to intelligently replace the traditional manual inspection process. Summary of the Invention
[0008] This invention provides a method for monitoring and evaluating the operation of equipment in a carbon dioxide storage site, thereby providing an intelligent and automated method for monitoring and evaluating carbon dioxide storage sites, and realizing the intelligent replacement of the traditional manual inspection process.
[0009] To achieve the above objectives, the present invention provides a method for monitoring and evaluating the operational status of equipment in a carbon dioxide storage site. The method includes: acquiring at least the following data during operation: carbon dioxide reservoir pressure, temperature, flow rate, operational vibration acceleration, corrosion rate, equipment sealing performance parameters, and external carbon dioxide concentration; preprocessing each of the above data and then performing multimodal feature fusion to filter out abnormal data; performing secondary anomaly detection on the filtered data based on a machine learning model, and calculating an equipment health index based on the valid data obtained after the secondary anomaly detection; and evaluating the current operational status of the equipment based on the equipment health index. The evaluation results are categorized as: normal, alert, warning, and / or fault.
[0010] As a preferred embodiment of the above technical solution, the preprocessing includes generating filler values by combining adjacent sensor data and historical data through a spatiotemporal correlation interpolation method when preprocessing the operating data.
[0011] The type of the adjacent sensor is the same as or different from the type of the sensor to which the currently preprocessed data belongs.
[0012] As a preferred embodiment of the above technical solution, multimodal feature fusion is performed to filter out abnormal data, including,
[0013] Extract time-domain and frequency-domain features from the preprocessed data. The time-domain features include mean, variance, and kurtosis, while the frequency-domain features include FFT spectral energy distribution. Specifically, in the process of extracting time-domain and frequency-domain features, multi-dimensional time-domain and frequency-domain features are extracted from data of the same type in the preprocessed data.
[0014] As a preferred embodiment of the above technical solution, when performing multimodal feature fusion on the preprocessed data, the weights of different sensor features are dynamically allocated based on the attention mechanism. After weighting the different sensor features, the data is input into the LSTM model, and the isolated forest algorithm is used to identify data points that are outside the normal operating range. Then, these data points that are outside the normal operating range are excluded.
[0015] As a preferred embodiment of the above technical solution, when performing multimodal feature fusion on the preprocessed data, the weights of different sensor features are dynamically allocated based on an attention mechanism according to different measurement modes and / or different carbon dioxide sequestration site characteristics. The weights are calculated using a scaled dot product attention model.
[0016] As a preferred embodiment of the above technical solution, a secondary anomaly detection is performed on the filtered data based on a machine learning model, including collecting data during normal operation of the device, training it with an autoencoder, and obtaining the standard deviation.
[0017] Real-time operating data of the equipment in the carbon dioxide storage site is collected, and the absolute value of the difference between the real-time operating data and the standard deviation is calculated. This absolute value is then divided by the standard deviation to obtain the data result. When this data result does not exceed 5% of the standard deviation, it is judged as normal; otherwise, it is abnormal.
[0018] As a preferred embodiment of the above technical solution, preferably, the equipment health index is calculated based on the valid data obtained after the secondary anomaly detection, including:
[0019] The equipment health index is obtained by summing the standardized scores of each data indicator of several valid data points and the weight coefficients corresponding to each valid data point.
[0020] As a preferred embodiment of the above technical solution, preferably, the range method is used to normalize each of the several valid data to obtain the standardized score.
[0021] As a preferred embodiment of the above technical solution, the weight coefficients are obtained by training a random forest model, acquiring the standardized scores corresponding to all devices under monitoring status from the historical dataset; the obtained standardized scores are labeled, with 0 indicating normal and 1 indicating fault; 200-300 decision trees are generated, and several features are randomly selected for node splitting in each tree; the total reduction in Gini impurity brought about by splitting the i-th feature in all trees is calculated, and the weight coefficients are obtained after normalization.
[0022] As a preferred embodiment of the above technical solution, preferably, when the equipment health index EHI ≥ 80, the evaluation result is normal; when 60 ≤ EHI < 80, the evaluation result is of concern; when 40 ≤ EHI < 60, the evaluation result is of warning; and when EHI < 40, the evaluation result is of fault.
[0023] This invention provides a method for monitoring and evaluating the operational status of equipment within a carbon dioxide storage site, comprising: acquiring operational data of equipment deployed within the carbon dioxide storage site; preprocessing the operational data and then performing multimodal feature fusion to filter abnormal data; performing secondary anomaly detection on the filtered data based on a machine learning model, and calculating an equipment health index based on the valid data obtained after the secondary anomaly detection; and evaluating the current operational status of the equipment based on the equipment health index; wherein the evaluation results are categorized as: normal, alert, warning, and / or fault.
[0024] The advantages of this invention are that it provides an intelligent and automated method for monitoring and evaluating carbon dioxide sequestration sites, enabling real-time monitoring, fault prediction, and intelligent maintenance of sequestration equipment. This method intelligently replaces traditional manual inspection processes, significantly improving the safety and reliability of carbon sequestration systems. It is applicable to large-scale geological sequestration projects and also solves the problem in this field where the lack of fusion analysis of multi-source data makes it difficult to detect latent equipment faults. These latent faults include, at least, slowly progressing corrosion faults. For example, cement that gradually loses its bonding ability due to the corrosive effect of carbon dioxide may create tiny leakage channels around the wellbore, leading to sequestration failures. Such corrosion and pressure changes are difficult to detect through routine inspections in the early stages, but can be detected in a timely manner through the equipment health index by using corresponding sensors / sensor combinations at appropriate locations in accordance with the technical solution of this invention, continuously monitoring changes in trace data / data combinations. Attached Figure Description
[0025] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0026] Figure 1 A flowchart illustrating a method for monitoring and evaluating the operational status of equipment within a carbon dioxide storage site, provided by this invention. Figure 1 .
[0027] Figure 2 A flowchart illustrating a method for monitoring and evaluating the operational status of equipment within a carbon dioxide storage site, provided by this invention. Figure 2 . Detailed Implementation
[0028] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0029] In response, this application proposes a method for monitoring and evaluating the operational status of equipment within a carbon dioxide sequestration site. Figure 1 This is a schematic diagram of a process provided for an embodiment of the present invention, such as... Figure 1 As shown, the following steps are included:
[0030] Step 101: Obtain the operating data of the equipment deployed in the carbon dioxide storage site.
[0031] Among them, acquiring operational data refers to collecting operational parameters generated by various sensors installed on equipment within the carbon dioxide storage site. Specifically, pressure sensors, temperature sensors, and flow sensors can be used to cover multi-dimensional information on the equipment's operational status.
[0032] The operational data includes pressure, temperature, flow rate, vibration acceleration, corrosion rate, equipment sealing performance parameters, and carbon dioxide leakage concentration.
[0033] Introducing equipment sealing performance parameters into data monitoring can solve the problem of micro-leakage under high pressure and corrosion environments. This is because carbon dioxide has strong permeability and corrosiveness in a supercritical state, and traditional pressure monitoring cannot accurately and timely detect the risk of seal failure.
[0034] Introducing carbon dioxide leakage concentration monitoring into the data collection process can solve the problem that conventional pressure sensors cannot accurately detect low-concentration carbon dioxide leaks, thus preventing the safety of the storage site from being compromised by accumulated leaks. This is because low-concentration carbon dioxide is not only colorless and odorless, but also poses a suffocation risk to maintenance personnel. Fiber optic distributed sensors are used to detect the carbon dioxide leakage concentration at potential leak points in the wellbore within the storage site. When the leakage concentration exceeds a threshold, an alarm is triggered, and a ventilation system is activated to replace the air inside the wellbore, preventing the accumulation of carbon dioxide.
[0035] Step 102: After preprocessing the running data, perform multimodal feature fusion to filter out abnormal data.
[0036] Preprocessing refers to imputing missing values and filtering noise in the raw operating data. Specifically, spatiotemporal correlation interpolation can be used to combine data from adjacent sensors and historical data to generate imputed values, thus avoiding misjudgments caused by anomalies in a single data source. Multimodal feature fusion refers to integrating time-domain and frequency-domain features extracted from different sensor data. Specifically, an attention mechanism can be used to dynamically assign weights to different sensor features, and the weighted input is then fed into an LSTM model to enhance the reliability of anomaly detection.
[0037] Traditional data interpolation methods may not be able to effectively utilize the spatiotemporal correlation between sensors, resulting in insufficient accuracy of the interpolated values. This is especially true when different sensor types are used, as a single interpolation method is difficult to adapt to data loss problems under complex working conditions.
[0038] To address this, this application further proposes a method for generating imputation values during the preprocessing of operational data. This method combines data from adjacent sensors with historical data using a spatiotemporal correlation interpolation approach. The types of adjacent sensors may be the same or different from the data type of the sensor to which the currently preprocessed data belongs. When adjacent sensors are of the same type as the current sensor, linear interpolation can be directly performed using the spatial data change trend. When the types are different, a cross-modal data mapping model needs to be established to normalize the values of different types of sensors, such as pressure and flow, into correlation coefficients. Furthermore, the time window for recalling historical data can be set to 5-10 minutes, using a sliding window mechanism to extract periodic temperature fluctuation characteristics. For example, when pressure sensor data is missing, real-time pressure gradient data collected from three adjacent temperature sensors can be combined with the average pressure change rate over the past 8 minutes at that location to form a composite interpolation value.
[0039] For example, in scenarios where temperature sensors experience data gaps, the real-time pressure fluctuation curves of pressure sensors within a 2-meter radius are first extracted, while historical temperature data from the past 6 minutes for those sensors is retrieved. By weighted fusion of the second derivative of the current pressure fluctuation curve and the linear fitting slope of the historical temperature data, imputation values are generated. During this process, the data from both pressure and temperature sensors are normalized and mapped to the 0-1 range. The weighting coefficients can be dynamically adjusted based on the sensor spacing; for example, the weight of the pressure data decreases by 15% for every 0.5-meter increase in spacing. This cross-sensor data complementarity mechanism allows for the reconstruction of missing parameters from spatially correlated heterogeneous data such as vibration and flow rates, even in the highly corrosive environment of carbon dioxide storage sites, keeping the data interpolation error rate within a low level, even if some sensors fail.
[0040] During data preprocessing, mutual information or maximum correlation minimum redundancy algorithms are used to filter the extracted time-domain and frequency-domain features, and a feature importance evaluation model is established to eliminate redundant and irrelevant features.
[0041] Step 103: Perform secondary anomaly detection on the filtered data based on the machine learning model, and calculate the equipment health index based on the valid data obtained after the secondary anomaly detection.
[0042] Secondary anomaly detection refers to re-identifying anomalies in the initially screened data. Specifically, this involves training an autoencoder to obtain the standard deviation and calculating the deviation of real-time data from the standard deviation to improve the accuracy of identifying complex data patterns. This step further enhances the accuracy of anomaly identification through the model's ability to learn complex data patterns. Specifically, an autoencoder is used to train normal operating data to obtain the standard deviation, and then the real-time collected operating data is compared with the standard deviation to determine whether the data is abnormal. This method can effectively identify abnormal patterns, avoiding the limitations of traditional fixed threshold methods.
[0043] Step 104: Assess the current operating status of the equipment based on the equipment health index; the assessment results are: normal, alert, warning, and / or fault.
[0044] The equipment health index is a quantitative indicator obtained by weighting effective data through standardized scores and weighted coefficients, transforming multidimensional data into quantifiable evaluation metrics. Specifically, it utilizes range normalization combined with weighted coefficients obtained from training a random forest model to objectively assess equipment status. The evaluation result refers to the levels assigned based on the equipment health index, specifically categorized into normal, watchful, warning, and fault states using preset thresholds, dynamically guiding equipment maintenance decisions. This addresses the issues of high subjectivity and low efficiency associated with traditional manual judgment.
[0045] The core innovation of this application lies in integrating multi-source sensor data and adopting multimodal feature fusion and secondary anomaly detection, combined with a dynamic evaluation system for equipment health index, to replace traditional single-sensor threshold alarms and manual inspections, thus solving the problems of insufficient comprehensive analysis capabilities and poor real-time performance.
[0046] The working process and principle of this application are as follows: This method integrates multi-source data with intelligent analysis technology to construct a complete equipment operation monitoring and evaluation system. First, acquiring operational data is fundamental. By deploying multiple types of sensors to cover multi-dimensional information about the equipment's operating status, the problem of incomplete coverage by traditional single-sensor data is solved. Second, the acquired operational data is preprocessed, followed by multi-modal feature fusion. The complementarity of data from different sensors enhances the reliability of anomaly detection. In this process, feature extraction and fusion techniques are used to extract time-domain and frequency-domain features from the preprocessed data. Based on an attention mechanism, the weights of different sensor features are dynamically allocated, achieving effective integration of multi-source heterogeneous data.
[0047] As a preferred embodiment, the solution of this application is specifically implemented as follows:
[0048] Within the carbon dioxide sequestration site, various types of sensors are deployed, including at least pressure sensors, temperature sensors, and vibration sensors, to collect operational data from the equipment. These sensors are installed at key equipment locations such as injection pumps, valves, and pipelines to comprehensively monitor the equipment's operational status.
[0049] After data acquisition, preprocessing is performed first. Preprocessing steps include data cleaning, noise reduction, and standardization. For missing data, a spatiotemporal correlation interpolation method is used, combining data from adjacent sensors with historical data to generate imputation values. For example, when pressure sensor P1 in a carbon dioxide storage site has missing data, real-time data from adjacent temperature sensor T1 and flow sensor F1 can be used, combined with historical data from sensor P1 for interpolation. Specifically, a spatial correlation model between P1 and T1, F1 can be established, and combined with the historical time-series characteristics of P1, to generate an interpolation value that comprehensively considers spatiotemporal factors. This method is applicable to situations where the data types of adjacent sensors are the same or different from the data being preprocessed. This application can effectively improve the accuracy and adaptability of data interpolation. By fully utilizing the spatial correlation between sensors and the time-series characteristics of historical data, this method can better handle data missing problems under complex operating conditions. Furthermore, by allowing data from different types of sensors to participate in interpolation, this approach overcomes the limitations of traditional interpolation methods, enhances the applicability of interpolation algorithms in heterogeneous sensor networks, and addresses the issue that due to the differences in feature representation capabilities among different types of sensor data, single-dimensional feature extraction cannot fully reflect changes in equipment operating status, leading to the risk of missed detections or misjudgments in anomaly data screening. This cross-modal data complementarity approach can significantly improve the reliability of missing value recovery, thereby providing a more accurate data foundation for subsequent anomaly detection and equipment health status assessment.
[0050] Next, multimodal feature fusion is performed on the preprocessed data to filter out outliers, including extracting time-domain features (such as mean, variance, and kurtosis) and frequency-domain features (such as FFT spectral energy distribution). For data of the same type, multi-dimensional time-domain and frequency-domain features are extracted.
[0051] In extracting time-domain features, the mean is obtained by calculating the arithmetic mean of the data sequence, the variance is obtained by calculating the squared mean of the data deviations from the mean, and the kurtosis is obtained by the ratio of the fourth-order central moment to the fourth power of the standard deviation. In the frequency-domain feature extraction process, the FFT spectral energy distribution is achieved by converting the time-domain signal to the frequency domain using a Fast Fourier Transform (FFT) and then statistically analyzing the energy proportions of different frequency bands.
[0052] Specifically, multi-dimensional time-domain and frequency-domain feature extraction is performed on data of the same type in the preprocessed data. For example, for temperature data, its mean, variance, kurtosis, and other statistical measures are calculated as time-domain features, while an FFT transformation is performed to extract the spectral energy distribution as frequency-domain features. Similarly, time-domain statistical features and frequency-domain energy features are extracted for pressure data. This multi-dimensional feature extraction method is applied to all types of sensor data, ensuring that each data type receives comprehensive feature representation.
[0053] When performing dual-dimensional feature extraction in both the time and frequency domains for the same type of data, the following implementation methods exist: For temperature data, the time domain features may include a mean of 25℃, a variance of 0.5, and a kurtosis of 3.2, while the frequency domain features may include 65% energy in the 0-10Hz frequency band; for pressure data, the time domain features may include a mean of 5MPa, a variance of 0.2, and a kurtosis of 2.8, while the frequency domain features may include 72% energy in the 10-20Hz frequency band. After generating imputed values using the spatiotemporal correlation interpolation method during the preprocessing stage, the imputed values and the original data participate together in feature calculation to ensure data integrity and feature continuity.
[0054] Specifically, the abnormal data screening process is divided into three levels: First, the mean of the time domain features is used to capture the overall data offset, the variance is used to identify fluctuation anomalies, and the kurtosis is used to detect spike pulses; second, the frequency domain features are used to analyze the abnormal energy distribution of periodic signals; and finally, the time domain and frequency domain features of the same data source are jointly judged.
[0055] For example, when the mean temperature data remains stable but the kurtosis suddenly increases to a critical threshold, it is identified as a transient anomaly; when the energy proportion of pressure data in a preset frequency band decreases to a critical threshold, it can be identified as a periodic anomaly. The number of feature dimensions for different types of data is dynamically adjusted according to their physical characteristics. For example, vibration data can have time-domain waveform factor indicators and frequency-domain harmonic distortion indicators added to form a five-dimensional feature vector. Through multi-dimensional feature fusion, the characterization error of changes in equipment operating status is controlled within a low range, and the anomaly detection recall rate can be significantly improved.
[0056] Through the above technical solution, this application achieves a comprehensive characterization of the equipment's operating status. Time-domain features reflect the overall level, fluctuation amplitude, and extreme value distribution of the data, while frequency-domain features reveal periodic or sudden signal characteristics. This multi-dimensional feature fusion method enhances the accuracy and robustness of anomaly data screening, effectively reducing the risk of missed detections and false positives. Therefore, this invention provides more reliable and comprehensive feature inputs for subsequent anomaly detection, improving the overall accuracy of monitoring and evaluation.
[0057] While the aforementioned approach proposes multimodal feature fusion to screen abnormal data and improve monitoring accuracy, the traditional fixed-weight feature fusion method struggles to adapt to dynamically changing operating conditions due to the significant differences in the correlation between data features collected by different sensors over time. This results in the weakening of key information or amplification of noise interference during abnormal data screening, thereby affecting the reliability of subsequent secondary anomaly detection.
[0058] Therefore, this invention further proposes a technical solution for dynamically allocating weights of different sensor features based on an attention mechanism when performing multimodal feature fusion on preprocessed data. The weights are calculated using a scaled dot product attention model to adapt to different measurement modes and storage site characteristics.
[0059] After weighting, features from different sensors are weighted and the fused features are input into an LSTM model. The LSTM model can capture the temporal dependencies of the data, which helps identify abnormal patterns. Then, the Isolation Forest algorithm is used to identify data points outside the normal operating range, and these data points are then excluded, completing the initial abnormal data screening.
[0060] Specifically, dynamic weight allocation is achieved through a scaled dot product attention model. This model calculates the attention scores of different sensor features as weight coefficients, with the weight coefficients controlled between 0.1 and 0.9 to avoid complete suppression or dominance of a single sensor feature. The weighted multimodal features are input into an LSTM model as a sequence with a time step of 20. The hidden layer dimension of this model is set to 64 to balance computational efficiency and feature representation capability. The Isolation Forest algorithm constructs isolation trees and calculates anomaly scores based on path length. When the anomaly score exceeds a threshold of 1.5, it is identified as an anomalous data point.
[0061] Under high-pressure sudden change conditions, the attention weight of the pressure sensor is automatically increased to enhance the contribution of pressure fluctuation characteristics, while the weight of the temperature sensor is decreased to suppress ambient temperature interference. When the weighted feature sequence is processed by the LSTM model, its gating mechanism can retain equipment state evolution information over a certain time step, such as capturing gradual anomaly patterns in pressure sensor readings that exceed the baseline value over several consecutive time steps. By calculating the distance difference between data points and the normal operating condition distribution using the isolated forest algorithm, short-term sudden anomalies that deviate from the normal data distribution by more than the standard deviation threshold, as well as long-term gradual anomalies that continuously deviate from the standard deviation, can be identified simultaneously. After the time-series features output by the LSTM are subjected to anomaly detection, these data points are identified as exceeding the normal operating condition range and excluded, effectively reducing noise interference in subsequent health index calculations.
[0062] Through the above technical solutions, this application achieves dynamic adaptability in multimodal data fusion. The attention mechanism can adaptively adjust weights based on the importance differences of real-time data features, effectively improving the expressive power of key information. The LSTM model captures the dynamic evolution of equipment operating status over time, enhancing the ability to model long-term trends. The isolated forest algorithm accurately identifies short-term sudden anomalies and long-term gradual anomalies, improving the accuracy and reliability of anomaly detection. Therefore, this solution adapts to complex and changing operating environments and achieves hierarchical filtering of abnormal data, providing a high-quality data foundation for subsequent equipment health status assessment.
[0063] This application further proposes to dynamically allocate the weights of different sensor features based on an attention mechanism according to different measurement modes and / or different carbon dioxide storage site characteristics. The weights are calculated using a scaled dot product attention model to address the problem that the dynamic weight allocation lacks adaptability to specific scenarios due to the different influences of different measurement modes or different storage site characteristics on sensor data characteristics, thus reducing the accuracy of abnormal data screening.
[0064] The measurement modes can be categorized into pressure measurement, temperature measurement, and vibration measurement modes. Site feature parameters cover geological structure porosity and reservoir pressure values. The scaled dot product attention model generates the original attention score by performing a dot product operation between the query vector and the key vector, and then scales it according to the dimension of the sensor features. Geological structure parameters are obtained from geological exploration reports, and reservoir pressure values are monitored in real time by pressure sensors. Different measurement modes correspond to different feature encoding methods. For example, the data window length is set to X seconds in the pressure measurement mode and Y seconds in the temperature measurement mode. During the attention weight calculation process, the geological structure parameters are encoded as conditional vectors and concatenated with the sensor features before being input into the attention layer.
[0065] Specifically, in the pressure measurement mode, the sensor feature vector is mapped to a query vector and a key vector of several dimensions. After obtaining the relevance score through dot product operation, the real-time reservoir pressure monitoring value is concatenated as a condition vector so that the attention score can reflect the adjustment requirements of the feature weight under the current pressure condition. When the porosity of the geological structure is lower than the threshold, the feature weight of the vibration sensor is increased to enhance the identification of mechanical loosening anomalies. A scaling factor is set to control the original attention score in the range of [-1,1] to avoid the gradient vanishing problem. Finally, the generated dynamic weight is weighted with the sensor features and then input into the LSTM model for time series modeling to improve the accuracy of anomaly data screening.
[0066] Through the above technical solution, this application can dynamically adjust the sensor feature weights according to different measurement modes and storage site characteristics, improving the targeting and adaptability of multimodal feature fusion. The introduction of the scaled dot product attention model effectively captures the correlation between sensor features, while controlling gradient stability through scaling operations, making it suitable for weight calculation of high-dimensional sensor data. This method significantly improves the accuracy of anomaly data screening and enhances the monitoring system's adaptability to complex working conditions.
[0067] In the subsequent secondary anomaly detection phase, data from normal equipment operation is first collected and trained using an autoencoder to obtain the standard deviation. Then, real-time operating data is collected, and the absolute value of the difference between the real-time operating data and the standard deviation is calculated, divided by the standard deviation to obtain the data result. If this data result does not exceed 5% of the standard deviation, it is considered normal; otherwise, it is considered abnormal.
[0068] The training data for the autoencoder can be selected from historical data of the device running continuously for more than 30 days, and the number of hidden layer nodes and input layer nodes can be set. During training, the mean squared error loss function is used for backpropagation optimization, and the number of iterations is set. The standard deviation can be calculated based on the statistical distribution of the error reconstructed from the training set. The difference between the real-time data and the benchmark value can be calculated using a sliding window mechanism, and the sampling frequency is kept consistent with that of the training data.
[0069] When real-time data fluctuations exceed a 5% threshold, an alarm signal is triggered and the health index recalculation mechanism is initiated. By combining dynamic benchmarks with statistical criteria, fluctuations caused by normal equipment wear and tear and sudden anomalies are effectively distinguished. For example, if abnormal vibration signals before the sealing ring fails are identified, the health index is recalculated. Compared with static threshold detection methods, the accuracy is improved, and it can better adapt to the dynamic changes in equipment operating status, reducing the probability of false alarms and missed alarms.
[0070] For example, in a carbon dioxide storage facility, data is collected during normal equipment operation, such as pressure, temperature, and flow rate. This data is input into an autoencoder for training, yielding the standard deviation. The autoencoder can employ a multi-layer neural network structure. During training, the learning rate, batch size, and number of iterations are set. After training, for each input feature, the standard deviation of its reconstruction error is calculated and used as the baseline value for that feature's standard deviation.
[0071] During actual equipment operation, real-time operational data is collected. For each feature, the absolute value of the difference between the real-time operational data and its corresponding standard deviation is calculated. This absolute value is then divided by the standard deviation to obtain a normalized deviation value. This normalized deviation value is compared to a preset threshold of 5%. If the normalized deviation value does not exceed 5%, the data point is considered normal; otherwise, it is considered abnormal. This judgment process is performed simultaneously on all features. Ultimately, if all features are judged to be normal, the equipment is considered to be in a normal state; if any feature is judged to be abnormal, an anomaly alarm is triggered.
[0072] Furthermore, based on the valid data after secondary anomaly detection, the equipment health index is calculated. The calculation method involves multiplying the standardized score of each valid data indicator by its corresponding weight coefficient and then summing the results. The standardized score is obtained by normalizing each valid data point using the range method. The weight coefficients are obtained through training a random forest model.
[0073] The standardized scores are normalized for each data indicator using the range method to eliminate dimensional differences. Weight coefficients are obtained through training a random forest model. Standardized scores labeled as normal or faulty in historical datasets are input into a random forest model consisting of 200-300 decision trees. Each tree randomly selects some features for node splitting, and finally, the proportion of each feature in the total reduction of Gini impurity is calculated and normalized. For example, the weight coefficient for temperature sensor data might be calculated as 0.35, while the weight coefficient for pressure sensor data might be 0.45, reflecting the different impacts of different parameters on equipment status. In the weighted summation process, the standardized score of each indicator is multiplied by its corresponding weight and then summed to form a health index in the range of 0-100.
[0074] Specifically, when calculating the equipment health index, the temperature, pressure, and vibration indicators in the valid data are first standardized using the range method, converting the original values into standardized scores in the 0-1 range. Then, weight coefficients generated by training a random forest model are assigned to each indicator; for example, the temperature indicator has a weight of 0.2, the pressure indicator has a weight of 0.5, and the vibration indicator has a weight of 0.3. The weighted summation process is implemented as: standardized temperature score × 0.2 + standardized pressure score × 0.5 + standardized vibration score × 0.3. The final result is adjusted to an index value of 0-100 through linear mapping. This calculation method highlights the diagnostic value of key indicators. For example, when the pressure indicator shows abnormal fluctuations, its higher weight coefficient will significantly decrease the health index, thus accurately reflecting the actual operating status of the equipment. Compared to the simple arithmetic mean method, this technical solution improves the sensitivity of the health index to equipment performance degradation while reducing the false alarm rate.
[0075] Through the above technical solutions, this application achieves precise quantification of the differences in contribution of different data indicators, avoiding the calculation bias of the health index caused by simple summation. Standardized scoring eliminates the incomparability between data of different dimensions, enabling all indicators to participate in the calculation on a unified scale. Introducing weighting coefficients quantifies the degree of influence of different data indicators on the equipment's health status, avoiding the weakening of important characteristics caused by simple arithmetic averaging. The weighted summation method retains the contribution characteristics of the original information of each indicator while highlighting the diagnostic value of key indicators through weight allocation. Therefore, the distinguishability and reliability of the health index are improved, providing data support for the accurate grading of equipment assessment results.
[0076] Finally, the current operating status of the equipment is assessed based on the calculated equipment health index. The assessment results are divided into four levels: when the health index is greater than or equal to 80, the assessment result is normal; when the health index is between 60 and 80, the assessment result is at risk; when the health index is between 40 and 60, the assessment result is a warning; and when the health index is less than 40, the assessment result is a malfunction.
[0077] Through the above-described scheme, this application achieves intelligent and automated monitoring and evaluation of equipment operation within carbon dioxide sequestration sites. This method improves the accuracy and timeliness of anomaly detection through comprehensive analysis of multi-source data. Multimodal feature fusion technology fully utilizes the complementarity of data from different sensors to capture complex fault characteristics that are difficult to identify with a single sensor. Secondary anomaly detection based on machine learning further enhances the accuracy of anomaly identification, enabling the identification of performance degradation trends for each piece of equipment. The calculation and grading evaluation mechanism of the health index enables quantitative management of equipment status, providing an objective basis for maintenance decisions. The adaptability of this method allows it to adapt to the equipment monitoring needs of different geological structures, providing technical support for the large-scale deployment of carbon sequestration projects.
[0078] Through the above technical solutions, this application also effectively solves the problem of traditional weight allocation methods relying on subjective experience. It automatically quantifies the impact of each monitoring feature on the equipment's health status using a random forest model. Based on a statistical method for reducing Gini impurity, it objectively reflects the actual contribution of different sensor data to equipment fault prediction, making the health index calculation results more closely reflect the actual operating state of the equipment. By integrating the feature importance assessment of multiple decision trees, it avoids the evaluation bias that may exist in a single model, improves the stability and reliability of the weight coefficients, and ultimately ensures that the equipment health assessment results have higher accuracy and timely warning capabilities.
[0079] The technical solution of this application will now be described in conjunction with specific embodiments, such as... Figure 2 As shown:
[0080] Step 201: Obtain the device's operating data.
[0081] Specifically, multiple types of sensors are deployed at key nodes of the sealing equipment, including at least pressure sensors for monitoring injection pressure and pipeline pressure fluctuations, temperature sensors for monitoring carbon dioxide gas temperature and equipment surface temperature, flow meters for monitoring CO2 flow, acoustic emission sensors for capturing pipeline crack signals, vibration acceleration sensors for monitoring abnormal pump vibration, and electrochemical sensors for detecting corrosion rates.
[0082] The data from the above sensors are collected in real time at a frequency of 1Hz-10Hz and transmitted to the cloud server via a low-power wide area network (LoRa / NB-IoT).
[0083] Furthermore, operational data includes pressure, temperature, flow rate, vibration acceleration, corrosion rate, equipment sealing performance parameters, and carbon dioxide leakage concentration. This data is crucial for real-time monitoring of equipment operating status and environmental conditions, providing detailed information about equipment health and potential risks.
[0084] Because carbon dioxide is highly permeable and corrosive in its supercritical state, traditional pressure monitoring cannot accurately and promptly detect the risk of seal failure. However, by monitoring sealing performance parameters, potential leaks can be detected earlier, allowing for preventative measures to be taken to ensure stable equipment operation and environmental safety.
[0085] By incorporating carbon dioxide leakage concentration into data monitoring, the system addresses the challenge of detecting carbon dioxide using conventional methods due to its colorless and odorless chemical properties, thus preventing the asphyxiation hazard to maintenance personnel. Utilizing fiber optic distributed sensors, the system can detect carbon dioxide leakage concentrations at potential leak points within the wellbore at the storage site. When the leakage concentration exceeds a preset safety threshold, the system automatically triggers an alarm and activates a ventilation system to replace the air within the wellbore, preventing carbon dioxide accumulation and ensuring the safety of personnel and the stability of the storage site.
[0086] Step 202: Preprocess the running data.
[0087] The initial basic processing stage of data preprocessing: After data encryption, it is transmitted to the edge computing node for outlier removal, noise filtering (using wavelet transform algorithm) and data normalization. For missing data, spatiotemporal correlation interpolation method is used to generate filler values by combining adjacent sensor data and historical data.
[0088] Feature selection stage of data preprocessing: After completing the aforementioned basic processing, the feature processing stage begins. This stage uses mutual information or maximum correlation minimum redundancy algorithms to extract features from the dataset obtained in the basic processing stage. The extracted time-domain and frequency-domain features are then selected, and a feature importance evaluation model is established. The aim is to eliminate redundant and irrelevant features, retaining only the most valuable features for subsequent analysis or modeling, thereby optimizing the feature set and improving the model's performance and efficiency.
[0089] Step 203: Perform multimodal fusion on the preprocessed data results.
[0090] The time-domain and frequency-domain features of the data are extracted. The time-domain features include mean, variance, and kurtosis, while the frequency-domain features include the FFT spectral energy distribution. In this embodiment, more than one pressure sensor capable of sensing pressure values at different locations is used; therefore, the average pressure value is adopted. If only one pressure sensor is used, the average pressure value is directly adopted. Thus, some data may have four dimensions, while others may have five, with the dimensions depending on the number of measuring devices. Subsequently, different sensor features are dynamically weighted using an attention mechanism and input into an LSTM (Long Short-Term Memory) network model to predict the device's state trend over the next 12 hours. Data points outside the normal operating range are identified using the Isolation Forest algorithm and then excluded.
[0091] When performing multimodal feature fusion on the data, weights for different sensor features are dynamically assigned based on an attention mechanism. The weights are calculated using a scaled dot product attention model. Different measurement modes are selected based on the site, and different weights are assigned according to the actual measurement mode.
[0092] Step 204: Use an autoencoder model to perform anomaly detection on the data after multimodal fusion screening.
[0093] This includes acquiring data during normal equipment operation, training this data on an autoencoder, and calculating the standard deviation of this data; acquiring real-time data of the equipment operating in the storage area, and dividing the absolute value of the difference between the real-time data and the standard deviation by the standard deviation. If the difference does not exceed 5%, it is considered normal; if it exceeds 5%, it is considered abnormal.
[0094] When real-time data fluctuations exceed the 5% threshold, an alarm signal can be directly triggered and the health index calculation mechanism can be started. By combining dynamic benchmarks with statistical data, fluctuations caused by normal wear and tear of equipment and sudden anomalies can be effectively distinguished. For example, if an abnormal vibration signal before the sealing ring fails is identified, the health index can be directly calculated (without secondary anomaly detection). Compared with the static threshold detection method, the accuracy is improved, and it can better adapt to the dynamic changes in the operating status of the equipment, reducing the probability of false alarms and missed alarms.
[0095] Step 205: Perform secondary anomaly detection based on the machine learning model, remove abnormal data, and calculate the equipment health index.
[0096] The formula for calculating the equipment health index is:
[0097]
[0098] Among them, S i W is the standardized score of the i-th indicator. i is the weight coefficient corresponding to the i-th indicator.
[0099] Specifically, S i The steps to obtain it are as follows:
[0100] The original equipment data of the i-th item is normalized using the range method to obtain the processed data;
[0101]
[0102] Where k is the slope factor, 1≤k≤1.5, and is usually chosen to be 1; S' i For data processing; c is the offset, 1≤c≤2, e is the Euler number.
[0103] W i The data was obtained by training a random forest model, and the steps are as follows:
[0104] Obtain the standardized score S for all devices in the monitoring state from the historical dataset. i ;
[0105] The obtained standardized score S i The labels are used, with 0 indicating normal and 1 indicating a fault.
[0106] Generate 200-300 decision trees, and randomly select n features for node splitting in each tree;
[0107] The total reduction in Gini impurity resulting from the splitting of the i-th feature in all trees is calculated, and the weight coefficient W is obtained after normalization. i .
[0108] Step 206: Assess the current operating status of the equipment based on the equipment health index. The assessment result is: Normal, Attention, Warning, or Fault.
[0109] The health assessment results are ultimately categorized into four types and used to guide operational and maintenance decisions.
[0110] Normal (EHI≥80): The assessment result is normal, the equipment is in optimal operating condition, and no intervention is required;
[0111] Attention (60≤EHI<80): The assessment result is at attention. The equipment has slight performance degradation. It is recommended to increase the monitoring frequency (e.g., from once / hour to once / 10 minutes).
[0112] Warning (40≤EHI<60): The assessment result is a warning, which clarifies the fault risk, triggers the automatic diagnostic program and generates a pre-maintenance work order;
[0113] Fault (EHI<40): The assessment result is a fault. Immediately implement shutdown protection and initiate emergency handling measures (such as closing the upstream valve).
[0114] To more intuitively observe the health assessment results, a heat map of the EHI of all equipment in the sealed site is displayed in the control center, with flashing red markings indicating faulty equipment with an EHI < 40.
[0115] The technical solution of the present invention will be described in conjunction with specific data. This embodiment provides a set of data on a high-pressure plunger injection pump and obtains the health assessment results of this device. For ease of explanation, the data collected by each sensing / sensing device provided in this embodiment has a dimension of 1 (calculated based on only 1 sensor).
[0116] The data that needs to be detected for a high-pressure plunger injection pump includes: vibration acceleration sensor to collect the vibration acceleration of the pump body bearing housing, acoustic emission sensor to collect the vibration value of the flange at the pump body outlet, pressure sensor to collect the pressure value of the pump body outlet pipeline, temperature sensor to collect the temperature of the bearings used in the pump body; electrochemical corrosion sensor to collect the corrosion rate of the inner wall of the pump body cavity, and flow meter to collect the carbon dioxide flow rate at the pump body outlet.
[0117] The data collected by the sensors on site are as follows:
[0118] Vibration acceleration sensor <![CDATA[8.2m / s 2 ]]> Acoustic emission sensor 1200 times / second pressure sensor 0.3Mpa Temperature sensor 82℃ Electrochemical corrosion sensor 0.12mm / year Flow meter 300L / min
[0119] The data was processed using the method described above, and the final health assessment results of the high-pressure plunger injection pump are as follows:
[0120] EHI = 38.2 (vibration) * 0.4 + 45.1 (temperature) * 0.2 + 72.3 (corrosion) * 0.15 = 52.1 (warning level).
[0121] If an early warning is issued, the pump speed needs to be reduced to below 85% of the rated value to extend its lifespan. The result should be sent to maintenance personnel and recorded.
[0122] It should be noted that in actual implementation, vibration acceleration sensors may include several vibration acceleration sensors installed at different locations on the equipment throughout the carbon dioxide storage site, as well as pressure sensors, temperature sensors, etc., installed at different locations. Health assessment is not based solely on the value of one sensor.
[0123] The technical solution of the present invention will be further explained with reference to specific data. This embodiment provides a set of data about a carbon dioxide compressor and obtains the health assessment results of this device. For ease of explanation, the data collected by each sensing / sensing device provided in this embodiment has a dimension of 1 (calculated based on only 1 sensor).
[0124] The data that needs to be monitored for a carbon dioxide compressor includes: vibration acceleration sensor to collect the vibration acceleration of the main shaft bearing, temperature sensor to collect the maximum temperature value of the cylinder, pressure differential sensor to collect the pressure difference between the upstream and downstream of the sealing gas, and laser absorption spectrometer to collect the carbon dioxide concentration on the compressor casing.
[0125] The data collected by the sensors on site are as follows:
[0126]
[0127]
[0128] The data was processed using the method described above, and the final health assessment results for the carbon dioxide compressor were obtained as follows:
[0129] EHI = 81 (vibration) * 0.25 + 72 (temperature) * 0.1 + 76 (sealing) * 0.35 + 78 (concentration) * 0.3 = 77.45 (level of concern).
[0130] Under the monitoring of the results, the dry gas seal of the carbon dioxide compressor needs to be monitored at regular intervals, and the results should be sent to the maintenance personnel and recorded.
[0131] It should be noted that in actual implementation, vibration acceleration sensors may include several vibration acceleration sensors installed at different locations on the equipment throughout the carbon dioxide storage site, as well as pressure sensors, temperature sensors, etc., installed at different locations. Health assessment is not based solely on the value of one sensor.
[0132] The beneficial effects of this invention are as follows:
[0133] I. The method provided by this invention involves multiple different types of sensors. When a single sensor fails, the system can maintain an evaluation accuracy of over 80% through other modal data.
[0134] Second, based on the health assessment results of operating equipment, the proportion of planned maintenance has increased from 40% to 90%, reducing the number of emergency repairs and the incidence of major leakage accidents has dropped to less than 0.1 times / year (the industry average is 2 times / year); the maintenance response time has been shortened from 4-6 hours of traditional manual analysis to less than 10 minutes, and the equipment health history data has been fully recorded to meet the audit requirements of the ISO27916 carbon sequestration safety standard.
[0135] Third, the tolerance for temporary over-limits caused by fluctuations in operating conditions is improved (such as a brief rise in temperature during sudden load changes without triggering false alarms), and the detection sensitivity for slowly developing corrosion-related faults is improved (the score decay mechanism amplifies minute trends); the data-driven weight allocation avoids subjective human bias, and the importance of features directly reflects the causal relationship between indicators and faults.
[0136] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for monitoring and evaluating the operational status of equipment within a carbon dioxide sequestration site, characterized in that, The following steps are involved: At least the following data should be obtained during operation of the equipment deployed in the carbon dioxide storage site: reservoir pressure, temperature, flow rate, operating vibration acceleration, corrosion rate, equipment sealing performance parameters, and external carbon dioxide concentration. After preprocessing each of the above data separately, multimodal feature fusion is performed to filter out outlier data; A secondary anomaly detection is performed on the filtered data based on a machine learning model, and the equipment health index is calculated based on the valid data obtained after the secondary anomaly detection. The current operating status of the equipment is assessed based on the equipment health index; the assessment results are categorized as: normal, alert, warning, and / or fault.
2. The method according to claim 1, characterized in that, The preprocessing includes generating filler values by combining adjacent sensor data and historical data using a spatiotemporal correlation interpolation method when preprocessing the operational data. The type of the adjacent sensor is the same as or different from the type of the sensor to which the currently preprocessed data belongs.
3. The method according to claim 2, characterized in that, The process of multimodal feature fusion and abnormal data filtering includes, Extract time-domain and frequency-domain features from the preprocessed data. The time-domain features include mean, variance, and kurtosis, while the frequency-domain features include FFT spectral energy distribution. Specifically, in the process of extracting time-domain and frequency-domain features, multi-dimensional time-domain and frequency-domain features are extracted from data of the same type in the preprocessed data.
4. The method according to claim 1, characterized in that, When performing multimodal feature fusion on the preprocessed data, the weights of different sensor features are dynamically allocated based on the attention mechanism. After weighting the different sensor features, the data is input into the LSTM model. The isolated forest algorithm is used to identify data points that are outside the normal operating range, and then these data points outside the normal operating range are excluded.
5. The method according to claim 4, characterized in that, When performing multimodal feature fusion on the preprocessed data, the weights of different sensor features are dynamically assigned based on the attention mechanism according to different measurement modes and / or different carbon dioxide sequestration site characteristics. The weights are calculated using a scaled dot product attention model.
6. The method according to claim 1, characterized in that, The secondary anomaly detection of the filtered data based on the machine learning model includes collecting data during normal operation of the device, training it with an autoencoder, and obtaining the standard deviation. Real-time operating data of the equipment in the carbon dioxide storage site is collected, and the absolute value of the difference between the real-time operating data and the standard deviation is calculated. This absolute value is then divided by the standard deviation to obtain the data result. When this data result does not exceed 5% of the standard deviation, it is judged as normal; otherwise, it is abnormal.
7. The method according to claim 2, characterized in that, The equipment health index is calculated based on the valid data obtained after secondary anomaly detection, including: The equipment health index is obtained by summing the standardized scores of each data indicator of several valid data points and the weight coefficients corresponding to each valid data point.
8. The method according to claim 7, characterized in that, The range method is used to normalize each of the several valid data points to obtain the standardized score.
9. The method according to claim 7, characterized in that, The weight coefficients mentioned above are obtained through training a random forest model. Obtain the standardized scores for all devices under monitoring status from the historical dataset; The obtained standardized scores are labeled, with 0 indicating normal and 1 indicating fault. Generate 200-300 decision trees, and randomly select several features for node splitting in each tree; The total reduction in Gini impurity resulting from the splitting of the i-th feature in all trees is calculated, and the weight coefficients are obtained after normalization.
10. The method according to claim 1, characterized in that, When the Equipment Health Index (EHI) is ≥ 80, the assessment result is normal; when 60 ≤ EHI < 80, the assessment result is at risk; when 40 ≤ EHI < 60, the assessment result is a warning. When the equipment health index EHI < 40, the assessment result is a failure.
Citation Information
Cited By
Automobile fault offline diagnosis method based on multi-sensor data fusion
CN121655888A
ScCO2 geological sequestration multi-mode safety evaluation system and method
CN122020433A
Multimodal Safety Assessment System and Method for scCO2 Geological Storage
CN122020433B