Intelligent monitoring method for CO2 leak sampling
By acquiring and analyzing CO2 leakage sampling data in real time, and dynamically controlling the sampling strategy, the problems of poor detection accuracy and efficiency in the existing technology are solved, and an efficient and reliable CO2 leakage sampling process is achieved.
Patent Information
- Application Number
- CN202411281715.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-13
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2044-09-13
AI Technical Summary
The prior art is difficult to effectively detect micro CO2 leakage points and cannot provide real-time data, resulting in inefficient sampling efficiency. Traditional methods rely on manual data recording, which is inefficient and error-prone.
The intelligent monitoring method is used to obtain sampling data in real time, calculate the actual sampling depth according to the preset strategy, and analyze the sampling data based on the preset sampling state classification model and analysis model, obtain the actual sampling state and temperature pressure change trend information, and dynamically regulate the sampling strategy or issue early warning information.
Real-time and accurate CO2 leakage detection and sampling status analysis are achieved, sampling efficiency and reliability are improved, and the risks of manual intervention and error are reduced.
Smart Images

Figure CN119223547B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of intelligent detection of carbon dioxide leakage, and in particular to an intelligent monitoring method for CO2 leakage sampling. Background Art
[0002] With the development of carbon capture and storage technology, CO2 flooding and storage projects are becoming increasingly popular. However, CO2 leakage is still one of the major challenges facing this technology; the ground monitoring equipment disclosed in traditional technology cannot effectively cover large areas of underground reservoirs, and it is difficult to detect tiny leakage points, resulting in the inability to provide real-time data, which makes it difficult for operators to understand the status of the sampling process in a timely manner, and it is difficult to accurately locate and assess the specific location and scale of the leakage, so that the sampling parameters cannot be adjusted according to real-time data, resulting in low sampling efficiency, or missing the best sampling time, and even causing the leakage problem to fail to be discovered and handled in time. In addition, traditional methods rely on manual data recording, which is not only inefficient but also prone to errors, limiting the comprehensive understanding of the leakage situation. Summary of the invention
[0003] In view of this, the embodiments of the present disclosure provide an intelligent monitoring method for CO2 leakage sampling, which can solve the problems existing in the prior art such as poor detection accuracy, poor efficiency, inability to quickly locate leakage points in real time, and inability to achieve comprehensive analysis.
[0004] In a first aspect, an embodiment of the present disclosure provides an intelligent monitoring method for CO2 leakage sampling, which specifically includes the following schemes:
[0005] Acquire sampling data in real time, wherein the sampling data includes actual temperature data and actual pressure data inside the sampling kettle and the ambient temperature of the sampling kettle;
[0006] Acquire the actual sampling depth according to the preset strategy and the sampling data;
[0007] Analyze the sampling data and the actual sampling depth based on a preset sampling state classification model to obtain an actual sampling state;
[0008] Analyze the sampled data based on a preset analysis model to obtain temperature and pressure change trend information;
[0009] The sampling strategy is dynamically adjusted based on the actual sampling state, the temperature and pressure change trend information, and the first preset strategy, and / or corresponding warning information is issued based on the actual sampling state, the temperature and pressure change trend information, and the second preset strategy.
[0010] Optionally, obtaining the actual sampling depth according to the preset strategy and the sampling data includes:
[0011] Obtaining a preset lowering length of the steel wire rope connected to the sampling kettle at the initial temperature through an encoder or a preset length sensor;
[0012] The ambient temperature of the sampling kettle in different areas is obtained based on the temperature sensor arranged on the lowered steel wire rope;
[0013] Based on the ambient temperature of each area segment and a preset temperature correction formula, a theoretical length of each area segment is obtained;
[0014] Superimposing the theoretical lengths of all the area segments to obtain the actual sampling depth;
[0015] The preset temperature correction formula is: L corrected =L measured ×(1+αΔT), where L corrected is the theoretical length of each zone segment, L measured is the preset lowering length, α is the linear expansion coefficient of the wire rope, and ΔT is the temperature change between the ambient temperature of each section and the initial temperature.
[0016] Optionally, analyzing the sampling data and the actual sampling depth based on a preset sampling state classification model to obtain the actual sampling state includes:
[0017] Determine the initial classification model;
[0018] Preprocessing the historical sampling data, training the initial classification model based on the preprocessed historical sampling data and a preset strategy, and using the initial classification model that meets the training requirements as the preset sampling state classification model;
[0019] Analyze the sampling data and the actual sampling depth based on the preset sampling state classification model to obtain the actual sampling state;
[0020] The historical sampling data includes several groups of historical parameters and corresponding status information, and the historical parameters include associated temperature data, pressure data and sampling depth data;
[0021] The actual sampling status includes any one of normal, warning, serious, and dangerous.
[0022] Optionally, the preset strategy includes any one of a cross-validation strategy and an interpolation reward strategy;
[0023] When the preset strategy is a cross-validation strategy, the initial classification model is trained based on the preprocessed historical sampling data and the preset strategy to obtain a preset sampling state classification model, including:
[0024] Dividing the preprocessed historical sampling data into a training set and a test set;
[0025] Optimize the target hyperparameters using a cross-validation method, train the initial classification model based on the optimized target hyperparameters, the training set, and the test set, and use the initial classification model that meets the conditions as the preset sampling state classification model;
[0026] The target hyperparameters include penalty parameters and kernel function parameters.
[0027] Optionally, when the preset strategy is an interpolation reward strategy, the initial classification model is trained based on the preprocessed historical sampling data and the preset strategy to obtain a preset sampling state classification model, including:
[0028] Determine the state space, action space, and target reward function of the configuration;
[0029] Training the initial classification model based on the state space, the action space, and the target reward function, and using the initial classification model that meets the training requirements as the preset sampling state classification model;
[0030] The state space is s, s=(T, P, D), T is the historical temperature data, P is the historical pressure data, and D is the historical sampling depth;
[0031] The action space is a, a∈{all types of historical execution actions}, and the historical execution actions include continuing normal sampling, adjusting parameters, or stopping sampling;
[0032] The target reward function is: R final =R total S(s,a), R total =α·R base +β·R interp -P uncert ; Among them, S(s,a) is the safety constraint function, R base is the basic reward, R interp is the interpolation reward, α is the weight of the base reward, β is the weight of the interpolation reward, P uncert is the uncertainty penalty term.
[0033] Optional, R base =w T ·f(ΔT)+w P ·f(ΔP)+w Qf(Q); where ΔT is the temperature deviation, ΔP is the pressure deviation, Q is the sampling quality, f(ΔT) is the scoring function (absolute value function or square function) for the absolute value or square value of the temperature deviation, f(ΔP) is the scoring function for the absolute value or square value of the pressure deviation, f(Q) is the scoring function for the absolute value or square value of the sampling quality, w T is the weight of temperature deviation, w P is the weight of pressure deviation, w Q is the weight of sampling quality;
[0034] R interp =GPR(s,a), GPR(s,a) is the interpolation reward value obtained based on the Gaussian process regression model;
[0035] P uncert =λ·σ 2 (s,a), where σ 2 (s, a) is the variance of the prediction uncertainty of the GPR model for the state and action, and λ is the penalty coefficient.
[0036] Optionally, the method further includes: optimizing the target reward function to obtain a data-driven reward function; training the initial classification model based on the state space, the action space, and the data-driven reward function, and using the initial classification model that meets the training requirements as the preset sampling state classification model;
[0037] The optimizing the target reward function to obtain the data-driven reward function includes: combining the knowledge of experts in the corresponding field and converting the expert rules into expert knowledge numerical rewards through a fuzzy logic system;
[0038] Determining a data-driven reward function based on the expert knowledge numerical reward and the target reward function;
[0039] The data-driven reward function includes: R = γ·R final +(1-γ)·R expert ; Among them, R expert is the expert knowledge numerical reward, γ is the weight coefficient, 0≤γ≤1.
[0040] Optionally, the method further includes: training and optimizing the preset sampling state classification model based on an offline reinforcement learning strategy, analyzing the sampling data and the actual sampling depth based on the trained and optimized preset sampling state classification model to obtain the actual sampling state;
[0041] The training and optimization of the preset sampling state classification model based on the offline reinforcement learning strategy includes:
[0042] Performing initial preprocessing on historical data to obtain a sample data set; the initial preprocessing includes time series alignment, standardization, and data enhancement;
[0043] Obtain importance weights based on the determined behavior strategy probability density and target strategy probability density;
[0044] The importance weight is ω(s,a): π b (a|s) is the probability density of behavior strategy, π e (a|s) is the target strategy probability density;
[0045] determining a truncated importance sampling based on the importance weight;
[0046] The truncated importance sampling is ω TIS (s,a):ω TIS (s,a)=min(ω(s,a),c), c is the cutoff threshold;
[0047] The weight of the sample data set is adjusted based on the importance weight, and the preset sampling state classification model is trained and optimized based on the sample data set after the weight adjustment, a preset offline evaluation method, a confidence interval estimation strategy, a preset optimization strategy, an offline-online hybrid evaluation strategy, and a model integration strategy.
[0048] Optionally, the preset offline evaluation method includes one or more of a direct method and a counterfactual estimation method;
[0049] The preset optimization strategy includes one or more of conservative strategy iteration and batch constrained deep Q learning.
[0050] Optionally, analyzing the sampled data based on a preset analysis model to obtain temperature and pressure change trend information includes:
[0051] Determine the preset analysis model;
[0052] The preset analysis model is trained based on historical time series data, and the sampled data is analyzed by the trained preset analysis model to obtain temperature and pressure change trend information;
[0053] The historical time series data includes historical temperature time series data and historical pressure time series data.
[0054] In a second aspect, the present disclosure also provides an intelligent monitoring system for CO2 leakage sampling, including:
[0055] The sampling module is used to obtain sampling data in real time, wherein the sampling data includes actual temperature data and actual pressure data inside the sampling kettle and the ambient temperature of the sampling kettle;
[0056] A depth acquisition module, used to acquire the actual sampling depth according to a preset strategy and the sampling data;
[0057] A state acquisition module, used to analyze the sampling data and the actual sampling depth based on a preset sampling state classification model to acquire the actual sampling state;
[0058] A temperature and pressure change trend acquisition module, used to analyze the sampled data based on a preset analysis model to obtain temperature and pressure change trend information;
[0059] The analysis module is used to dynamically adjust the sampling strategy based on the actual sampling state, the temperature and pressure change trend information, and the first preset strategy, or to issue corresponding warning information based on the actual sampling state, the temperature and pressure change trend information, and the second preset strategy.
[0060] In a third aspect, the embodiments of the present disclosure further provide a computer device, which adopts the following technical solution:
[0061] The computer device comprises:
[0062] at least one processor; and,
[0063] a memory communicatively connected to the at least one processor; wherein,
[0064] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute any of the above-mentioned intelligent monitoring methods for CO2 leakage sampling.
[0065] In a fourth aspect, an embodiment of the present disclosure further provides a computer-readable storage medium, which stores computer instructions, and the computer instructions are used to enable a computer to execute any of the above-mentioned intelligent monitoring methods for CO2 leakage sampling.
[0066] In a fifth aspect, an embodiment of the present disclosure further provides a computer program product, including a computer program / instruction, which implements the steps of any of the above methods when executed by a processor.
[0067] The intelligent monitoring method for CO2 leakage sampling disclosed in this embodiment can ensure subsequent real-time analysis feedback by acquiring sampling data in real time; then the actual sampling depth is acquired according to the preset strategy and sampling data, and the external influence in the sampling environment is taken into account to ensure that accurate sampling depth information is acquired, so as to provide reliable basic information for subsequent decision-making; then the sampling data and the actual sampling depth are analyzed based on the preset sampling state classification model to acquire the actual sampling state, and the sampling data is analyzed based on the preset analysis model to obtain the temperature and pressure change trend information, so that the actual sampling state and the temperature and pressure change trend information can be automatically and quickly acquired, and the risk of leakage is reduced. Low manual intervention; finally, the sampling strategy is dynamically adjusted based on the actual sampling status, temperature and pressure change trend information, and the first preset strategy, or the corresponding early warning information is issued based on the actual sampling status, temperature and pressure change trend information, and the second preset strategy. Intelligent analysis can be performed based on the actual monitoring situation to effectively improve the accuracy and reliability of sampling; the method disclosed in the present application ensures the acquisition of high-fidelity CO2 leakage samples through real-time monitoring and intelligent control; real-time monitoring of the status in the kettle, timely detection of abnormal situations, and reduction of sampling risks; high degree of automation, reduction of manual operations, and improvement of sampling efficiency, providing more accurate and reliable basic data for CO2 leakage assessment.
[0068] The above description is only an overview of the technical solution of the present invention. In order to more clearly understand the technical means of the present invention, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present invention more obvious and easy to understand, the following preferred embodiments are specifically cited and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0070] Figure 1 A schematic flow chart of an intelligent monitoring method for CO2 leakage sampling provided in an embodiment of the present disclosure.
[0071] Figure 2 A schematic flow chart of a method for obtaining an actual sampling depth provided in an embodiment of the present disclosure.
[0072] Figure 3 A schematic flow chart of a method for acquiring an actual sampling state provided in an embodiment of the present disclosure.
[0073] Figure 4A flowchart of a method for obtaining a preset sampling state classification model based on a cross-validation strategy provided in an embodiment of the present disclosure.
[0074] Figure 5 A flowchart of a method for obtaining a preset sampling state classification model based on a preset offline evaluation method provided in an embodiment of the present disclosure.
[0075] Figure 6 A flowchart of a method for training and optimizing a preset sampling state classification model based on an offline reinforcement learning strategy provided in an embodiment of the present disclosure.
[0076] Figure 7 A schematic flow chart of a method for obtaining temperature and pressure change trend information provided in an embodiment of the present disclosure.
[0077] Figure 8 A schematic diagram of the structure of a computer device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION
[0078] The embodiments of the present disclosure are described in detail below with reference to the accompanying drawings.
[0079] It should be clear that the following embodiments of the present disclosure are described by specific specific examples, and those skilled in the art can easily understand other advantages and effects of the present disclosure from the contents disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all of the embodiments. The present disclosure can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present disclosure. It should be noted that the following embodiments and features in the embodiments can be combined with each other in the absence of conflict. Based on the embodiments in the present disclosure, all other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of the present disclosure.
[0080] It should be noted that various aspects of the embodiments within the scope of the appended claims are described below. It should be apparent that the aspects described herein may be embodied in a wide variety of forms, and any specific structure and / or function described herein is merely illustrative. Based on the present disclosure, it should be understood by those skilled in the art that an aspect described herein may be implemented independently of any other aspect, and two or more of these aspects may be combined in various ways. For example, any number of aspects described herein may be used to implement the device and / or practice the method. In addition, other structures and / or functionalities other than one or more of the aspects described herein may be used to implement this device and / or practice this method.
[0081] It should also be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present disclosure. The drawings only show components related to the present disclosure rather than being drawn according to the number, shape and size of components in actual implementation. In actual implementation, the type, quantity and proportion of each component may be changed arbitrarily, and the component layout may also be more complicated.
[0082] Additionally, in the following description, specific details are provided to facilitate a thorough understanding of the examples. However, it will be understood by those skilled in the art that the aspects described may be practiced without these specific details.
[0083] Reference Figure 1 , the present application discloses an intelligent monitoring method for CO2 leakage sampling, comprising:
[0084] S100, acquiring sampling data in real time, the sampling data including actual temperature data and actual pressure data inside the sampling kettle and the ambient temperature of the sampling kettle.
[0085] In this application, a sampling kettle is used to collect CO2 leakage samples; a temperature sensor and a pressure sensor are installed in the sampling kettle to monitor the temperature and pressure inside the sampling kettle in real time; specifically, in actual monitoring, the temperature sensor and the pressure sensor continuously monitor the state of the inner cavity of the sampling kettle, and the thick film integrated circuit collects sensor data at a preset frequency.
[0086] The thick film integrated circuit in this embodiment is a high temperature resistant (150° C.) thick film integrated circuit, which can ensure stable operation in harsh environments.
[0087] Through this step, real-time data collection can be provided, making the monitoring of sampling conditions more timely and accurate; high-precision sensors can provide reliable data and reduce the risks caused by sensor failure or data errors; at the same time, the impact of ambient temperature changes on the sampling process can be tracked in real time.
[0088] S200, obtaining an actual sampling depth according to a preset strategy and sampling data.
[0089] S300, analyzing the sampling data and the actual sampling depth based on a preset sampling state classification model to obtain the actual sampling state.
[0090] The sampling status can be automatically classified by presetting the sampling status classification model, reducing manual intervention and improving efficiency. The sampling status can be obtained in real time, which helps to quickly identify and handle abnormal situations.
[0091] S400, analyzing the sampled data based on a preset analysis model to obtain temperature and pressure change trend information.
[0092] The preset analysis model can timely discover the changing trends of temperature and pressure, help predict potential problems, and provide a basis for strategy adjustments and early warnings.
[0093] S500, dynamically adjusting the sampling strategy based on the actual sampling state, the temperature and pressure change trend information, and the first preset strategy, and / or issuing corresponding warning information based on the actual sampling state, the temperature and pressure change trend information, and the second preset strategy.
[0094] The sampling strategy includes one or more of sampling depth and sampling time.
[0095] In this step, the first situation is: the sampling strategy can be dynamically adjusted according to real-time data to optimize the sampling effect, or, dynamic regulation is not triggered, but early warning information is issued in time to help take measures to deal with possible abnormal or dangerous situations; at the same time, it also includes the second situation, which can not only dynamically adjust the sampling strategy according to the first preset strategy, but also issue corresponding early warning information according to the second preset strategy.
[0096] The intelligent monitoring method for CO2 leakage sampling disclosed in this embodiment can ensure subsequent real-time analysis feedback by acquiring sampling data in real time; then, the actual sampling depth is acquired according to the preset strategy and sampling data, and the external influence in the sampling environment is taken into account to ensure that accurate sampling depth information is acquired, so as to provide reliable basic information for subsequent decision-making; then, the sampling data and the actual sampling depth are analyzed based on the preset sampling state classification model to acquire the actual sampling state, and the sampling data is analyzed based on the preset analysis model to obtain the temperature and pressure change trend information, which can automatically and quickly acquire the actual sampling state and the temperature and pressure change trend information, thereby reducing manual intervention; finally, the sampling strategy is dynamically adjusted based on the actual sampling state, the temperature and pressure change trend information, and the first preset strategy, or the corresponding warning information is issued based on the actual sampling state, the temperature and pressure change trend information, and the second preset strategy, and intelligent analysis can be performed based on the actual monitoring situation to effectively improve the accuracy and reliability of sampling.
[0097] Specifically, by dynamically adjusting the sampling strategy to ensure that the sampling process adapts to the actual environment and conditions, the real-time monitoring and early warning system can quickly identify and respond to potential problems, reduce risks, use data analysis results to make scientific decisions, and improve the overall efficiency and reliability of the system; this method effectively improves the accuracy and intelligence level of the CO2 leakage sampling process, thereby optimizing the sampling effect and risk management.
[0098] For the dynamic control of the sampling strategy, specifically, an adjustment instruction may be sent to the sampling system through the controller to achieve dynamic optimization of parameters such as sampling depth and sampling time.
[0099] For data transmission, carrier communication technology is used to transmit data to the ground controller through a seven-core armored cable, breaking through the technical bottleneck of information redundancy in long-distance communications underground and achieving stable data transmission at a depth of 4,000 meters.
[0100] Further, the surface controller receives and processes the transmitted data for data analysis, sampling system control and feedback.
[0101] In this embodiment, the transmitted data is received and processed by the ground controller. Specifically, a high-performance embedded processor, such as ARM Cortex-A53, can be used to realize real-time data processing. An FPGA-based parallel processing architecture can be used to improve the data processing speed. At the same time, real-time storage, backup and visual display of data can be realized.
[0102] The order of S400 and S300 can be changed, and the present application does not limit the steps for obtaining the actual sampling status and the temperature and pressure change trend information.
[0103] Furthermore, corresponding warning information is issued based on the actual sampling status, temperature and pressure change trend information, and the second preset strategy; specifically, multi-level thresholds can be set to achieve early warning of abnormal conditions.
[0104] Among them, the multi-level thresholds set include: 1) Normal threshold: no alarm is triggered within the normal working range; 2) Warning threshold: when the temperature or pressure approaches the limit, a warning signal is issued to prompt relevant personnel to pay attention; 3) Early warning threshold: if the temperature or pressure exceeds the slight abnormal range, an early warning message is issued in time; 4) Alarm threshold: when the situation is serious and exceeds the safe working range, an alarm is triggered and corresponding measures are taken.
[0105] Alternatively, multiple warning levels may be set, such as normal, warning, severe, and dangerous. Specifically, a specific threshold may be defined for each parameter (such as temperature, pressure) at each level.
[0106] During the analysis and judgment, the real-time data of the sampling system is continuously monitored, and the real-time value of each parameter is compared with the predefined threshold value; the status of multiple parameters is comprehensively considered, and rules are formulated to determine the overall status of the system, such as using the status of the most serious parameter as the overall status; based on the results of the comprehensive evaluation, the corresponding level of warning is triggered; different response measures are designed for different levels of warnings, such as sending notifications, initiating emergency procedures, etc.
[0107] Furthermore, dynamic threshold adjustments can be made, including: regularly analyzing historical data to evaluate the rationality of current thresholds; combining expert knowledge and statistical analysis to dynamically adjust thresholds to improve the accuracy of early warnings.
[0108] Furthermore, all warning events can be recorded, including timestamp, warning level, trigger parameters and other information; warning logs can be analyzed regularly for system performance evaluation and improvement.
[0109] Through these detailed steps, the CO2 leakage sampling intelligent monitoring method can more accurately classify the sampling status, predict the temperature and pressure change trend, and issue multi-level warnings in time. This method can significantly improve the reliability and efficiency of the system and provide strong technical support for the monitoring of CO2 leakage in sealed storage.
[0110] Reference Figure 2 , the method for obtaining the actual sampling depth includes:
[0111] S210, obtaining a preset lowering length of the steel wire rope at an initial temperature through an encoder or a preset length sensor.
[0112] S220, obtaining the ambient temperature of the lowered steel wire rope in different sections based on the temperature sensor.
[0113] By obtaining the ambient temperature of each zone segment, adjustments can be made for different environmental conditions to ensure that the theoretical length of each zone segment is more accurate.
[0114] Dividing the entire process into multiple area segments for analysis can make the entire measurement process more detailed and better understand the impact of each area on the actual sampling depth.
[0115] S230, obtaining a theoretical length of each area segment based on the ambient temperature of each area segment and a preset temperature correction formula.
[0116] S240, superimposing the theoretical lengths of all the area segments to obtain the actual sampling depth.
[0117] The preset temperature correction formula is: L corrected =L measured ×(1+αΔT), where L corrected is the theoretical length of each zone segment after temperature correction, L measured is the preset lowering length, α is the linear expansion coefficient of the wire rope, and ΔT is the temperature change between the ambient temperature of each section and the initial temperature.
[0118] In this embodiment, temperature can cause the material to expand or contract. The use of the temperature correction formula can effectively compensate for the measurement error caused by the change in ambient temperature and improve the accuracy of the actual sampling depth measurement. Different temperature environments may be encountered at different depths. Through real-time data monitoring and correction, the actual state of the wire rope during the lowering process can be more realistically reflected.
[0119] Accurate depth measurement can ensure the rational use of wire ropes and reduce equipment damage or accidents caused by inaccurate lowering depths. This method, combined with an encoder or preset length sensor, a temperature sensor and a correction formula, can more accurately obtain the actual sampling depth of the wire rope, thereby improving the accuracy, safety and efficiency of the measurement, and ultimately achieving cost reduction and rational use of resources. This method has important application value in actual operations.
[0120] Furthermore, if necessary, the tension data on the wire rope can be obtained based on the tension sensor, and the equipment attitude angle can be obtained based on the gyroscope to meet the comprehensive collection requirements of measured data.
[0121] Furthermore, the Kalman filter algorithm can be used to fuse multi-source data to improve the accuracy of depth calculation.
[0122] Reference Figure 3 , the method for obtaining the actual sampling state specifically includes:
[0123] S310, determining an initial classification model.
[0124] Specifically, the initial classification model may be a model based on machine learning, for example, a decision tree, a support vector machine (SVM) or a neural network.
[0125] For model settings, the parameters of the initial classification model may correspond to the depth of the tree, the kernel function type of the SVM, or the number of layers of the neural network and the number of neurons in each layer.
[0126] The initial classification model provides a benchmark against which subsequent model training and optimization can be evaluated; by selecting and setting the initial classification model, you can ensure a structured starting point.
[0127] S320, preprocessing the historical sampling data, training the initial classification model based on the preprocessed historical sampling data and a preset strategy, and using the initial classification model that meets the training requirements as a preset sampling state classification model.
[0128] The historical sampling data includes several groups of historical parameters and corresponding status information, and the historical parameters include associated temperature data, pressure data and sampling depth data; the preset strategy includes any one of a cross-validation strategy and an interpolation reward strategy.
[0129] Specifically, preprocessing can include data cleaning, standardization, feature selection, etc.
[0130] Among them, data cleaning can be done by removing missing values and outliers to ensure data quality; standardization refers to standardizing the selected features to ensure that all features are on the same scale; feature selection can be done by selecting features that play an important role in classification and reducing dimensions to improve the efficiency of the model.
[0131] In this step, the preprocessed data can improve the training effect of the model and make the model's prediction of the actual sampling status more accurate; through training and optimization, the overfitting of the model to the training data can be reduced and the performance on new data can be improved, so that the model can adapt to various changes and situations encountered in the actual sampling process.
[0132] S330, analyzing the sampling data and the actual sampling depth based on a preset sampling state classification model to obtain the actual sampling state.
[0133] The actual sampling status includes any one of normal, warning, serious, and dangerous.
[0134] Through this step, the sampling status can be monitored in real time and the status prediction can be carried out to detect potential problems in time; the prediction based on the classification model can accurately judge the sampling status, avoid errors in human judgment, help make decisions quickly, and take corresponding measures to ensure the safety and effectiveness of the operation.
[0135] The method for obtaining the actual sampling status disclosed in this embodiment can analyze the sampling data more accurately and ensure accurate judgment of the actual sampling status by training and optimizing the classification model; the sampling status can be obtained in time and corresponding measures can be taken, which is helpful to prevent and reduce potential safety hazards that may occur during the sampling process; the accurate prediction of the sampling status can help to reasonably allocate resources and improve the efficiency and effectiveness of the sampling process; at the same time, the data analysis results and predictions provided by the system can support decision makers to make more scientific decisions and improve the intelligence level of the overall workflow.
[0136] Reference Figure 4 When the preset strategy is a cross-validation strategy, the initial classification model is trained based on the preprocessed historical sampling data and the preset strategy to obtain a preset sampling state classification model, including:
[0137] A100, divides the preprocessed historical sampling data into training set and test set.
[0138] Specifically, the preprocessed historical sampling data may be divided according to a certain ratio (eg, 80% as a training set and 20% as a test set).
[0139] Another common ratio is 70% training set and 30% test set.
[0140] In some cases, a portion of the training set can be further divided into a validation set for tuning the model's hyperparameters.
[0141] In order to avoid division bias, a random division method is usually used to ensure the representativeness of the training set and test set.
[0142] The training set is used to train the model and adjust the model parameters, and the test set is used to evaluate the performance of the model and ensure the generalization ability of the model. By dividing the data into training sets and test sets, the overfitting of the model can be effectively detected.
[0143] A200 uses the cross-validation method to optimize the target hyperparameters, trains the initial classification model based on the optimized target hyperparameters, training set, and test set, and uses the initial classification model that meets the conditions as the preset sampling state classification model.
[0144] Among them, the target hyperparameters include penalty parameters and kernel function parameters.
[0145] When the preset sampling state classification model is the SVM model, the penalty parameter C controls the width of the classification boundary and the tolerance for classification errors. A larger C value means a lower tolerance for misclassification, which may lead to overfitting; a smaller C value allows more misclassifications, which may lead to underfitting.
[0146] For the kernel function parameter γ, in the model using the kernel function, γ controls the range of influence of the data point on the decision boundary. A larger γ value will cause the model to focus on a smaller area, which may lead to overfitting; a smaller γ value will cause the model to focus on a larger area, which may lead to underfitting.
[0147] In this step, optimizing hyperparameters through cross-validation can significantly improve the predictive performance of the model. Cross-validation can provide a more accurate estimate of the generalization ability of the model and reduce the risk of overfitting. Hyperparameter optimization and cross-validation ensure the reliability and stability of the model.
[0148] Specifically, the model performance is evaluated on the test set, and indicators such as confusion matrix and classification report are used to evaluate the model's accuracy, precision, recall, etc.
[0149] The method for obtaining the preset sampling state classification model disclosed in this embodiment, through cross-validation and hyperparameter optimization, the obtained preset sampling state classification model has better performance on training data and test data, and can more accurately predict the sampling state; the cross-validation method can effectively detect the performance of the model on different data subsets, and improve the stability and robustness of the model; precise hyperparameter adjustment enables the model to better capture data features and improve the prediction accuracy of the actual sampling state; randomly dividing the training set and the test set and using cross-validation can effectively reduce the bias of data division and ensure the fairness and effectiveness of model training; the preset sampling state classification model with good generalization ability can provide more reliable data support for actual operations and improve the overall decision-making level.
[0150] Furthermore, the cross-validation method includes one or more of a k-fold cross-validation method, a leave-one-out cross-validation method, and a stratified cross-validation method.
[0151] Among them, the k-fold cross-validation method includes: dividing the data set into k folds, using k-1 folds for training each time, and using the remaining fold for validation; repeating k times, each time selecting a different fold as the validation set, and the final model performance evaluation value is the average of the k validation results.
[0152] The leave-one-out cross validation method is a special k-fold cross validation, where k is equal to the total number of samples in the dataset. Each time one sample is used as the validation set, and the rest of the samples are used as the training set. This method is computationally expensive, but can provide a more accurate performance evaluation.
[0153] When the stratified cross-validation method is used for classification problems, it ensures that the proportion of samples in each category in each fold is consistent with the proportion in the entire dataset.
[0154] Specifically, the process of optimizing the target hyperparameters, that is, using cross-validation to optimize the hyperparameters, includes: 1) selecting candidate hyperparameter sets: such as different combinations of C and γ values; 2) evaluating each set of hyperparameters: calculating the performance indicators of the model under each set of hyperparameters through cross-validation (such as accuracy, F1 score, etc.); 3) selecting the best hyperparameters: selecting the hyperparameter combination with the best performance based on the cross-validation results.
[0155] Further, the model is trained on the entire training set using the best selected hyperparameters. Make sure the data used includes all possible features and labels to fully train the model; select an algorithm suitable for the task (such as support vector machine, decision tree, neural network, etc.) and train it using the best hyperparameters.
[0156] Evaluate the performance of the final trained model on the test set and calculate various performance indicators (such as accuracy, recall, F1 score, etc.); if the performance on the test set is poor, further adjustments can be made, such as adding features, adjusting the model structure, collecting more data, etc.
[0157] In the case of an independent validation set, after training the model on the training set, the validation set can be used for final tuning and confirmation; multiple validations can be performed using different validation sets to ensure the stability and generalization ability of the model. Through these steps, a machine learning model with good generalization performance can be effectively trained to ensure its performance in practical applications.
[0158] Reference Figure 5 When the preset strategy is the interpolation reward strategy, the initial classification model is trained based on the preprocessed historical sampling data and the preset strategy to obtain the preset sampling state classification model, including:
[0159] B100, determine the configured state space, action space and target reward function.
[0160] The state space is s, s = (T, P, D), T is the historical temperature data, P is the historical pressure data, and D is the historical sampling depth.
[0161] The action space is a, a∈{all types of historical execution actions}. Historical execution actions include continuing normal sampling, adjusting parameters or stopping sampling. Of course, historical execution actions can also include other actions to ensure that the action space covers all possible operation options.
[0162] The target reward function is: R final =R total S(s,a), R total =α·R base +β·R interp -P uncert ; Among them, S(s,a) is the safety constraint function, R base is the basic reward, R interp is the interpolation reward, α is the weight of the base reward, β is the weight of the interpolation reward, P uncert is the uncertainty penalty term.
[0163] Safety constraint function, that is, defining safety constraints, for example, when the detected CO2 concentration exceeds a certain threshold, sampling must be stopped immediately.
[0164] S(s,a)=ifCO2 concentration>threshold,then 0else 1.
[0165] The base reward is a reward for normal operations, which defines the base reward value in the absence of special circumstances.
[0166] The interpolation reward is a reward for processing data interpolation, which defines the reward value when the system needs to interpolate data.
[0167] The uncertainty penalty term is a penalty for the uncertainty of the system, and is defined as a penalty term for uncertainty that may occur during the operation process.
[0168] The weight parameters are the weights of the basic reward and the interpolation reward, respectively, to control the relative importance of the two in the total reward; the safety constraint function is used to ensure the safety of executing actions. Define the constraint function value of state s and action a on safety.
[0169] In this step, the definitions of state space and action space are clarified so that the model training has a clear direction; the target reward function is combined with the basic reward, interpolation reward and uncertainty penalty term to optimize the overall performance of the system and ensure that the reward is consistent with the system goal; the safety constraint function can ensure that the actions taken during model training are safe and reduce potential risks.
[0170] Furthermore, multiple levels of thresholds may be set for each indicator, such as normal, warning, and danger.
[0171] B200, trains the initial classification model based on the state space, action space, and target reward function, and uses the initial classification model that meets the training requirements as the preset sampling state classification model.
[0172] Specifically, state-action pairs can be constructed based on the defined state space and action space. The actions taken in each state and their effects in the historical data are collected. Then, the reward value of each state-action pair is calculated based on the target reward function, using the defined base reward, interpolation reward, and uncertainty penalty terms.
[0173] For model training, you can choose a suitable reinforcement learning algorithm (such as Q-learning, deep Q network DQN or policy gradient method) to train the classification model. During the training process, you can use state-action pairs and reward values to train the model so that the model can learn the best strategy by optimizing the reward function. For model verification, you can verify the performance of the model on the test set to ensure that it can accurately predict the sampled state and make reasonable decisions.
[0174] The method disclosed in this embodiment uses an interpolation reward strategy so that the classification model can be adaptively optimized in various situations to improve the accuracy of prediction; the model can make more appropriate decisions based on the reward mechanism to ensure the effectiveness and safety of the sampling process; by introducing safety constraints and uncertainty penalty items, the risks in operations can be effectively managed and reduced; the overall solution improves the intelligence level of the system and makes the sampling process more scientific and efficient.
[0175] In summary, the interpolation reward strategy can significantly improve the performance of the classification model and optimize the decision quality and system stability in the actual sampling process through precise reward mechanism and training method.
[0176] In this embodiment, α determines the importance of the base reward in the final comprehensive reward, and β determines the influence of the interpolation reward in the comprehensive reward. S(s,a) is a safety constraint function used to ensure that the action a taken in a certain state s does not violate the safety constraint.
[0177] This formula combines multiple reward mechanisms to calculate the total reward value, which takes into account the basic reward, interpolation reward and uncertainty penalty. This comprehensive reward mechanism can guide the agent (i.e., model) to learn the best strategy. The α and β in the formula allow for flexible adjustment of the impact of different reward sources so that the agent can effectively learn and optimize the strategy; when it is necessary to optimize system performance, the comprehensive reward mechanism can help meet the basic goals while also taking into account some additional goals and system uncertainties; in automated control systems, it can ensure that the system meets the requirements of safety and stability while achieving its goals.
[0178] In this embodiment, the basic reward R base and the interpolation reward R interp Represent the main reward signal and the supplementary reward signal respectively; the uncertainty penalty term P uncert It is used to reduce the impact of system uncertainty or instability on the total reward. Through this formula, the weights of each reward and penalty item can be adjusted according to the needs of specific applications, thereby achieving better performance and stability in practical applications.
[0179] Specifically, the basic reward is: R base =w T ·f(ΔT)+w P ·f(ΔP)+w Q ·f(Q).
[0180] Among them, ΔT is the temperature deviation, which can also be understood as the temperature change. For example, if the system needs to maintain a stable temperature, ΔT can represent the difference between the current temperature and the target temperature. f(ΔT) is a scoring function (i.e., an absolute value function or a square function) for the absolute value or square value of the temperature deviation ΔT, which is usually used to map the temperature change to a reward value. For example, a function can be used to map the magnitude of the temperature change to a positive or negative reward.
[0181] ΔP is the pressure deviation, which can also be understood as the pressure change. For example, if the system needs to maintain a stable pressure, ΔP can represent the difference between the current pressure and the target pressure. f(ΔP) is a scoring function for the absolute value or square value of the pressure deviation, which is used to map the pressure change to a reward value.
[0182] Q is the sampling quality, and f(Q) is the scoring function for the absolute value or square value of the sampling quality Q.
[0183] w T is the weight of temperature deviation, w P is the weight of pressure deviation, w Q These three weights determine the importance of each factor in the basic reward.
[0184] Assuming that in some cases we do not have complete sampling data or the data is noisy, the GPR model can be used for reward interpolation.
[0185] R interp =GPR(s,a), GPR(s,a) is the interpolated reward value obtained by reward interpolation based on the Gaussian process regression (GPR) model, that is, the RBF kernel function is used to predict the reward of the unknown state; specifically, the Gaussian process regression (GPR) model predicts the corresponding interpolated reward based on the given state-action pair (s,a), (s,a) represents the state s of the system and the action a taken, and the Gaussian process regression (GPR) model uses this information to estimate the interpolated reward, that is, the predicted reward value, which is usually used to supplement the basic reward function to provide more comprehensive reward information.
[0186] Use Gaussian process regression (GPR) for reward interpolation, including: selecting kernel functions: such as radial basis function (RBF) kernel; training GPR model: using known state-reward pairs; predicting rewards for unknown states: R interp =GPR(s,a).
[0187] Furthermore, the reward interpolation R interp It can be adjusted dynamically. As the system runs and data collection increases, the GPR model can be continuously updated to provide more accurate reward predictions.
[0188] Gaussian Process Regression (GPR) is used to learn from existing reward data and interpolate unobserved state-action pairs. interp It is the reward value predicted by the GPR model, which is used to supplement the basic reward function and provide more comprehensive reward information. This method can handle data sparsity and uncertainty, and improve the accuracy of the reward function and the performance of the system.
[0189] The goal of reward interpolation based on the Gaussian Process Regression (GPR) model is to predict or estimate the interpolation reward R for a state-action pair (s, a) through the GPR model. interp ,This process is to make the reward function more accurately reflect the actual situation of the system, especially when facing uncertainty or sparse data.
[0190] P uncert =λ·σ 2 (s,a), where σ 2 (s,a) is the variance of the GPR model's prediction uncertainty for state s and action a. In the GPR model, the variance reflects the model's uncertainty about the predicted value. The larger the variance, the higher the model's uncertainty about the predicted value. Usually, this happens when the training data is sparse or the state-action pair has not been fully explored.
[0191] λ is the penalty coefficient, which is used to adjust the intensity of the uncertainty penalty term. It determines the weight of uncertainty in the total penalty. A larger λ value will increase the impact of uncertainty on the total reward, prompting the system to pay more attention to and reduce uncertainty. A smaller λ value will reduce the impact of uncertainty on the total reward, making the system less sensitive to uncertainty.
[0192] In this embodiment, the uncertainty penalty term P uncert is designed to handle and quantify uncertainty in predictions; specifically, it is used to penalize state-action pairs for which the model’s prediction uncertainty is high, thereby prompting the system to act more cautiously in these areas of high uncertainty.
[0193] In this embodiment, the parameters of the reward function can also be dynamically adjusted by monitoring the performance of the system (e.g., the number and accuracy of detected CO2 leakage events). A sliding time window can be used to capture the effects on different time scales.
[0194] Furthermore, the intelligent monitoring method for CO2 leakage sampling disclosed in the present application also includes: optimizing the target reward function to obtain a data-driven reward function; training the initial classification model based on the state space, action space, and data-driven reward function, and using the initial classification model that meets the training requirements as a preset sampling state classification model.
[0195] Specifically, the target reward function is optimized to obtain a data-driven reward function, including:
[0196] Combining the knowledge of experts in the corresponding field, the expert rules are converted into expert knowledge numerical rewards through the fuzzy logic system;
[0197] Determine the data-driven reward function based on expert knowledge numerical rewards and target reward functions;
[0198] The data-driven reward function includes: R = γ·R final +(1-γ)·R expert ; Among them, R expert is the expert knowledge numerical reward, γ is the weight coefficient, 0≤γ≤1.
[0199] Where γ represents the data-driven reward (R final ) and expert knowledge reward (R expert ) in the overall reward. For example, γ = 0.7 means that the data-driven reward accounts for 70% of the overall reward, while the expert knowledge reward accounts for 30%.
[0200] Methods for selecting γ include: selecting a suitable γ value through experiments and cross-validation; this usually involves evaluating the performance of the system under different γ values to find the optimal balance; in some cases, γ can be set based on the experience or domain knowledge of experts. For example, if expert knowledge is particularly important in a specific field, a higher weight may be given; in some applications, γ can be dynamically adjusted based on real-time feedback from the system. For example, when data-driven rewards perform poorly, the weight of expert knowledge rewards can be increased, and vice versa.
[0201] Specifically, combining domain expert knowledge and converting it into expert knowledge numerical rewards through a fuzzy logic system may include: 1) Collecting expert knowledge: Collaborating with domain experts to collect knowledge and experience about CO2 leakage monitoring and sampling, including reward and penalty rules in different scenarios; 2) Establishing a fuzzy logic system: Using a fuzzy logic system to convert expert rules into numerical rewards. For example, experts may provide suggestions on different sampling conditions, such as "increase rewards under high temperature conditions" or "reduce rewards when pressure is abnormal"; 3) Defining fuzzy variables and rules: defining fuzzy variables (such as "high temperature" or "low pressure") and fuzzy rules (such as "if the temperature is high and the pressure is low, the reward increases"), and converting these rules into numerical form.
[0202] Example: Fuzzy variables: temperature, pressure; fuzzy rule: if temperature > 75°C and pressure < 100kPa, the reward increases by 10%. This solution ensures that the reward function takes into account the actual experience and expertise of domain experts, enhancing the practical application capabilities of the model; the fuzzy logic system converts complex expert rules into quantified rewards, improving the systematicity and consistency of processing.
[0203] In this embodiment, the optimization of the data-driven reward function can significantly improve the prediction and decision-making capabilities of the model in practical applications, integrate expert knowledge into the reward function, improve the performance of the model in specific fields, and make it more in line with actual needs; adjust the influence of the reward through the weight coefficient so that the model can adapt to different operating environments and conditions; the improved reward function can more accurately reflect the key factors in the sampling process, thereby optimizing decision-making quality and operational efficiency; combine data-driven and expert knowledge to make the intelligent monitoring method more intelligent and efficient, and adapt to complex scenarios in practical applications.
[0204] Overall, the optimization of the data-driven reward function can comprehensively improve the performance and application effect of the intelligent monitoring method in CO2 leakage sampling by combining actual data and expert knowledge.
[0205] Furthermore, the intelligent monitoring method for CO2 leakage sampling disclosed in the present application also includes: training and optimizing a preset sampling state classification model based on an offline reinforcement learning strategy, analyzing the sampling data and actual sampling depth based on the trained and optimized preset sampling state classification model to obtain the actual sampling state.
[0206] Reference Figure 6 The method for training and optimizing the preset sampling state classification model based on the offline reinforcement learning strategy specifically includes:
[0207] C100, perform initial preprocessing on historical data to obtain a sample data set; the initial preprocessing includes time series alignment, standardization, and data enhancement.
[0208] Specifically, historical data is aligned according to timestamps to ensure that data at all time points are consistent; during processing, data from different sources at the same time point are ensured to be correctly aligned.
[0209] Standardize each feature so that they have the same scale. Common standardization methods include Z-score standardization (converting data to a standard normal distribution with a mean of 0 and a variance of 1); sample data can be enhanced, such as increasing sample diversity through interpolation, data smoothing, etc.
[0210] In this step, time alignment ensures data consistency, standardization reduces dimensionality effects, and data enhancement increases sample diversity, improving data quality and model training effects. Data preprocessing reduces training deviations caused by different data sources or insufficient data.
[0211] C200, obtains the importance weight based on the determined behavior strategy probability density and target strategy probability density.
[0212] The importance weight is ω(s,a): π b (a|s) is the probability density of behavior strategy, π e (a|s) is the target strategy probability density.
[0213] The behavior strategy refers to the probability of taking a specific action in a given state, and its probability density function is estimated using the behavior strategy recorded in the historical data. The target strategy refers to the desired strategy in the optimization process, which is usually the current prediction strategy of the model, and its probability density function is estimated using the strategy in the training process.
[0214] In this step, the importance weights can accurately reflect the differences between the behavioral strategy and the target strategy, which helps to correct and optimize the model; by calculating the weights, historical data can be better utilized for strategy optimization, improving the adaptability and effectiveness of the model.
[0215] C300, truncated importance sampling based on importance weights.
[0216] Truncated importance sampling is ω TIS (s,a):ω TIS (s,a)=min(ω(s,a),c), where c is the cutoff threshold, which is used to limit the range of importance weights, thereby reducing the impact of extreme weights on the evaluation.
[0217] The cutoff threshold is used to limit the range of importance weights to reduce the impact of extreme weights on the evaluation.
[0218] In this step, by limiting extreme weights, the impact of extreme values in importance sampling on model training is reduced, thereby improving the stability of training; truncation processing can reduce weight variance and avoid model training instability caused by extreme weights of a few samples.
[0219] C400 adjusts the weight of the sample data set based on the importance weight, and trains and optimizes the preset sampling state classification model based on the sample data set after the weight adjustment, the preset offline evaluation method, the confidence interval estimation strategy, the preset optimization strategy, the offline-online hybrid evaluation strategy, and the model integration strategy.
[0220] Specifically, the weight of the sample data set can be adjusted according to the truncated importance weight, and the adjusted sample weight is used to train the model; the adjusted sample data is evaluated using a preset offline evaluation method; the confidence interval estimation strategy is used to evaluate the prediction accuracy of the model to ensure the reliability of the model; optimization strategies (such as gradient descent, policy gradient) are applied to train and optimize the model; offline and online data are combined for evaluation to ensure the performance of the model in a real environment; the prediction results of multiple models are integrated to improve the overall prediction effect and stability.
[0221] In this step, the performance and reliability of the model are optimized through sample weight adjustment, offline evaluation, confidence interval estimation, etc. The offline-online hybrid evaluation strategy and model integration strategy can comprehensively consider multiple aspects of the model and improve the overall performance of the model.
[0222] The method disclosed in this embodiment can effectively improve the training effect of the model through importance weighting and truncation processing, so that it has higher accuracy in practical applications; in this embodiment, measures such as data preprocessing, importance sampling and weight adjustment enhance the stability and reliability of the model and reduce the impact of extreme values on training; by optimizing the reward function and combining expert knowledge, the model is more in line with the real scene in practical applications, and the application effect of the model is improved; the offline reinforcement learning method enables the model to be optimized through historical data, improves the level of intelligence, and adapts to different operating environments and needs; the offline-online hybrid evaluation strategy and model integration strategy ensure the performance and stability of the model in different environments, and improve the overall performance. Through the implementation of these steps, the sampling state classification model can be effectively optimized to make it perform better and more reliable in practical applications.
[0223] The preset offline evaluation method includes one or more of a direct method and a counterfactual estimation method.
[0224] Specifically, the direct method includes: training a Q function approximator: Q θ (s,a)≈Q π(s, a); Use supervised learning to minimize the mean square error: L(θ) = E[(r+γ·max a′ Q θ (s′,a′)-Q θ (s,a)) 2 ].
[0225] By training a Q-function approximator to estimate the value function of the target policy, this approach directly uses processed samples to estimate the performance of the policy.
[0226] The counterfactual estimation method includes: using a doubly robust estimator: DR = DM + w(s,a)·(r + γV(s′)-Q(s,a)). Where V(s) = E a [Q(s,a)] combines the advantages of behavioral strategy and target strategy.
[0227] By combining the information of both the behavioral policy and the target policy using a doubly robust estimator, this approach provides more accurate estimates of policy value in offline data evaluation and takes into account the distribution differences between policies.
[0228] Confidence interval estimation includes: 1) using the Bootstrap method to generate multiple data subsets from the data; 2) independently evaluating the strategy on each subset; 3) calculating the confidence interval of the evaluation result. The Bootstrap method can provide the confidence interval of the strategy evaluation result, reflecting the uncertainty of the evaluation result.
[0229] Specifically, the Bootstrap method is used to generate multiple data subsets from the data, including: 1) randomly extracting several samples from the original data set to form multiple new subsets, and the samples of each subset are extracted from the original data with replacement, which means that the same sample may appear multiple times in a subset or may not appear; generating multiple subsets. This process will generate multiple (usually hundreds or thousands) of data subsets. Through these subsets, different possibilities of the data can be simulated, so as to have a more comprehensive understanding of the stability of the estimation results.
[0230] 2) Perform independent strategy evaluation on each subset, that is, for each generated subset, perform strategy evaluation separately and calculate the performance indicators (such as value function or return) of the target strategy. Specifically, it includes: training a strategy model on each subset and evaluating its performance. The result of the evaluation is an estimate of the performance of the target strategy on this specific subset. By independently evaluating each subset, a series of estimated values of strategy performance can be obtained, which reflect the performance of the strategy under different data subsets.
[0231] 3) Calculate the confidence interval of the evaluation results, that is, calculate the confidence interval by performing statistical analysis on the evaluation results obtained on each subset. Specifically, it includes: calculating the mean and standard error, that is, calculating the mean and standard error of the evaluation results on all subsets; determining the confidence interval range, that is, determining the confidence interval of the evaluation results based on these statistics. For example, a 95% confidence interval means that 95% of the time, the true strategy performance indicator will fall within this interval.
[0232] The confidence interval provides the range of the evaluation results and their uncertainty. This helps to understand the reliability of the strategy performance. If the confidence interval is narrow, it means that the evaluation results are relatively stable; if the confidence interval is wide, it means that the uncertainty of the results is high. The Bootstrap method helps estimate the distribution and stability of the strategy evaluation results by generating multiple data subsets and independently evaluating the strategy on these subsets. The confidence interval provides a credible range of the evaluation results and reflects the uncertainty of the results.
[0233] In this embodiment, the sample weights are adjusted to compensate for the distribution differences between strategies to make the evaluation more accurate; the performance of the target strategy is actually evaluated through offline strategy evaluation methods (such as direct methods and counterfactual estimation); the reliability of the evaluation results is quantified through confidence interval estimation; the strategy is improved based on the evaluation results to improve the actual performance of the strategy; and offline and online methods are combined for hybrid evaluation to verify and optimize the strategy and improve the performance of the strategy in the actual environment.
[0234] The preset optimization strategies include one or more of conservative policy iteration and batch constrained deep Q learning.
[0235] Conservative Policy Iteration (CPI) includes:
[0236] π new = arg max π E s [min(1+∈,w(s,a))·A πold (s,π(s))].
[0237] Among them, π new is the new strategy, that is, the strategy we hope to improve; ∈ is the radius of the trust region, which limits the extreme value of the weight and can avoid over-reliance on extreme samples, thereby achieving a more conservative strategy update. In the process of strategy improvement, the weight w(s, a) will be truncated within the range of 1+ε. πold (s, π(s)) is the advantage function under the old strategy πold, which is used to measure the quality of taking an action π(s) relative to other actions in a given state s.
[0238] Using batch constrained deep Q-learning (BCQ) involves training a generative model G(s) to imitate the behavioral policy; when the policy is improved, the actions are restricted to the output range of G(s) to ensure the effectiveness of the policy.
[0239] The offline-online hybrid evaluation strategy combines offline evaluation and limited online interaction, specifically including: designing a safe exploration mechanism to allow limited online interactions to verify offline evaluation results; using importance-weighted offline evaluation results to guide online exploration to reduce the risk of interacting with the environment; implementing an uncertainty-based active learning strategy to prioritize the exploration of state-action pairs with high uncertainty to enhance the learning ability and adaptability of the model.
[0240] The model integration strategy is used to improve the stability and accuracy of policy evaluation, including: training multiple independent Q function approximators and using integration methods to reduce the deviation of a single model; using model integration averaging to reduce the deviation of a single model and improve the stability and accuracy of policy evaluation; using integration variance as a measure of uncertainty to further improve the reliability of policy evaluation.
[0241] Through these detailed technical steps, the strategy for CO2 leak sampling status analysis can be effectively evaluated and improved in an offline environment, taking into account key factors such as data distribution bias, uncertainty, and safety constraints. This approach can significantly improve the reliability and performance of the system while minimizing the risks in actual operation.
[0242] Reference Figure 7 The method for obtaining the temperature and pressure change trend information specifically includes the following steps:
[0243] S410, determining a preset analysis model.
[0244] Specifically, you can choose a model suitable for processing time series data, such as long short-term memory network (LSTM), recurrent neural network (RNN), ARIMA (autoregressive integrated moving average model), etc.
[0245] This step includes model selection, initial setting of model parameters, and model structure design. Selecting a suitable analysis model can better handle the time series characteristics of temperature and pressure data and improve analysis accuracy; through reasonable parameter setting and structure design, the performance of the model in practical applications can be improved and more accurate trend prediction can be obtained.
[0246] Furthermore, the stationarity of the time series data can be checked in the preprocessing stage, using the Augmented Dickey-Fuller test (ADF). If the data is not stationary, difference processing is performed to make it stationary.
[0247] S420, training a preset analysis model based on historical time series data, analyzing the sampled data using the trained preset analysis model to obtain temperature and pressure change trend information;
[0248] The historical time series data includes historical temperature time series data and historical pressure time series data.
[0249] Specifically, the preprocessed historical time series data is input into a preset analysis model for training, and an appropriate optimization algorithm (such as Adam optimizer) and loss function (such as mean square error MSE) can be used for training.
[0250] For the prediction of temperature and pressure changes, the trained model is used to analyze the new sampling data to predict the temperature and pressure change trends. Example: Input the sampling data of the last week into the trained model to obtain the predicted temperature and pressure trends.
[0251] In this step, training based on historical data enables the model to learn real data patterns, thereby improving the accuracy of predictions; it can process the latest sampling data and provide real-time temperature and pressure change trends to support dynamic decision-making; it comprehensively considers the time series data of temperature and pressure to provide comprehensive trend information for subsequent analysis and decision-making.
[0252] If the preset analysis model is the ARIMA model, model identification includes: analyzing the autocorrelation function (ACF) and partial autocorrelation function (PACF) graphs; and determining the order (p, d, q) of the ARIMA model based on the characteristics of the ACF and PACF graphs.
[0253] Model fitting includes: fitting the ARIMA model using a certain order (p, d, q); estimating model parameters.
[0254] Model diagnosis includes: checking the model residuals to ensure that they have white noise characteristics; if the residuals do not meet the white noise characteristics, the model parameters need to be readjusted.
[0255] For forecasting, use the fitted model to make short-term predictions of future temperature and pressure; calculate confidence intervals for the forecasts.
[0256] Model updating can also be performed, specifically by regularly (such as daily or weekly) updating the model with new observational data; re-evaluating model parameters and adjusting the model structure if necessary.
[0257] The scheme disclosed in this embodiment can improve the accuracy of temperature and pressure change trend prediction by selecting and training appropriate analysis models, ensuring that the prediction results are in line with actual conditions; it can effectively identify and predict temperature and pressure change trends, help to provide early warning of potential equipment problems or environmental changes, and support maintenance and optimization decisions; make full use of historical time series data to obtain more reliable trend information through data-driven model training; according to different needs and data characteristics, the parameters and structure of the model can be adjusted to improve the adaptability and flexibility of the analysis; it can analyze and predict real-time data, support dynamic monitoring and real-time decision-making, and improve the system's responsiveness and operational efficiency.
[0258] Furthermore, the intelligent monitoring method for CO2 leakage sampling disclosed in the present application also includes strategy optimization, using the optimized strategy to train the preset sampling state classification model, analyzing the sampling data and actual sampling depth based on the trained model to obtain the actual sampling state.
[0259] Specifically, policy optimization includes one or more of conservative policy optimization, batch constrained deep Q learning, conservative Q learning, model-based counterfactual policy optimization, and model integration policy optimization.
[0260] Specifically, Conservative Policy Optimization (CPO) includes: a. Objective function design, that is, maximizing the expected return while ensuring that the policy changes are within a controllable range.
[0261] Among them, the objective function is: θ is the policy parameter, γ is the discount factor, and r t is the reward and KL is the Kullback-Leibler divergence.
[0262] b. Constraints, that is, adding safety constraints to ensure that the new strategy does not cause the system to enter a dangerous state.
[0263] The formula is: E[g(s,a)]≤ξ, g(s,a) is the security measurement function, and ξ is the security threshold.
[0264] c. Optimization algorithm, that is, using the Trust Region Policy Optimization (TRPO) or Proximal Policy Optimization (PPO) algorithm.
[0265] The formula is: where δ is the maximum allowed value of the KL divergence.
[0266] Optimizations for batch-constrained deep Q-learning (BCQ), including:
[0267] a. Behavior cloning network, that is, training a generative model G(s) to imitate the behavior strategy in the dataset. The loss function is: Loss G =E[(aG(s)) 2 ].
[0268] bQ function approximation, that is, training two Q networks to reduce over-estimation bias.
[0269] The formula is: Among them, θ′ is the target network parameter.
[0270] c. Action selection, that is, during inference, n candidate actions are generated from G(s) and the action with the highest Q value is selected.
[0271] The formula is: * = arg max a {Q(s,a)|a∈{G(s)+∈,∈~N(0,Φ)},|∈|<ρ}. Where Φ is the noise covariance and ρ is the maximum disturbance range.
[0272] For conservative Q-learning (CQL) optimization, including:
[0273] aQ function regularization is to add a conservative regularization term based on the standard Bellman update.
[0274] Loss function: Loss CQL =Loss TD +α·(E s [log∑ a exp(Q(s,a))]-E s,a~D [Q(s,a)]), where α is the regularization coefficient and D is the offline dataset.
[0275] b. Adaptive α adjustment, that is, dynamically adjusting α to balance conservatism and optimization goals.
[0276] Formula: α=α+η·(τ-E s [log∑ a exp(Q(s,a))]+E s,a~D [Q(s,a)]). Where η is the learning rate and τ is the target value.
[0277] Model-based Counterfactual Policy Optimization (MBCPO) includes:
[0278] a. Dynamic model learning, that is, training a probabilistic dynamic model P(s'|s, a) to simulate the environment; loss function: LossP = E[(s′ - P(s,a)) 2 + λ KL ·KL(P(s′∣s,a) || P data (s′∣s,a))。
[0279] b. Counterfactual trajectory generation, i.e., using the learned dynamics model to generate counterfactual trajectories.
[0280] Formula: τ cf = {(s t , a t , r t , s t+1 ) | st+1 ∼ P(·|st, at), at ∼ π(·|st)}.
[0281] c. Policy optimization, i.e., jointly optimizing the policy based on real data and counterfactual data.
[0282] Formula: J(θ) = E D [R(τ)] + λ cf ·E τcf [R(τ cf )]; where R(τ) is the cumulative reward of the trajectory, and λ cf is the weight of counterfactual data.
[0283] For model ensemble policy optimization, it specifically includes:
[0284] a. Model ensemble, i.e., training multiple independent policy models and making the final decision using ensemble methods (such as voting or weighted average).
[0285] b. Bayesian optimization, i.e., using Bayesian optimization to adjust hyperparameters such as learning rate, regularization coefficient, etc.
[0286] Formula: θ* = arg max θ E[Performance|Data, θ].
[0287] c. Progressive training, i.e., starting from simple tasks, gradually increasing the difficulty, and optimizing the policy using the idea of curriculum learning.
[0288] d. Robustness enhancement, i.e., adding adversarial samples during training to improve the robustness of the policy.
[0289] Loss function formula: Loss robust = E[Loss(s, a)] + λ·max δ Loss(s + δ, a), where δ is the bounded adversarial perturbation.
[0290] By implementing these detailed strategy optimization techniques, the decision strategy for CO2 leak sampling state analysis can be effectively improved in an offline environment. These methods pay special attention to handling the peculiarities of offline data, such as distribution shift, limited exploration range, etc., while considering safety and robustness. By combining these techniques, the performance and reliability of the system can be significantly improved, providing smarter and more reliable decision support for CO2 leak monitoring.
[0291] Among them, conservative strategy optimization designs the objective function that takes safety constraints into consideration and uses the TRPO or PPO algorithm for optimization to ensure that strategy changes are within a controllable range.
[0292] Batch constrained deep Q learning introduces a behavior cloning network to imitate the behavior strategy in the dataset, uses a double Q network to reduce over-estimation bias, and selects actions through generative models and noise perturbations during inference.
[0293] Conservative Q-learning adds a conservative regularization term to the standard TD error and uses an adaptive adjustment mechanism to balance conservatism and optimization objectives.
[0294] Model-based counterfactual policy optimization learns a probabilistic dynamics model to simulate the environment, generate counterfactual trajectories, and jointly optimize the policy based on real data and counterfactual data.
[0295] Ensemble and Tune Improve decision stability using model ensembles, apply Bayesian optimization to tune hyperparameters, and implement progressive training and robustness enhancement techniques.
[0296] These techniques elaborate on how to optimize the decision-making strategy for CO2 leak sampling state analysis in an offline environment. They pay special attention to the peculiarities of offline data, such as distribution shift and limited exploration range, while also considering safety and robustness. This comprehensive approach can significantly improve the performance and reliability of the system.
[0297] In this application, real-time parameters such as temperature and pressure are input into the trained model, and the model outputs the current sampling status judgment and recommended actions.
[0298] The application also includes: online fine-tuning and adaptation, specifically, collecting new data and feedback in actual applications; regularly using new data to fine-tune the model to improve its adaptability.
[0299] Furthermore, the present application also includes performance evaluation and iteration, specifically, regularly evaluating the judgment accuracy and decision quality of the system; iteratively optimizing algorithms and models based on the evaluation results and expert feedback.
[0300] In this application, through offline learning and interpolation rewards, the system can better handle unseen scenarios; by making full use of historical data and reducing reliance on real-time trial and error, offline learning avoids the risks that may be brought by online learning and can handle complex nonlinear relationships and dynamic changes.
[0301] By applying offline reinforcement learning technology with interpolation rewards, the CO2 leak sampling system can analyze the temperature and pressure changes in the kettle more intelligently and accurately, and make more reliable sampling status judgments. This method combines data-driven learning capabilities and adaptability to unknown situations, and is expected to significantly improve the overall performance and reliability of the system.
[0302] The application discloses an intelligent monitoring method for CO2 leakage sampling, which details the entire process from data collection to model training, explains how to use interpolation technology to handle missing or uncertain state transitions and improve the generalization ability of the model, emphasizes the advantages of offline learning in terms of security and data efficiency, and describes how to apply the trained model to real-time status judgment and how to maintain the adaptability of the model through online fine-tuning.
[0303] In this application, uncertainty estimation was introduced to enhance the reliability of the model in complex environments. This approach combines advanced machine learning technology with the specific needs of CO2 leak sampling, and has the potential to significantly improve the intelligence level and decision-making ability of the system.
[0304] In the present application, the sampling strategy is dynamically adjusted based on the actual sampling state, the temperature and pressure change trend information, and the first preset strategy, or the corresponding warning information is issued based on the actual sampling state, the temperature and pressure change trend information, and the second preset strategy, specifically including:
[0305] 1) Develop a decision-making system based on fuzzy logic to achieve intelligent adjustment of sampling parameters;
[0306] 2) Establish a multi-objective optimization model to balance sampling efficiency, sample quality and safety;
[0307] 3) Design a hierarchical alarm mechanism, including sound and light alarms and remote notification functions.
[0308] The development of a decision-making system based on fuzzy logic includes: designing a fuzzy logic system to process temperature and pressure trend data and automatically adjust sampling parameters based on these data. Fuzzy logic systems use fuzzy rules and membership functions to handle uncertainty and fuzzy information.
[0309] Example: If the temperature is too high and the pressure is unstable, increase the sampling frequency, if the temperature is normal but the pressure is low, decrease the sampling frequency.
[0310] Membership functions include: Membership functions that define different temperature and pressure ranges, indicating how much they fall into the categories of "high", "medium", and "low".
[0311] The fuzzy logic system can handle the uncertainty and ambiguity in the data, making the decision more flexible and adaptable to the actual situation; it can automatically adjust the sampling parameters according to the real-time analysis results to improve the efficiency and quality of sampling.
[0312] For establishing a multi-objective optimization model, the objectives are defined including sampling efficiency, sample quality and safety. Sampling efficiency is to optimize the sampling speed and frequency, sample quality is to ensure sample representativeness and accuracy, and safety is to prevent potential safety hazards and risks during the sampling process.
[0313] Specifically, a mathematical optimization model is established to comprehensively consider these objectives and seek the best balance point; for example, methods such as linear programming, nonlinear programming or genetic algorithms can be used.
[0314] Example: An optimization model can balance the trade-offs between sampling frequency, sample size, and sampling time to ensure that sampling efficiency is improved while maintaining sample quality and safety.
[0315] In this step, multiple objectives can be comprehensively balanced to ensure that sampling efficiency is met without sacrificing sample quality or safety; by optimizing the model, the optimal sampling parameters can be obtained to improve the overall operation effect and safety.
[0316] The design of a hierarchical alarm mechanism specifically includes: 1) being able to set up hierarchical alarms; 2) setting up a remote notification function.
[0317] Among them, the graded alarms include minor alarms, moderate alarms and severe alarms. Minor alarms alert operators by flashing lights or slight sounds. Medium alarms issue more obvious sound and light alarms and display the alarm information on the control panel. Severe alarms emit high-decibel sounds, flashing red lights, and automatically send notifications to the mobile phones or computers of relevant personnel.
[0318] The notification methods of the remote notification function include: setting up SMS, email or app notifications to send alarm information to remote devices.
[0319] Example: When the temperature or pressure exceeds the set threshold, the system automatically sends an alarm SMS or email to the on-duty personnel or maintenance team.
[0320] Through this step, alarms and notifications can be issued in a timely manner, which helps to quickly respond to potential problems and reduce equipment failures or safety risks. The graded alarm mechanism can take different response measures according to the severity of the problem, improving the flexibility and effectiveness of the alarm system.
[0321] In this application, the AES-256 encryption algorithm is used to protect the transmitted and stored data, which can achieve data verification and error correction to ensure data integrity. A data backup and recovery mechanism can be established to prevent data loss in unexpected situations.
[0322] Furthermore, it is possible to optimize the intelligent analysis algorithm to improve the prediction and decision-making capabilities of the system; it is also possible to develop multi-parameter integrated monitoring technology to achieve real-time monitoring of more indicators such as CO2 concentration and pH value; at the same time, it is also possible to explore integration with other intelligent systems, such as geological models, risk assessment systems, etc., to achieve more comprehensive CO2 storage supervision. This intelligent monitoring method provides an innovative solution for CO2 leakage sampling, significantly improving the accuracy, reliability and efficiency of sampling. The application of this technology will provide strong support for the safety assessment and management of CO2 storage projects, and promote the further development of carbon capture and storage technology.
[0323] In the prior art, with the development of carbon capture and storage (CCS) technology, the use of carbon dioxide (CO2) to drive oil production and storage is becoming more and more common. Carbon capture and storage technology aims to reduce greenhouse gas emissions and slow down global warming by capturing carbon dioxide and storing it underground. However, in this process, carbon dioxide may accidentally leak into the surface or groundwater, which may cause harm to the environment and human health. In the method disclosed in this application, through real-time monitoring and intelligent control, high-fidelity CO2 leakage samples are ensured; the state in the kettle is monitored in real time, abnormal conditions are discovered in time, and sampling risks are reduced; the degree of automation is high, manual operations are reduced, and sampling efficiency is improved, providing more accurate and reliable basic data for CO2 leakage assessment.
[0324] The second aspect of the present application discloses an intelligent monitoring system for CO2 leakage sampling, comprising:
[0325] The sampling module is used to obtain sampling data in real time, and the sampling data includes the actual temperature data and the actual pressure data inside the sampling kettle and the ambient temperature of the sampling kettle;
[0326] A depth acquisition module is used to obtain the actual sampling depth according to the preset strategy and sampling data;
[0327] A state acquisition module is used to analyze the sampling data and the actual sampling depth based on a preset sampling state classification model to obtain the actual sampling state;
[0328] A temperature and pressure change trend acquisition module is used to analyze the sampled data based on a preset analysis model to obtain temperature and pressure change trend information;
[0329] The analysis module is used to dynamically adjust the sampling strategy based on the actual sampling state, the temperature and pressure change trend information, and the first preset strategy, or to issue corresponding warning information based on the actual sampling state, the temperature and pressure change trend information, and the second preset strategy.
[0330] The computer device according to the embodiment of the present disclosure includes a memory and a processor. The memory is used to store non-temporary computer-readable instructions. Specifically, the memory may include one or more computer program products, and the computer program product may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may, for example, include a random access memory (RAM) and / or a cache memory (cache), etc. The non-volatile memory may, for example, include a read-only memory (ROM), a hard disk, a flash memory, etc.
[0331] The processor may be a central processing unit (CPU) or other forms of processing units having data processing capabilities and / or instruction execution capabilities, and may control other components in the computer device to perform desired functions. In one embodiment of the present disclosure, the processor is used to run the computer-readable instructions stored in the memory, so that the computer device performs all or part of the steps of the intelligent monitoring method for CO2 leakage sampling of the aforementioned embodiments of the present disclosure.
[0332] Those skilled in the art should be able to understand that in order to solve the technical problem of how to obtain a good user experience, the present embodiment may also include well-known structures such as a communication bus and an interface, and these well-known structures should also be included in the protection scope of the present disclosure.
[0333] like Figure 8 A schematic diagram of the structure of a computer device provided in an embodiment of the present disclosure is shown, which is a schematic diagram of the structure of a computer device suitable for implementing the embodiment of the present disclosure. Figure 8 The computer device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.
[0334] like Figure 8 As shown, the computer device may include a processor (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) or a program loaded from a storage device into a random access memory (RAM). In the RAM, various programs and data required for the operation of the computer device are also stored. The processor, ROM, and RAM are connected to each other via a bus. An input / output (I / O) interface is also connected to the bus.
[0335] Typically, the following devices can be connected to the I / O interface: input devices such as sensors or visual information acquisition devices; output devices such as display screens; storage devices such as tapes, hard disks, etc.; and communication devices. The communication device can allow the computer device to communicate with other devices (such as edge computing devices) wirelessly or by wire to exchange data. Figure 8 A computer device having various devices is shown, but it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed instead.
[0336] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program contains program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device, or installed from a storage device, or installed from a ROM. When the computer program is executed by a processor, all or part of the steps of the intelligent monitoring method for CO2 leak sampling of an embodiment of the present disclosure are executed.
[0337] For detailed description of this embodiment, reference may be made to the corresponding descriptions in the aforementioned embodiments, which will not be repeated here.
[0338] According to the computer-readable storage medium of the embodiment of the present disclosure, non-transitory computer-readable instructions are stored thereon. When the non-transitory computer-readable instructions are executed by a processor, all or part of the steps of the intelligent monitoring method for CO2 leakage sampling of the above-mentioned embodiments of the present disclosure are executed.
[0339] The above-mentioned computer-readable storage media include, but are not limited to: optical storage media (e.g., CD-ROM and DVD), magneto-optical storage media (e.g., MO), magnetic storage media (e.g., magnetic tape or mobile hard disk), media with built-in rewritable non-volatile memory (e.g., memory card) and media with built-in ROM (e.g., ROM box).
[0340] For detailed description of this embodiment, reference may be made to the corresponding descriptions in the aforementioned embodiments, which will not be repeated here.
[0341] The basic principles of the present disclosure are described above in conjunction with specific embodiments. However, it should be noted that the advantages, strengths, effects, etc. mentioned in the present disclosure are only examples and not limitations, and it cannot be considered that these advantages, strengths, effects, etc. are required by each embodiment of the present disclosure. In addition, the specific details disclosed above are only for the purpose of illustration and ease of understanding, and are not limitations. The above details do not limit the present disclosure to the necessity of adopting the above specific details to be implemented.
[0342] In the present disclosure, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply that there is any such actual relationship or order between these entities or operations. The block diagrams of the devices, devices, equipment, and systems involved in the present disclosure are only illustrative examples and are not intended to require or imply that they must be connected, arranged, and configured in the manner shown in the block diagram. As will be appreciated by those skilled in the art, these devices, devices, equipment, and systems can be connected, arranged, and configured in any manner. Words such as "including", "comprising", "having", etc. are open words, referring to "including but not limited to", and can be used interchangeably with them. The words "or" and "and" used here refer to the words "and / or" and can be used interchangeably with them, unless the context clearly indicates otherwise. The words "such as" used here refer to the phrase "such as but not limited to", and can be used interchangeably with them.
[0343] Additionally, as used herein, "or" used in a list of items beginning with "at least one" indicates a separate list, so that, for example, a list of "at least one of A, B, or C" means A or B or C, or AB or AC or BC, or ABC (i.e., A and B and C). Furthermore, the word "exemplary" does not mean that the example described is preferred or better than other examples.
[0344] It should also be noted that in the system and method of the present disclosure, each component or each step can be decomposed and / or recombined. Such decomposition and / or recombination should be regarded as equivalent solutions of the present disclosure.
[0345] Various changes, substitutions, and modifications of the techniques described herein may be made without departing from the teachings defined by the appended claims. Furthermore, the scope of the claims of the present disclosure is not limited to the specific aspects of the processes, machines, manufactures, compositions of events, means, methods, and actions described above. Currently existing or later to be developed processes, machines, manufactures, compositions of events, means, methods, or actions that perform substantially the same functions or achieve substantially the same results as the corresponding aspects described herein may be utilized. Thus, the appended claims include such processes, machines, manufactures, compositions of events, means, methods, or actions within their scope.
[0346] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use the present disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of the present disclosure. Therefore, the present disclosure is not intended to be limited to the aspects shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.
[0347] The above description has been given for the purpose of illustration and description. In addition, this description is not intended to limit the embodiments of the present disclosure to the forms disclosed herein. Although multiple example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, changes, additions and sub-combinations thereof.
Claims
1. An intelligent monitoring method for CO2 leakage sampling, characterized in that: include: Acquire sampling data in real time, wherein the sampling data includes actual temperature data and actual pressure data inside the sampling kettle and the ambient temperature of the sampling kettle; Acquire the actual sampling depth according to the preset strategy and the sampling data; Analyze the sampling data and the actual sampling depth based on a preset sampling state classification model to obtain an actual sampling state; Analyze the sampled data based on a preset analysis model to obtain temperature and pressure change trend information; The sampling strategy is dynamically adjusted based on the actual sampling state, the temperature and pressure change trend information, and the first preset strategy, and / or corresponding warning information is issued based on the actual sampling state, the temperature and pressure change trend information, and the second preset strategy.
2. The intelligent monitoring method for CO2 leakage sampling according to claim 1 is characterized in that: The obtaining of the actual sampling depth according to the preset strategy and the sampling data includes: Obtaining a preset lowering length of the steel wire rope connected to the sampling kettle at the initial temperature through an encoder or a preset length sensor; The ambient temperature of the sampling kettle in different areas is obtained based on the temperature sensor arranged on the lowered steel wire rope; Based on the ambient temperature of each area segment and a preset temperature correction formula, a theoretical length of each area segment is obtained; Superimposing the theoretical lengths of all the area segments to obtain the actual sampling depth; The preset temperature correction formula is: L corrected =L measured ×(1+αΔT), where L corrected is the theoretical length of each zone segment, L measured is the preset lowering length, α is the linear expansion coefficient of the wire rope, and ΔT is the temperature change between the ambient temperature of each section and the initial temperature.
3. The intelligent monitoring method for CO2 leakage sampling according to claim 2 is characterized in that: The analyzing the sampling data and the actual sampling depth based on the preset sampling state classification model to obtain the actual sampling state includes: Determine the initial classification model; Preprocessing the historical sampling data, training the initial classification model based on the preprocessed historical sampling data and a preset strategy, and using the initial classification model that meets the training requirements as the preset sampling state classification model; Analyze the sampling data and the actual sampling depth based on the preset sampling state classification model to obtain the actual sampling state; The historical sampling data includes several groups of historical parameters and corresponding status information, and the historical parameters include associated temperature data, pressure data and sampling depth data; The actual sampling status includes any one of normal, warning, serious, and dangerous.
4. The intelligent monitoring method for CO2 leakage sampling according to claim 3 is characterized in that: The preset strategy is a cross-validation strategy; The training of the initial classification model based on the preprocessed historical sampling data and a preset strategy to obtain a preset sampling state classification model includes: Dividing the preprocessed historical sampling data into a training set and a test set; Optimize the target hyperparameters using a cross-validation method, train the initial classification model based on the optimized target hyperparameters, the training set, and the test set, and use the initial classification model that meets the conditions as the preset sampling state classification model; The target hyperparameters include penalty parameters and kernel function parameters.
5. The intelligent monitoring method for CO2 leakage sampling according to claim 3 is characterized in that: The preset strategy is an interpolation reward strategy; The training of the initial classification model based on the preprocessed historical sampling data and a preset strategy to obtain a preset sampling state classification model includes: Determine the state space, action space, and target reward function of the configuration; Training the initial classification model based on the state space, the action space, and the target reward function, and using the initial classification model that meets the training requirements as the preset sampling state classification model; The state space is s, s=(T, P, D), T is the historical temperature data, P is the historical pressure data, and D is the historical sampling depth; The action space is a, a∈{all types of historical execution actions}, and the historical execution actions include continuing normal sampling, adjusting parameters, or stopping sampling; The target reward function is: R final =R total S(s,a), R total =α·R base +β·R interp −P uncert ; Where S(s,a) is the safety constraint function, R base As a basic reward, R interp is the interpolation reward, α is the weight of the base reward, β is the weight of the interpolation reward, P uncert is the uncertainty penalty term.
6. The intelligent monitoring method for CO2 leakage sampling according to claim 5 is characterized in that: R base =w T ·f(ΔT)+w P ·f(ΔP)+w Q f(Q); where ΔT is the temperature deviation, ΔP is the pressure deviation, Q is the sampling quality, f(ΔT) is the scoring function for the absolute value or square value of the temperature deviation (absolute value function or square function), f(ΔP) is the scoring function for the absolute value or square value of the pressure deviation, f(Q) is the scoring function for the absolute value or square value of the sampling quality, w T is the weight of temperature deviation, w P is the weight of pressure deviation, w Q is the weight of sampling quality; R interp =GPR(s,a), GPR(s,a) is the interpolation reward value obtained based on the Gaussian process regression model; P uncert =λ·σ 2 (s,a), where σ 2 (s,a) is the variance of the GPR model’s prediction uncertainty for states and actions, and λ is the penalty coefficient.
7. The intelligent monitoring method for CO2 leakage sampling according to claim 6 is characterized in that: It also includes: optimizing the target reward function to obtain a data-driven reward function; training the initial classification model based on the state space, the action space, and the data-driven reward function, and using the initial classification model that meets the training requirements as the preset sampling state classification model; The optimizing the target reward function to obtain the data-driven reward function includes: combining the knowledge of experts in the corresponding field and converting the expert rules into expert knowledge numerical rewards through a fuzzy logic system; Determining a data-driven reward function based on the expert knowledge numerical reward and the target reward function; The data-driven reward function includes: R = γ·R final +(1−γ)·R expert ; Among them, R expert is the expert knowledge numerical reward, γ is the weight coefficient, 0≤γ≤1.
8. The intelligent monitoring method for CO2 leakage sampling according to claim 7, characterized in that: Also includes: The preset sampling state classification model is trained and optimized based on an offline reinforcement learning strategy, and the sampling data and the actual sampling depth are analyzed based on the preset sampling state classification model after training and optimization to obtain the actual sampling state; The training and optimization of the preset sampling state classification model based on the offline reinforcement learning strategy includes: Performing initial preprocessing on historical data to obtain a sample data set; the initial preprocessing includes time series alignment, standardization, and data enhancement; Obtain importance weights based on the determined behavior strategy probability density and target strategy probability density; The importance weight is : , is the behavior strategy probability density, is the target strategy probability density; determining a truncated importance sampling based on the importance weight; The truncated importance sampling is : , c is the cutoff threshold; The weight of the sample data set is adjusted based on the importance weight, and the preset sampling state classification model is trained and optimized based on the sample data set after the weight adjustment, a preset offline evaluation method, a confidence interval estimation strategy, a preset optimization strategy, an offline-online hybrid evaluation strategy, and a model integration strategy.
9. The intelligent monitoring method for CO2 leakage sampling according to claim 8, characterized in that: The preset offline evaluation method includes one or more of a direct method and a counterfactual estimation method; The preset optimization strategy includes one or more of conservative strategy iteration and batch constrained deep Q learning.
10. The intelligent monitoring method for CO2 leakage sampling according to claim 1, characterized in that: The analyzing the sampled data based on a preset analysis model to obtain temperature and pressure change trend information includes: Determine the preset analysis model; The preset analysis model is trained based on historical time series data, and the sampled data is analyzed by the trained preset analysis model to obtain temperature and pressure change trend information; The historical time series data includes historical temperature time series data and historical pressure time series data.
Citation Information
Patent Citations
Method for determining sampling scheme, semiconductor substrate measurement apparatus, lithographic apparatus
CN113906347A
Method and system for monitoring leakage of carbon dioxide underground sealed gas
CN114994243A