Intelligent Operation and Maintenance System and Device for Gathering and Transportation Based on Missing Data Filling Method
Through the data acquisition, grouping, real-time prediction and missing value supplement modules in the intelligent operation and maintenance system of the integrated transportation intelligent operation and maintenance system, the LSTM model and isolated forest algorithm are used to solve the problem of data missing in the oilfield integrated transportation system, real-time accurate filling and abnormal detection of data are realized, and operation and maintenance efficiency and data integrity are improved.
Patent Information
- Application Number
- CN202510271975.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-03-10
AI Technical Summary
The lack of data during the data collection process in the oilfield collection and transportation system leads to monitoring blind spots, affecting equipment health management and process optimization, increasing production safety risks and operating costs. The traditional data filling method has large errors and cannot adapt to changes in external factors in real time.
A smart operation and maintenance system based on missing data filling method is adopted, including data acquisition, grouping, real-time prediction and missing value supplement modules. Data prediction and abnormal detection are performed through the LSTM model and the isolated forest algorithm, and comprehensive calculation is performed with pipeline data in the same group to update and fill missing data in real time.
It improves the accuracy and reliability of data prediction, reduces errors, enhances the adaptability and functionality of the system, ensures data integrity and continuous operation of the system, and supports operation and maintenance decision-making.
Smart Images

Figure CN119784365B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of intelligent operation and maintenance of oilfield gathering and transportation systems, and specifically relates to a gathering and transportation intelligent operation and maintenance system and device based on a missing data filling method. Background Art
[0002] In an oilfield gathering and transportation system, monitoring and management rely on a large number of sensors and data acquisition devices, which monitor temperature, pressure, and flow parameters during the production process in real time. However, complex operating environments and technical limitations can lead to data loss during the data acquisition process. Data loss can result in monitoring blind spots, affecting equipment health management and process optimization, and increasing production safety risks and operating costs;
[0003] Traditional oilfield gathering and transportation data interpolation methods mostly establish prediction equations based on the operating historical data of oilfield pipelines to predict missing data. However, due to different external factors during the daily pipeline operation process, it is easy to have large errors in filling missing pipeline data during the gathering and transportation process only through historical data, and it is impossible to group pipelines and perform real-time prediction according to the actual pipeline specifications during daily operation, resulting in problems of low practicability and functionality;
[0004] This case proposes a gathering and transportation intelligent operation and maintenance system and device based on a missing data filling method to solve the above technical problems. Summary of the Invention
[0005] The present invention aims to solve at least one of the technical problems existing in the prior art. For this purpose, the present invention proposes a gathering and transportation intelligent operation and maintenance system and device based on a missing data filling method, and solves the above technical problems by improving the detection method and processing method.
[0006] To achieve the above object, the present invention adopts the following technical solutions:
[0007] A gathering and transportation intelligent operation and maintenance system based on a missing data filling method, comprising a data acquisition module, a gathering and transportation grouping module, a real-time prediction module, and a missing value supplement module;
[0008] The data acquisition module includes a pressure sensor, a temperature sensor, a flow sensor, and a water cut sensor, which are used to collect real-time pressure data, temperature data, flow data, and crude oil water cut data of the transportation pipeline during the oilfield gathering and transportation process, and at the same time collect oilfield gathering and transportation reference data, including pipeline specifications and start-stop states, and transmit them to the subsequent modules;
[0009] The gathering and transportation grouping module groups the daily operating pipelines based on the collected current oilfield gathering and transportation pipeline specifications and start-stop status, establishes a daily operating pipeline database based on MySQL, groups the daily enabled pipelines based on pipeline specifications, divides the same-type daily operating pipelines into the same file, and records the real-time gathering and transportation data collected by the data acquisition module;
[0010] The real-time prediction module, for the pipelines in the same file, respectively establishes the prediction models for the daily pressure, temperature, flow rate, and water cut data of each pipeline based on the real-time gathering and transportation data of each pipeline in the file, and synchronously performs stage verification and update on the prediction models based on the real-time recorded pipeline data;
[0011] The missing value filling module, based on the pipeline data prediction models in the same file, comprehensively fills the missing data of the pipelines with missing data by combining the output values of the prediction models with missing data and the output values of the prediction models of other pipelines in the same file.
[0012] Further, the real-time prediction module, for the pipelines in the same file, respectively establishes the prediction models for the daily pressure, temperature, flow rate, and water cut data of each pipeline based on the real-time gathering and transportation data of each pipeline in the file, and synchronously performs stage verification and update on the prediction models based on the real-time recorded pipeline data. The specific steps are as follows:
[0013] For the pipeline data in the same grouped file in the daily operating pipeline database, respectively establish the prediction models for pressure, temperature, flow rate, and water cut data based on the real-time data of each pipeline during daily operation;
[0014] Collect the real-time monitoring data of different pipelines in each file at time intervals of m, and perform stage verification and update on the prediction models established on the same day.
[0015] Further, for the pipeline data in the same grouped file in the daily operating pipeline database, respectively establish the prediction models for pressure, temperature, flow rate, and water cut data based on the real-time data of each pipeline during daily operation. The specific steps are as follows:
[0016] For the pipeline data in the same grouped file, attach timestamps to the pressure, temperature, flow rate, and water cut data of each pipeline. Based on each pipeline in the same grouped file, respectively collect the pressure, temperature, flow rate, and water cut data at the first T moments after the pipeline runs on the same day. Divide the T moments into N equal parts, and respectively establish the prediction models for the pressure, temperature, flow rate, and water cut data of different pipelines in the same grouped file based on the LSTM model. The specific steps are as follows:
[0017] Collect the data at the first T moments after the pipeline runs on the same day, and record it as {X t}, where {Xt}={P t ,A t ,F t ,W t}, where P t , A t 、F t , W t Respectively represent the pressure, temperature, flow rate and water content data of the current pipeline;
[0018] Divide the data at time T into N equal parts, each containing the time step Label each piece of data as
[0019] Each piece of data As input, each time step The input dimension of the model is set to 4, the data set is divided into training set, validation set and test set, and the mean square error is used as the loss function. The algorithm formula is:
[0020]
[0021] Among them, y t , They represent the true value and the predicted value respectively. The Adam optimizer is used to further update the model parameters to obtain the pressure, temperature, flow rate, and water content data prediction model for different pipelines in the same group.
[0022] Furthermore, the real-time monitoring data of different pipelines in each archive is collected based on the time interval m, and the prediction model established on that day is verified and updated in stages. The specific steps are:
[0023] Based on the time interval m, the time point t is obtained respectively ma , at each time point t ma The data in different pipeline files in the same group are extracted, including the actual values of pipeline pressure, pipeline temperature, pipeline flow and water content data.
[0024] Set the time point t ma Input into the LSTM model trained on the corresponding pipeline on that day to obtain the time point t ma Current pipeline data prediction value Calculate the actual value based on the mean absolute error method With the predicted value The error of is, and its algorithm formula is:
[0025]
[0026] Among them, α ma Represents the time point tma The mean absolute error value between the actual value and the predicted value at a certain time, and based on the mean absolute error threshold θ for α ma Make a determination:
[0027] When α ma ≥ θ, it means that the current pipeline data prediction model needs to be updated;
[0028] When α ma < θ, it means that the current pipeline data prediction model does not need to be updated and continues to be used;
[0029] For α ma ≥ θ, collect the new data since the time t ma Divide it equally based on N equal parts again, and retrain the LSTM model using the updated dataset again, and replace the old model with the new prediction model as the data prediction model of the current pipeline.
[0030] Furthermore, a same-day anomaly detection module is also set in the real-time prediction module. Based on the pipeline same-type data within the same group, real-time anomaly detection is performed at every 2T moment. The specific steps are as follows:
[0031] Collect the real-time pressure, temperature, flow rate, and water cut data per minute before 2T of the current day from all pipeline files within the same group, and divide the original dataset based on the data type respectively. Analyze the anomaly parameters for each dataset through the Isolation Forest method. The specific steps are as follows:
[0032] Randomly extract a sample subset from the original dataset, with a size of ψ. For each tree, starting from the root node, randomly select a feature and a random split value on this feature, and divide the dataset into left and right child nodes. Recursively repeat the above process for each child node until reaching the preset tree depth or there is only one data point in the child node;
[0033] For each data point x, calculate its path length h(x) in each tree, and average the path lengths of all trees to obtain the average path length Calculate the anomaly score s(x,u), and its algorithm formula is:
[0034]
[0035] Among them, u represents the total number of samples, that is, the number of all data points participating in the anomaly detection, c(u) is the sample constant, H(i) represents the i-th harmonic number. When s(x,u) is greater than the detection threshold, it means that there is an anomaly in the pipeline data points collected on the current day. Mark the abnormal pipelines in the same group file, and isolate the LSTM models of the abnormal data types of the abnormal pipelines separately, and do not transmit the output value to the subsequent modules.
[0036] Further, the missing value supplement module predicts models based on pipeline data in the same file, and fills the missing data of the pipeline with missing data by combining the predicted model output value of the pipeline with missing data and the predicted model output values of other pipelines in the same file. The specific steps are as follows:
[0037] Collect the real-time monitoring data of all pipelines in the same file, and determine the data missing time point t γ And the pipeline with missing data. Based on the type of missing data, through the LSTM model of the data related to the type of the pipeline with missing data, by inputting the missing time point t γ Obtain the predicted value of the missing pipeline data, and combine the predicted model output values of other pipelines in the same file at the missing time point t γ To comprehensively generate the missing data value of the pipeline with missing data.
[0038] Further, the step of collecting the real-time monitoring data of all pipelines in the same file, determining the data missing time point t γ And the pipeline with missing data. Based on the type of missing data, through the LSTM model of the data related to the type of the pipeline with missing data, by inputting the missing time point t γ Obtain the predicted value of the missing pipeline data, and combine the predicted model output values of other pipelines in the same file at the missing time point t γ To comprehensively generate the missing data value of the pipeline, the specific steps are as follows:
[0039] Substitute the data missing time point t γ Into the LSTM prediction model of the relevant type data in the pipeline with missing data to obtain the predicted value of the missing pipeline data At the same time, substitute the missing time point t γ Into the LSTM prediction model of the relevant type data in other pipelines in the same grouped file to obtain Where G is the pipeline other than the pipeline with missing data in the current group;
[0040] Set the weight value ω g , ω1 + ω2 +... + ω g = 1, where g is the number of all pipelines in the current group. Calculate the missing data value of the pipeline based on the weight value. The specific algorithm formula is:
[0041]
[0042] Among them, Is the missing data value of the pipeline.
[0043] Furthermore, the gathering and transportation grouping module groups the daily operating pipelines based on the collected current oilfield gathering and transportation pipeline specifications and start-stop status, establishes a daily operating pipeline database based on MySQL, groups the daily enabled pipelines based on pipeline specifications, divides the same-type daily operating pipelines into the same file, and records the real-time gathering and transportation data collected by the data acquisition module. The specific steps are as follows:
[0044] Based on the collected current oilfield gathering and transportation pipeline specifications and the daily oilfield gathering and transportation pipeline task list, establish a file for the enabled oilfield gathering and transportation pipelines, and store the pipeline specifications in the file, including pipeline diameter and pipeline material;
[0045] For the pipelines operating on the same day, group the pipelines with the same diameter and material, establish a daily operating pipeline database based on MySQL, and put the pipeline files grouped into the same group in the same group in the daily operating pipeline database;
[0046] Receive the real-time pipeline data collected by the sensors in the data acquisition module and record it in the file of the corresponding pipeline.
[0047] The gathering and transportation intelligent operation and maintenance device based on the missing data filling method is applied to the gathering and transportation intelligent operation and maintenance system based on the missing data filling method, including a memory, a processor, and a program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps in the gathering and transportation intelligent operation and maintenance system based on the missing data filling method of the present invention.
[0048] Compared with the prior art, the beneficial effects of the present invention are:
[0049] 1. In the present invention, based on the gathering and transportation plan, the daily operating pipelines are grouped based on the specification model, and daily gathering and transportation data prediction equations are established for all pipelines in the same group respectively, avoiding the influence of external factors on pipeline data prediction, and enhancing the accuracy of pipeline data prediction;
[0050] 2. In the present invention, by periodically updating the already constructed data prediction model based on a fixed time interval, the prediction model can adapt to changes in a timely manner, continuously and accurately reflect the latest characteristics and laws of the data, maintain good prediction performance, enhance the accuracy of subsequent missing data filling, and correct errors to make the model prediction results closer to the true values;
[0051] 3. In the present invention, by combining the prediction results of the missing data at the missing time points of other prediction equations in the same group, and combining the predicted values of the missing data of the pipelines with missing data, the missing data values of the pipelines with missing data are obtained through comprehensive calculation, realizing the filling of the missing data of the pipelines with missing data. Combining multiple prediction results can reduce the sensitivity of the model to extreme values or abnormal inputs and improve the robustness of the prediction results;
[0052] 4. In the present invention, by grouping the pipelines enabled on the same day based on the specification type, and comprehensively calculating by combining the results of other prediction equations in the same group and the prediction value of the pipeline with missing data itself, multiple information sources can be fully utilized, avoiding the limitations and biases of a single prediction model, making the filled value closer to the true value, thereby improving the reliability of the data;
[0053] 5. In the present invention, by combining the real-time data of each pipeline in the same group on the same day, a comprehensive determination of possible abnormal data values is made based on the data within the same group on the same day, avoiding the limitations of judging abnormal values through historical data. Each pipeline in the same group on the same day is in a similar operating environment and conditions, and there is a strong correlation and comparability between their data. By comprehensively considering these real-time data to determine abnormal values, more information can be used for judgment, thereby reducing the possibility of misjudgment;
[0054] 6. In the present invention, based on the real-time data of each pipeline in the same group on the same day to determine abnormal values, a rapid comprehensive determination can be made based on the real-time data, without waiting for the collection and analysis of a large amount of historical data, enabling faster decision-making and response, and enhancing the adaptability and functionality;
[0055] The entire intelligent operation and maintenance system and device for gathering and transportation based on the missing data filling method can intelligently manage the pipeline data in the oilfield gathering and transportation process through real-time data acquisition and the establishment of a prediction model on the same day, improving the operation and maintenance efficiency. Through the missing value supplement module, the system can effectively handle the problem of missing data, ensuring the integrity of the data and the continuous operation of the system. The real-time prediction module can dynamically update the prediction model, improving the accuracy and reliability of the prediction, and providing support for operation and maintenance decision-making. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] Figure 1 It is a block diagram of the intelligent operation and maintenance system for gathering and transportation based on the missing data filling method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0057] The technical solutions of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0058] Embodiment 1:
[0059] As Figure 1 shown, the intelligent operation and maintenance system for gathering and transportation based on the missing data filling method includes a data acquisition module, a gathering and transportation grouping module, a real-time prediction module, and a missing value supplement module;
[0060] A data acquisition module, including a pressure sensor, a temperature sensor, a flow sensor, and a water cut sensor, is used to collect real-time pressure data, temperature data, flow data, and crude oil water cut data of the transportation pipeline during the oilfield gathering and transportation process. At the same time, it collects the basic data of the oilfield gathering and transportation, including pipeline specifications and start-stop status, and transmits them to the subsequent modules;
[0061] A gathering and transportation grouping module groups the daily operating pipelines based on the collected current oilfield gathering and transportation pipeline specifications and start-stop status, establishes a daily operating pipeline database based on MySQL, groups the daily enabled pipelines based on pipeline specifications, divides the same-type daily operating pipelines into the same file, and records the real-time gathering and transportation data collected by the data acquisition module. The specific steps are as follows:
[0062] Based on the collected current oilfield gathering and transportation pipeline specifications and the daily oilfield gathering and transportation pipeline task list, establish a file for the enabled oilfield gathering and transportation pipelines, and store the pipeline specifications in the file, including pipeline diameter and pipeline material;
[0063] For the pipelines operating on the same day, group the pipelines with the same diameter and material, establish a daily operating pipeline database based on MySQL, and put the pipeline files grouped into the same group in the daily operating pipeline database;
[0064] Receive the real-time pipeline data collected by the sensors in the data acquisition module and record it in the corresponding pipeline file;
[0065] A real-time prediction module, for the pipelines in the same file, based on the real-time gathering and transportation data of each pipeline in the file, respectively establishes prediction models for the daily pressure, temperature, flow, and water cut data of each pipeline, and synchronously verifies and updates the prediction models based on the real-time recorded pipeline data. The specific steps are as follows:
[0066] For the pipeline data in the same group file in the daily operating pipeline database, based on the real-time data of each pipeline during the daily operation, respectively establish prediction models for pressure, temperature, flow, and water cut data. The specific steps are as follows:
[0067] For the pipeline data in the same group file, attach time stamps to the pressure, temperature, flow, and water cut data of each pipeline. Based on each pipeline in the same group file, respectively collect the pressure, temperature, flow, and water cut data at the first T moments after the daily pipeline operation. Divide the T moments into N equal parts, and respectively establish prediction models for the pressure, temperature, flow, and water cut data of different pipelines in the same group file based on the LSTM model. The specific steps are as follows:
[0068] Collect the data at the first T moments after the daily pipeline operation and record it as {Xt}, where {X t} = {P t , A t , F t , W t}, where P t , A t , F t , W t respectively represent the pressure, temperature, flow rate, and water cut data of the current pipeline;
[0069] Divide the data at time T into N equal parts respectively, and the time step included in each part is Mark each part of the data as
[0070] Take each part of the data as the input, where each time step Set the input dimension of the model to 4, divide the data set into a training set, a validation set, and a test set, and at the same time use the mean square error as the loss function. Its algorithm formula is:
[0071]
[0072] where, y t , respectively represent the true value and the predicted value. Use the Adam optimizer to further update the model parameters to obtain the prediction models for the pressure, temperature, flow rate, and water cut data of different pipelines within the same group.
[0073] It should be noted that the training set is used for model parameter learning, the validation set is used for adjusting hyperparameters, and the test set is used for evaluating model performance. The division ratio is 70% for the training set, 15% for the validation set, and 15% for the test set. The value of the first T moments is generally 120 minutes, and N equal parts are usually set to 120, that is, the real-time pressure data, temperature data, flow rate data, and water cut data of the pipeline obtained through the sensor are used as one data point per minute. The first T moments and N equal parts can also be adjusted according to the actual situation. Through the trained LSTM model, the predicted values of the pressure, temperature, flow rate, and water cut in the future period can be output. Based on the LSTM model, prediction models for the pressure, temperature, flow rate, and water cut data of different pipelines in the same group file are established respectively to realize real-time prediction and dynamic update of the data. At the same time, based on the daily pipeline operation conditions, the model of the current day is established, which can unify the external influences on the pipeline caused by other external factors, such as weather and humidity, to ensure the accuracy of the current data prediction.
[0074] Collect the real-time monitoring data of different pipelines in each file based on the time interval m, and conduct stage verification and update on the prediction model established on the current day. The specific steps are as follows:
[0075] Obtain time points t based on time interval m respectively ma , at each time point t ma Extract the data in different pipeline files within the same group, including the actual values of pipeline pressure, pipeline temperature, pipeline flow rate, and water content data
[0076] Input the time point t ma into the LSTM model trained for the corresponding pipeline on the same day, and obtain the predicted value of the current pipeline data at time point t ma Calculate the error between the actual value and the predicted value based on the mean absolute error method. Its algorithm formula is:
[0077]
[0078] where α ma represents the mean absolute error value between the actual value and the predicted value at time point t ma . Determine α ma based on the mean absolute error threshold θ:
[0079] When α ma ≥θ, it means that the prediction model of the current pipeline data needs to be updated;
[0080] When α ma <θ, it means that the prediction model of the current pipeline data does not need to be updated and continues to be used;
[0081] For α ma ≥θ, collect the new data since time t ma , re - evenly divide it based on N equal parts, and retrain the LSTM model using the updated dataset, and replace the old model with the new prediction model as the data prediction model of the current pipeline.
[0082] It should be noted that based on different data types, different mean absolute error thresholds θ need to be set. Calculate the statistics such as the mean and standard deviation of the daily data of pressure, flow rate, temperature, and water content respectively, and set the mean absolute error threshold θ through the empirical method. When the mean absolute error of the model for a certain variable exceeds its specific threshold, it triggers the phased update of the model related to that variable, so as to always maintain the prediction accuracy of the model. The time interval m needs to be set according to the actual usage situation, usually set to 20 minutes, and can also be increased or decreased as needed.
[0083] The real-time prediction module is also provided with a same-day anomaly detection module, which performs real-time anomaly detection at every 2T moment based on the same-type data of pipelines within the same group. The specific steps are as follows:
[0084] Collect the real-time pressure, temperature, flow rate, and water cut data per minute before 2T of the same day from all pipeline files within the same group, and divide the original data set based on the data type respectively. Analyze the anomaly parameters of each data set through the Isolation Forest method. The specific steps are as follows:
[0085] Randomly extract a sample subset from the original data set, with a size of ψ. For each tree, starting from the root node, randomly select a feature and a random splitting value on this feature, and divide the data set into left and right child nodes. Recursively repeat the above process for each child node until reaching the preset tree depth or there is only one data point in the child node;
[0086] For each data point x, calculate its path length h(x) in each tree, and average the path lengths of all trees to obtain the average path length Calculate the anomaly score s(x,u), and its algorithm formula is:
[0087]
[0088] Among them, u represents the total number of samples, that is, the number of all data points participating in the anomaly detection, c(u) is the sample constant, and H(i) represents the i-th harmonic number. When s(x,u) is greater than the detection threshold, it means that the pipeline data points collected on the same day are abnormal. Mark the abnormal pipelines in the same group file, and isolate the LSTM model of the abnormal data type of the abnormal pipelines separately, and do not transmit the output value to the subsequent modules.
[0089] It should be noted that in the Isolation Forest, the closer the anomaly score s(x,u) is to 1, the more likely the data point is an outlier; the closer the anomaly score s(x,u) is to 0, the more likely it is normal data. The detection threshold is generally set to 0.8, and can also be adjusted between 0.75 and 0.93 according to the actual situation. The T value in 2T needs to be consistent with the T value when constructing the LSTM model. Isolating the LSTM model of the abnormal data type of the abnormal pipelines separately and not transmitting the output value to the subsequent modules can reduce the error when performing missing value filling later.
[0090] Embodiment 2:
[0091] The missing value filling module, based on the pipeline data prediction model in the same file, comprehensively fills the missing data of the pipelines with missing data by combining the output value of the prediction model with missing data and the output values of the prediction models of other pipelines in the same file. The specific steps are as follows:
[0092] Collect the real-time monitoring data of all pipelines in the same file, and determine the data missing time point t γ and the pipelines with missing data. Based on the type of missing data, through the LSTM model of the relevant type data of the pipelines with missing data, by inputting the missing time point t γ obtain the predicted values of the missing pipeline data. Combine the predicted model output values of other pipelines in the same file at the missing time point t γ to comprehensively generate the missing data values of the pipelines with missing data. The specific steps are as follows:
[0093] Substitute the data missing time point t γ into the LSTM prediction model of the relevant type data in the pipelines with missing data to obtain the predicted values of the missing pipeline data At the same time, substitute the missing time point t γ into the LSTM prediction model of the relevant type data in other pipelines in the same grouped file to obtain where G is the pipelines other than the pipelines with missing data in the current group;
[0094] Set the weight value ω g , ω1 + ω2 +... + ω g = 1, where g is the number of all pipelines in the current group. Calculate the missing data values of the pipelines based on the weight values. The specific algorithm formula is:
[0095]
[0096] where, is the missing data value of the pipeline.
[0097] It should be noted that ω1 is generally set to 0.5, and the remaining weight values are the evenly divided values based on the number of pipelines other than the pipelines with missing data in the current group. By combining the prediction results of the missing data at the missing time point by other prediction equations in the same group, and combining the predicted values of the missing data of the pipelines with missing data, comprehensively calculate the missing data values of the pipelines to fill the missing data of the pipelines with missing data and ensure the integrity and accuracy of the data.
[0098] Example 3:
[0099] The gathering and transportation intelligent operation and maintenance device based on the missing data filling method is applied to the gathering and transportation intelligent operation and maintenance system based on the missing data filling method, including a memory, a processor, and a program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps in the gathering and transportation intelligent operation and maintenance system based on the missing data filling method of the present invention.
[0100] The intelligent operation and maintenance system and device for gathering and transportation based on the missing data filling method, when in use, through pressure sensors, temperature sensors, flow sensors, and water cut sensors, are used to collect real-time pressure data, temperature data, flow data, and crude oil water cut data during the oilfield gathering and transportation process. Based on the pipelines enabled for oilfield gathering and transportation every day, they are grouped according to the oilfield specifications, and a prediction model is established for the pipelines within the group in combination with real-time data;
[0101] The prediction model is updated stage by stage based on the actual measured values. At the same time, the prediction models of all pipelines within the same group are integrated, and the missing data is comprehensively calculated to obtain the missing data of the pipelines with missing data.
[0102] In the embodiments provided by the present invention, it should be understood that the disclosed equipment, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division, and there may be other division methods in actual implementation; the modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the method of this embodiment.
[0103] The above embodiments are only used to illustrate the technical method of the present invention and not to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical method of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical method of the present invention.
Claims
1. An intelligent operation and maintenance system for gathering and transportation based on a missing data filling method, characterized in that: It includes a data acquisition module, a gathering and transportation grouping module, a real-time prediction module, and a missing value supplementation module; The data acquisition module includes a pressure sensor, a temperature sensor, a flow sensor, and a water cut sensor, which are used to collect the real-time pressure data, temperature data, flow data, and crude oil water cut data of the transportation pipeline during the oilfield gathering and transportation process. At the same time, it collects the oilfield gathering and transportation reference data, including pipeline specifications and start-stop states, and transmits them to the subsequent modules; The gathering and transportation grouping module groups the daily operating pipelines based on the collected current oilfield gathering and transportation pipeline specifications and start-stop states, establishes a daily operating pipeline database based on MySQL, groups the daily enabled pipelines based on pipeline specifications, divides the same-type daily operating pipelines into the same file, and records the real-time gathering and transportation data collected by the data acquisition module; The real-time prediction module, for the pipelines in the same file, based on the real-time gathering and transportation data of each pipeline in the file, respectively establishes the prediction models for the daily pressure, temperature, flow, and water cut data of each pipeline, and based on the real-time recorded pipeline data, synchronously conducts stage verification and update of the prediction models; The missing value supplementation module, based on the pipeline data prediction models in the same file, comprehensively fills the missing data of the pipelines with missing data by combining the output values of the prediction models with missing data and the output values of the prediction models of other pipelines in the same file. The specific steps are as follows: Collect the real-time monitoring data of all pipelines in the same file and determine the data missing time point t γ And the pipelines with missing data. Based on the type of missing data, through the LSTM model of the data related to the pipelines with missing data, by inputting the missing time point t γ Obtain the predicted values of the missing pipeline data, and combine the output values of the prediction models of other pipelines in the same file at the missing time point t γ To comprehensively generate the missing data values of the pipelines with missing data, the specific steps are as follows: Substitute the data missing time point t γ into the LSTM prediction model of the relevant type of data in the data missing pipeline to obtain the predicted value of the missing pipeline data At the same time, substitute the missing time point t γ into the LSTM prediction model of the relevant type of data in other pipelines in the same grouped file to obtain where G is the pipeline other than the data missing pipeline in the current group; Set the weight value ω g , ω1 + ω2 +... + ω g = 1, where g is the number of all pipelines in the current group, and the pipeline missing data value is calculated based on the weight value. The specific algorithm formula is as follows: Among them, is the pipeline missing data value.
2. The intelligent operation and maintenance system for gathering and transportation according to the missing data filling method of claim 1, characterized in that: The real-time prediction module, for the pipelines in the same file, based on the real-time gathering and transportation data of each pipeline in the file, respectively establishes the prediction models for the daily pressure, temperature, flow, and water cut data of each pipeline, and based on the real-time recorded pipeline data, synchronously conducts stage verification and update of the prediction models. The specific steps are as follows: For the pipeline data in the same grouped file in the daily operating pipeline database, based on the real-time data of each pipeline during the daily operation, respectively establish the prediction models for pressure, temperature, flow, and water cut data; Collect the real-time monitoring data of different pipelines in each file at time interval m, and conduct stage verification and update of the prediction models established on the same day.
3. The intelligent operation and maintenance system for gathering and transportation according to the missing data filling method as claimed in claim 2, characterized in that: For the pipeline data in the same grouped file in the daily operating pipeline database, based on the real-time data of each pipeline during the daily operation, respectively establish the prediction models for pressure, temperature, flow, and water cut data. The specific steps are as follows: For the pipeline data in the same grouped file, attach time stamps to the pressure, temperature, flow, and water cut data of each pipeline. Based on each pipeline in the same grouped file, respectively collect the pressure, temperature, flow, and water cut data at the first T moments after the daily pipeline operation. Divide the T moments into N equal parts, and respectively establish the prediction models for the pressure, temperature, flow, and water cut data of different pipelines in the same grouped file based on the LSTM model. The specific steps are as follows: Collect the data of the pipeline operation at the previous T moment on the same day and record it as {X t}, where {X t} = {P t , A t , F t , W t}, where P t , A t , F t , W t represent the pressure, temperature, flow rate, and water content data of the current pipeline respectively; Divide the data at time T into N equal parts respectively, and the time steps included in each part are Mark each part of the data as Take each piece of data as input, where at each time step Set the input dimension of the model to 4, divide the dataset into training set, validation set and test set, and at the same time use mean squared error as the loss function. Its algorithm formula is: where y t and represent the true value and the predicted value respectively. The Adam optimizer is used to further update the model parameters to obtain the prediction models for the pressure, temperature, flow rate, and water cut data of different pipelines within the same group.
4. The intelligent operation and maintenance system for gathering and transportation according to the missing data filling method as claimed in claim 3, wherein: The step of collecting the real-time monitoring data of different pipelines in each file at time interval m and conducting stage verification and update of the prediction models established on the same day is as follows: Obtain time points t respectively based on time interval m ma , at each time point t ma extract data from different pipeline files within the same group, including the actual values of pipeline pressure, pipeline temperature, pipeline flow rate, and water content data Input the time point t ma into the LSTM model trained for the corresponding pipeline on the same day, and obtain the predicted value of the current pipeline data at the time point t ma Based on the mean absolute error method, calculate the error between the actual value and the predicted value , and its algorithm formula is: Among them, α ma represents the mean absolute error value between the actual value and the predicted value at time point t ma . Based on the mean absolute error threshold θ, α ma is determined as follows: When α ma ≥ θ, it means that the current pipeline data prediction model needs to be updated; When α ma < θ, it means that the current pipeline data prediction model does not need to be updated and continues to be used; For α ma When ≥ θ, collect the new data since the collection time t ma Evenly divide it again based on N equal parts, retrain the LSTM model using the updated data set, and replace the old model with the new prediction model as the data prediction model for the current pipeline.
5. The intelligent operation and maintenance system for gathering and transportation according to the method for filling missing data as claimed in claim 4, wherein: In the real-time prediction module, a same-day anomaly detection module is also set up. Based on the same-type data of pipelines within the same group, real-time anomaly detection is performed at every 2T moment. The specific steps are as follows: Collect the real-time pressure, temperature, flow rate, and water cut data per minute before the 2T moment of the same day from all pipeline files within the same group, and divide the original data set based on the data type respectively. Analyze the anomaly parameters of each data set through the isolation forest method. The specific steps are as follows: Randomly extract a sample subset from the original data set, with a size of ψ. Starting from the root node of each tree, randomly select a feature and a random splitting value on this feature to divide the data set into left and right child nodes. Recursively repeat the above process for each child node until the preset tree depth is reached or there is only one data point in the child node; For each data point x, calculate its path length h(x) in each tree, and average the path lengths over all trees to obtain the average path length Calculate the anomaly score s(x,u), and its algorithm formula is as follows: Among them, u represents the total number of samples, that is, the number of all data points participating in the anomaly detection, c(u) is the sample constant, H(i) represents the i-th harmonic number. When s(x, u) is greater than the detection threshold, it means that there is an anomaly in the pipeline data points collected on the same day. Mark the abnormal pipelines in the same group file, and separately isolate the LSTM model of the abnormal data type of the abnormal pipelines, and do not transmit the output value to the subsequent modules.
6. The intelligent operation and maintenance system for gathering and transportation according to the missing data filling method as claimed in claim 1, wherein: The gathering and grouping module groups the daily operating pipelines based on the collected current oilfield gathering and transportation pipeline specifications and start-stop status, establishes a daily operating pipeline database based on MySQL, groups the daily enabled pipelines based on the pipeline specifications, divides the same-type daily operating pipelines into the same file, and records the real-time gathering and transportation data collected by the data acquisition module. The specific steps are as follows: Based on the collected current oilfield gathering and transportation pipeline specifications and the daily oilfield gathering and transportation pipeline task list, establish files for the enabled oilfield gathering and transportation pipelines, and store the pipeline specifications in the files, including pipeline diameter and pipeline material; For the pipelines operating on the same day, group the pipelines with the same diameter and material into one group, establish a daily operating pipeline database based on MySQL, and put the pipeline files grouped into one group into the same group in the daily operating pipeline database; Receive the real-time pipeline data collected by the sensors in the data acquisition module and record it in the files of the corresponding pipelines.
7. An intelligent operation and maintenance device for gathering and transportation based on a missing data filling method, characterized in that, The device is applied to a gathering and transportation intelligent operation and maintenance system based on a missing data filling method, and includes a memory and a processor: The memory is used for non-temporary storage of computer-readable instructions; The processor is used to run the computer-readable instructions; Among them, when the computer-readable instructions are run by the processor, the steps in the system according to any one of claims 1-6 are executed.
Citation Information
Patent Citations
Energy management system data filling method and architecture based on time sequence analysis
CN119377203A
Industrial equipment data intelligent management method based on big data algorithm
CN119577681A